Abstract
Abstract Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinders progress. We present LLM4Mat-Bench, the largest benchmark to date for evaluating the performance of LLMs in predicting the properties of crystalline materials. LLM4Mat-Bench contains about 1.9 M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7 M, 615.5 M, and 3.1B tokens in total for each modality, respectively. We use LLM4Mat-Bench to fine-tune models with different sizes, including LLM-Prop and MatBERT, and provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction7 7 The Benchmark and code can be found at: https://github.com/vertaix/LLM4Mat-Bench. .
| Original language | English (US) |
|---|---|
| Article number | 020501 |
| Journal | Machine Learning: Science and Technology |
| Volume | 6 |
| Issue number | 2 |
| DOIs | |
| State | Published - Jun 30 2025 |
All Science Journal Classification (ASJC) codes
- Software
- Human-Computer Interaction
- Artificial Intelligence
Keywords
- benchmarks
- crystalline materials
- large language models
- materials property prediction
Fingerprint
Dive into the research topics of 'LLM4Mat-bench: benchmarking large language models for materials property prediction'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver