about
LLMs for Materials and Chemistry: 34 Real-World Examples (arxiv.org)
15 points by yz-exodao on May 7, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A global hackathon produced 34 projects using AI text models across chemistry and materials research, from predicting molecular properties to reading papers and running lab workflows. The survey finds these tools help most in low-data and interdisciplinary work, though reliability and reproducibility remain unsolved.

Abstract · 34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery

Large Language Models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 34 total projects developed during the second annual Large Language Model Hackathon for Applications in Materials Science and Chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Yoel Zimmermann, Adib Bazgir, Alexander Al-Feghali, Mehrad Ansari, Joshua Bocarsly, L. Catherine Brinson, Yuan Chiang, Defne Circi, Min-Hsueh Chiu, Nathan Daelman, Matthew L. Evans, Abhijeet S. Gangan, et al.
arXiv:2505.03049 · cs.LG, cond-mat.mtrl-sci · submitted May 5, 2025 · updated May 16, 2025
abstract · pdf · html · arXiv admin note: substantial text overlap with arXiv:2411.15221. This paper is a refinement and analysis of the raw project submissions from arXiv:2411.15221

add comment on HN

This is a great resource, when I was in Saudi Arabia a few months ago at the global AI conference I gave a talk on AI in materials science - and while I was talking an LLM read a paper, created a hypothesis and used our API to create and run experiments - all live on our workbench behind me while I gave the talk.

LLMs aren't perfect by a long shot, but at the very least they can attempt to reproduce the paper, take a lot of the grunt work out of establishing a baseline, and let the scientist focus on making changes and pushing forwards with new experiments, rather than getting bogged down in replication of existing work.