about
ChemBench: Evaluating LLMs Against Expert Chemists [New Results] (arxiv.org)
2 points by kjappelbaum on Nov 6, 2024 | hide | past | pdf | 1 comment on HN

In plain words: An automated test with over 2,700 chemistry questions compares AI models' knowledge and reasoning with expert human chemists. The best models beat the top humans on average, yet stumble on simple tasks and sound more confident than they should.

Abstract · Are large language models superhuman chemists?

Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. However, we possess only a limited systematic understanding of the chemical capabilities of LLMs, which would be required to improve models and mitigate potential harm. Here, we introduce "ChemBench," an automated framework for evaluating the chemical knowledge and reasoning abilities of state-of-the-art LLMs against the expertise of chemists. We curated more than 2,700 question-answer pairs, evaluated leading open- and closed-source LLMs, and found that the best models outperformed the best human chemists in our study on average. However, the models struggle with some basic tasks and provide overconfident predictions. These findings reveal LLMs' impressive chemical capabilities while emphasizing the need for further research to improve their safety and usefulness. They also suggest adapting chemistry education and show the value of benchmarking frameworks for evaluating LLMs in specific domains.

Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu, Martiño Ríos-García, Benedict Emoekabu, Aswanth Krishnan, Tanya Gupta, Mara Schilling-Wilhelmi, Macjonathan Okereke, Anagha Aneesh, Amir Mohammad Elahi, Mehrdad Asgari, et al.
arXiv:2404.01475 · cs.LG, cond-mat.mtrl-sci, cs.AI, physics.chem-ph · submitted Apr 1, 2024 · updated Nov 1, 2024
abstract · pdf · html

add comment on HN

We've released a new version of ChemBench (https://arxiv.org/abs/2404.01475), a framework measuring LLMs' chemistry capabilities against human experts.

We Evaluated 2,788 expert-curated chemistry questions across undergrad/grad topics and provide a human baseline from 19 chemistry experts (mostly MS/PhD level). Leading LLMs (like o1, Claude 3.5) significantly outperformed human experts on average

Performance varies by topic:

- Strong: General chemistry, calculations

- Weak: Safety/toxicity, analytical chemistry

Notable: Models excel at textbook problems but struggle with reasoning that requires to map text to reasoning about 2D/3D structures.

Interesting observations:

- Tool-augmented systems (with literature search) didn't improve performance much

- Open source models (Llama-3.1-405B) approaching closed source performance

- Models provide unreliable confidence estimates

- Performance correlates with model size but not molecular complexity

The results suggest both impressive capabilities and clear limitations of current LLMs in chemistry, with implications for chemistry education and scientific tooling.

Code/Data: https://github.com/lamalab-org/chem-bench Interactive results: https://chembench.org