In plain words: A chat model is wired to 18 chemistry tools, letting it look up facts, run calculations, and plan steps instead of guessing. Left alone, it made an insect repellent and three catalysts and proposed a new light-absorbing molecule, tasks plain chat models struggle with.
Abstract
Over the last decades, excellent computational chemistry tools have been developed. Integrating them into a single platform with enhanced accessibility could help reaching their full potential by overcoming steep learning curves. Recently, large-language models (LLMs) have shown strong performance in tasks across domains, but struggle with chemistry-related problems. Moreover, these models lack access to external knowledge sources, limiting their usefulness in scientific applications. In this study, we introduce ChemCrow, an LLM chemistry agent designed to accomplish tasks across organic synthesis, drug discovery, and materials design. By integrating 18 expert-designed tools, ChemCrow augments the LLM performance in chemistry, and new capabilities emerge. Our agent autonomously planned and executed the syntheses of an insect repellent, three organocatalysts, and guided the discovery of a novel chromophore. Our evaluation, including both LLM and expert assessments, demonstrates ChemCrow's effectiveness in automating a diverse set of chemical tasks. Surprisingly, we find that GPT-4 as an evaluator cannot distinguish between clearly wrong GPT-4 completions and Chemcrow's performance. Our work not only aids expert chemists and lowers barriers for non-experts, but also fosters scientific advancement by bridging the gap between experimental and computational chemistry.
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, Philippe Schwaller
arXiv:2304.05376 · physics.chem-ph, stat.ML · submitted Apr 11, 2023 · updated Oct 2, 2023
abstract · pdf · html · Experimental results
"Molecular similarity calculator The primary function of this tool is to evaluate the similarity between two molecules, utilizing the Tanimoto similarity measure 76 based on the ECFP2 molecular fingerprints 77 of both input molecules."
Molecular similarity is a particular area of expertise for me. I was curious what they cited. It's:
[76] TT, T. An elementary mathematical theory of classification and prediction; 1958
His full name is "Taffee Tadashi Tanimoto", with surname "Tanimoto". The 1958 paper says "T. T. Tanimoto", so that should be "Tanimoto, T. T." to match the citation style.
Also, it's highly unlikely they've read that paper. It's an internal IBM Technical Report which requires an ILL request, which is how I got it. Worldcat lists only 8 libraries with a copy.
If they have read it they would be unusual. In talking with people I found that most people cite it not because they read it but because other people cite it, so they see it as the appropriate citation.
The citation for the version published in the scientific literature is Rogers D.J., Tanimoto T.T. (October 1960). "A Computer Program for Classifying Plants". Science. 132 (3434): 1115–8. https://www.science.org/doi/10.1126/science.132.3434.1115
For a bit of science, task 4 in A.1.3 starts by telling "CN1CCC(CC1)=C1C2=C(SC=C2)C(=O)CC2=CC=CC=C12" is toxic.
PubChem tells me that structure is ketotifen, https://pubchem.ncbi.nlm.nih.gov/#query=CN1CCC(CC1)%3DC1C2%3... , and quotes DrugBank "In the US, it is now used in an over-the-counter ophthalmic formulation for the treatment of itchy eyes associated with allergies" . See also https://en.wikipedia.org/wiki/Ketotifen .
How is a program supposed to make it less toxic when it's already a widely used medicine?
There's perhaps an important issue - what does "toxic" mean? Something applied to your skin might be safe, but toxic if eaten in a large amount.
Task 5 in A.1.4 has the software respond that 20 grams of grain alcohol can be purchased from A2B Chem for 143 USD.
First, I found I couldn't find that entry in A2B Chem, but more importantly, that price seems rather high, yes?
Neither of which was caught in the student evaluations.