about
ChemCrow: Augmenting large-language models with chemistry tools (arxiv.org)
30 points by famouswaffles on Apr 17, 2023 | hide | past | pdf | 18 comments on HN

In plain words: A chat model is wired to 18 chemistry tools, letting it look up facts, run calculations, and plan steps instead of guessing. Left alone, it made an insect repellent and three catalysts and proposed a new light-absorbing molecule, tasks plain chat models struggle with.

Abstract

Over the last decades, excellent computational chemistry tools have been developed. Integrating them into a single platform with enhanced accessibility could help reaching their full potential by overcoming steep learning curves. Recently, large-language models (LLMs) have shown strong performance in tasks across domains, but struggle with chemistry-related problems. Moreover, these models lack access to external knowledge sources, limiting their usefulness in scientific applications. In this study, we introduce ChemCrow, an LLM chemistry agent designed to accomplish tasks across organic synthesis, drug discovery, and materials design. By integrating 18 expert-designed tools, ChemCrow augments the LLM performance in chemistry, and new capabilities emerge. Our agent autonomously planned and executed the syntheses of an insect repellent, three organocatalysts, and guided the discovery of a novel chromophore. Our evaluation, including both LLM and expert assessments, demonstrates ChemCrow's effectiveness in automating a diverse set of chemical tasks. Surprisingly, we find that GPT-4 as an evaluator cannot distinguish between clearly wrong GPT-4 completions and Chemcrow's performance. Our work not only aids expert chemists and lowers barriers for non-experts, but also fosters scientific advancement by bridging the gap between experimental and computational chemistry.

Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, Philippe Schwaller
arXiv:2304.05376 · physics.chem-ph, stat.ML · submitted Apr 11, 2023 · updated Oct 2, 2023
abstract · pdf · html · Experimental results

add comment on HN

I know it's a picayune comment, but bad references bug me. The preprint mentions:

"Molecular similarity calculator The primary function of this tool is to evaluate the similarity between two molecules, utilizing the Tanimoto similarity measure 76 based on the ECFP2 molecular fingerprints 77 of both input molecules."

Molecular similarity is a particular area of expertise for me. I was curious what they cited. It's:

[76] TT, T. An elementary mathematical theory of classification and prediction; 1958

His full name is "Taffee Tadashi Tanimoto", with surname "Tanimoto". The 1958 paper says "T. T. Tanimoto", so that should be "Tanimoto, T. T." to match the citation style.

Also, it's highly unlikely they've read that paper. It's an internal IBM Technical Report which requires an ILL request, which is how I got it. Worldcat lists only 8 libraries with a copy.

If they have read it they would be unusual. In talking with people I found that most people cite it not because they read it but because other people cite it, so they see it as the appropriate citation.

The citation for the version published in the scientific literature is Rogers D.J., Tanimoto T.T. (October 1960). "A Computer Program for Classifying Plants". Science. 132 (3434): 1115–8. https://www.science.org/doi/10.1126/science.132.3434.1115

For a bit of science, task 4 in A.1.3 starts by telling "CN1CCC(CC1)=C1C2=C(SC=C2)C(=O)CC2=CC=CC=C12" is toxic.

PubChem tells me that structure is ketotifen, https://pubchem.ncbi.nlm.nih.gov/#query=CN1CCC(CC1)%3DC1C2%3... , and quotes DrugBank "In the US, it is now used in an over-the-counter ophthalmic formulation for the treatment of itchy eyes associated with allergies" . See also https://en.wikipedia.org/wiki/Ketotifen .

How is a program supposed to make it less toxic when it's already a widely used medicine?

There's perhaps an important issue - what does "toxic" mean? Something applied to your skin might be safe, but toxic if eaten in a large amount.

Task 5 in A.1.4 has the software respond that 20 grams of grain alcohol can be purchased from A2B Chem for 143 USD.

First, I found I couldn't find that entry in A2B Chem, but more importantly, that price seems rather high, yes?

Neither of which was caught in the student evaluations.

Surprisingly, we find that GPT-4 as an evaluator cannot distinguish between clearly wrong GPT-4 completions and GPT-4 + ChemCrow performance. There is a significant risk of misuse of tools like ChemCrow and we discuss their potential harms.

I think this inability to distinguish what is clearly wrong will turn out to be the Achilles heel of LLM's and will limit usage to non-critical "fun"/"entertainment" type use cases.

> I think this inability to distinguish what is clearly wrong will turn out to be the Achilles heel of LLM's and will limit usage to non-critical "fun"/"entertainment" type use cases.

GPT 4 was released for GA March 14th, 34 days ago. Yet ironically everyone is now a LLM expert making broad claims that it’s not good enough for this that or the other. The cognitive dissonance is staggering.

These aren't mystery systems. While the sheer volume of data makes individual outputs difficult to understand how they were derived, the systems operation are well understood. These are statistical models of language. They don't have the ability to reason or reflect. Hence bullshit. It'll be interesting to see if the work around augmenting/plugins is able to reduce the amount of bullshitting, but I suspect the sophistication mis-match between LLMs and supervisory systems means there will always be a baseline level of bullshit.
>the systems operation are well understood

That's like saying human behavior is well understood because we know how neurons communicate signals. It's too low level to be useful, hence psychology.

>They don't have the ability to reason or reflect.

Yes they do

https://selfrefine.info/

https://arxiv.org/abs/2303.11366

Amazes me how much this is downvoted despite providing citations and being essentially correct (barring a tedious philosophical debate and assuming good faith) - and in a thread specifically calling out "armchair ML experts"!
so...can we split knowledge part and reasoning/improving/refining part, and make LLM to smaller one?
Theoretically...maybe. But even if you could genuinely separate reasoning from knowledge, we simply don't know how.
GPT2, GPT3, and ChatGPT had a bullshit problem. Turns out the bullshit problem is still in GPT4, but it just says the bullshit more elegantly.

I think it's safe to predict that ChatGPT5 will even more eloquently speak bullshit.

--------

This is useful for entertainment, creative writing, fiction, etc. Etc. But it's a critical flaw in other applications.

I don't have plus, but 3.5 is already pretty impressive at bullshitting

sometimes it tells me bullshit so plausible and so eloquently expressed you can't help but being impressed, despite knowing for a fact it's a pack of lies

they're also original lies you wouldn't find anywhere online

Yeh it’s shocking, the level of bullshit is almost at the level of my coworkers. I’m just waiting for the day we can get humans to stop hallucinating even after 40 years of deep learning.
I think the potential for generating elaborate urban legends of this technology is phenomenal. I wonder if this will become a game changer in the future on this feature alone.
I know someone who is a non native speaker based in another country who has struggled using Google translate to accomplish his job. With a specialized prompt he is able to translate from his native language into English with alacrity. This has revolutionized his job. The outputs capture the nuances of idiomatic language across both languages.

That’s not a toy. That’s a universal translator that’s eluded us for decades.

I would also note that the next obvious step for the LLM is to combine its semantic understanding with specialized models and expert systems and other ML to build ensembles like the one discussed here. Yes, LLM isn’t sufficient to do everything. But it’s sufficient to do the things we’ve been unable to do with these specialized systems. Layering in semantic reasoning systems, agent based goal finding and optimization, and other capabilities in a feedback loop with the LLM, will see remarkable capabilities unlock.

I’m always surprised how little imagination I see on H.N. at times, and there is a real impetus to call monumental breakthroughs nothingburgers.

I most certainly agree with your last point - I do think that many of us are underestimating the impact of LLMs due to certain issues.

That said, I think your example does highlight how the LLMs are most useful when the operator knows what they are doing. In this case your friend appears to be competent at their job - they just need help with translation and phrasing. If your friend also needed the LLM to do the problem solving for them, I'd be more skeptical about its use.

That’s why I think problem solving is better done by the 70 years of work on exceptional problem solving engines. Other than language LLMs also have amazing ability to synthesize semantics and generalize concepts, and restructuring into multiple types of output and receiving multiple types of input. LLMs are remarkable at many things, but not everything. It just happens I believe they’re remarkable at precisely the things our remarkable AI inventions of the past are not. The synthesis of the ensemble of reasoning work done over the last near century will yield a lot of the unrealized benefits of that work to now.
One interesting task for LLMs would be to have them construct so-called 'green chemistry' alternatives to common industrial synthesis pathways. The term should really be 'closed-loop synthesis' in which the only thing being produced is the desired product, and all waste molecules get recycled back into the starting products.

However, those would likely either be novel synthesis steps or novel ways of arranging known synthesis steps, and I'm not sure LLMs would be able to generate such pathways as there'd be no existing examples in the literature.

A similar problem would be the design of novel catalysts for industrial chemical processes that are more efficient and cheaper to manufacturer (and perhaps safer to handle) compared to existing ones.

I'm building a competing system like this right now as we speak for my side project https://atomictessellator.com

It's really fun, although running some of these LLMs in my home lab is really hard, I have 4TB of ram here across 3 machines, and 12TB of storage, and running Galactica120B is particularly challenging.

Anyone who wants to start with chain-of-thought reasoning + different types of memory, a good place to start is here: https://python.langchain.com/

I am probably not the target audience, but it would be cool to have some more descriptive text there, and an about or FAQ page to tell more about what it can do and what it knows or doesn’t know in the databases it has access to