about
Rule-based NLP system beats LLM for analysis of psychiatric clinical notes (arxiv.org)
120 points by PaulHoule on Apr 4, 2024 | hide | past | pdf | 19 comments on HN

In plain words: Two tools were tested for pulling social support and isolation mentions from psychiatric notes: one uses word lists and rules, the other a language model. The rule-based tool scored higher (0.89 vs 0.65 at one hospital) because it followed the human labeling rules exactly.

Abstract · Extracting Social Support and Social Isolation Information from Clinical Psychiatry Notes: Comparing a Rule-based NLP System and a Large Language Model

Background: Social support (SS) and social isolation (SI) are social determinants of health (SDOH) associated with psychiatric outcomes. In electronic health records (EHRs), individual-level SS/SI is typically documented as narrative clinical notes rather than structured coded data. Natural language processing (NLP) algorithms can automate the otherwise labor-intensive process of data extraction. Data and Methods: Psychiatric encounter notes from Mount Sinai Health System (MSHS, n=300) and Weill Cornell Medicine (WCM, n=225) were annotated and established a gold standard corpus. A rule-based system (RBS) involving lexicons and a large language model (LLM) using FLAN-T5-XL were developed to identify mentions of SS and SI and their subcategories (e.g., social network, instrumental support, and loneliness). Results: For extracting SS/SI, the RBS obtained higher macro-averaged f-scores than the LLM at both MSHS (0.89 vs. 0.65) and WCM (0.85 vs. 0.82). For extracting subcategories, the RBS also outperformed the LLM at both MSHS (0.90 vs. 0.62) and WCM (0.82 vs. 0.81). Discussion and Conclusion: Unexpectedly, the RBS outperformed the LLMs across all metrics. Intensive review demonstrates that this finding is due to the divergent approach taken by the RBS and LLM. The RBS were designed and refined to follow the same specific rules as the gold standard annotations. Conversely, the LLM were more inclusive with categorization and conformed to common English-language understanding. Both approaches offer advantages and are made available open-source for future testing.

Braja Gopal Patra, Lauren A. Lepow, Praneet Kasi Reddy Jagadeesh Kumar, Veer Vekaria, Mohit Manoj Sharma, Prakash Adekkanattu, Brian Fennessy, Gavin Hynes, Isotta Landi, Jorge A. Sanchez-Ruiz, Euijung Ryu, Joanna M. Biernacka, et al.
arXiv:2403.17199 · cs.CL · submitted Mar 25, 2024
abstract · pdf · html · 2 figures, 3 tables

add comment on HN

This work is based on a fine-tuned Google Palm Model from 2022. I'm not sure if this is a fair comparison to the latest groundbreaking series of LLMs
In view of the fact that we have been experiencing a breakthrough in the public perception of LLMs for 1.5 years and that the resources for their further development have increased explosively, it is indeed questionable to publish this now in March '24. A review based on the latest, most powerful LLMs would be urgently needed.
FLAN-T5-XL, a 3B model, to be precise.
I would've expected so. Rule-based systems in NLP can express parts of speech, coreferences, modifiers…it's almost too easy to do something like "Extract all subject-verb-object clauses where the object is connected to the subject by, say, a possessive pronoun.
So you're thinking of expressions like "Winston said he had noticed a continued decrease in his suicidal ideation" (real example). I don't know if I'd agree that it's clear a rule-based system would be able to catch everything like this. Consider:

1a. (original) Winston said he had noticed a continued decrease in his suicidal ideation.

1b. (simplified) Winston noticed a decrease in his suicidal ideation.

2. Winston's suicidal ideation decreased.

3. There was a decrease in the suicidal ideation Winston felt.

4. Suicidal ideation decreased in the patient.

Many more syntactic repackagings of similar propositional content are possible, too. For perfect recall, rule-based systems need to grapple with this, and additionally, in a real setting you're probably getting predicted (rather than gold-standard) POS tags, relations, etc., adding additional noise.

My expectation (as someone who's mildly bearish on LLMs, btw) would be that if an LLM appears to be doing worse than a rule-based system then you probably haven't tried some low-hanging fruit yet such as adjusting prompting strategies, fine-tuning, using different pretraining data, etc.

Your 1a and 1b say completely different things, and 2, 3, and 4 are even further away.

In 1a, someone says something. In 4, you seem to have an objective fact about a patient.

They used Flan T5-XL, a 3B model from 2022. Very weird choice.
My computational linguistics professor in grad school (who went to do do NLP research at Google) always said, "Do the dumb thing first!"
You have to admit that expecting general LLMs to automatically be domain experts everywhere is kind of a dumb thing, though!
They beat a 2-year old 3B LLM. I bet even GPT3.5 will beat their model, even without RAG.
Any specifically intentioned and designed software should outperform a generalist approach to solving a solution set. The primary difference is in the level of brittleness between the two - I mean AlphaGO beat Lee Sedol at GO, but Lee has a much more advanced self-driving ability.
I think the more appropriate comparison would be between a rule-based system and an LLM specifically built on clinical data, e.g. Gatortron (https://arxiv.org/ftp/arxiv/papers/2203/2203.03540.pdf)
Of all the models to perform such a comparison, Flan-T5-XL, a 3B parameter model from 2022 is a genuinely baffling choice. The paper is nearly worthless for this reason alone.
It seems the authors implicitly chose the Flan-T5-XL model based on a previous paper they refer where it outperformed prompted engineered ChatGPT 3.5 and 4 on a slightly different but similar task. This should definitely need to be further explained and possibly confirmed on their tasks.

"Our best-performing fine-tuned models outperformed zero- and few-shot performance of ChatGPT-family models [...]" from Large language models to identify social determinants of health in electronic health records (https://arxiv.org/abs/2308.06354)

Flan-T5 is actually not a bad model, even today (for fine-tuning)
I don't think it's a bad model. It's just strange to use it as the ambassador for LLM performance as it's nowhere near SOTA.
That a rule-based system is able to outperform a generalist model on a highly-specialized task is not exactly surprising.

What is missing from the discussion (and this is no fault of the paper—this isn’t their focus) is how much cost and effort goes into building the specialist system, vs. using a generalist model.

Cost of building is higher. Cost of data processing probably much lower.
isn't this the exact usage of medpalm?