In plain words: Instead of picking among many entity types, it asks one yes-or-no question per word, learning from many types to spot new ones. With no examples it scored 35% on the usual score, beating earlier systems and matching GPT-3 while 1000 times smaller.
Abstract · From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
Supervised named entity recognition (NER) in the biomedical domain depends on large sets of annotated texts with the given named entities. The creation of such datasets can be time-consuming and expensive, while extraction of new entities requires additional annotation tasks and retraining the model. To address these challenges, this paper proposes a method for zero- and few-shot NER in the biomedical domain. The method is based on transforming the task of multi-class token classification into binary token classification and pre-training on a large amount of datasets and biomedical entities, which allow the model to learn semantic relations between the given and potentially novel named entity labels. We have achieved average F1 scores of 35.44% for zero-shot NER, 50.10% for one-shot NER, 69.94% for 10-shot NER, and 79.51% for 100-shot NER on 9 diverse evaluated biomedical entities with fine-tuned PubMedBERT-based model. The results demonstrate the effectiveness of the proposed method for recognizing new biomedical entities with no or limited number of examples, outperforming previous transformer-based methods, and being comparable to GPT3-based models using models with over 1000 times fewer parameters. We make models and developed code publicly available.
Miloš Košprdić, Nikola Prodanović, Adela Ljajić, Bojana Bašaragin, Nikola Milošević
arXiv:2305.04928 · cs.CL, cs.AI, cs.IR · submitted May 5, 2023 · updated Aug 25, 2024
abstract · pdf · html · Collaboration between Bayer Pharma R&D and Serbian Institute for Artificial Intelligence Research and Development. Artificial Intelligence in Medicine (2024)
This paper is extremely similar in domain (same corpus, etc - but that's not surprising since everyone uses these) but they're leaning heavily on the pretraining allowing capabilities towards few and zero shot, which is already well understood. Ultimately I think it's a good resource to use for the code, and if the API ends up being easier to change some of the internals on compared to the many options out there such as scispacy, or any of the pipelines used to achieve Pubtator, then it's a welcome addition.
My assessment is that this is a useful alternative where there are many solutions, but mostly an engineering product, and quite far away from any scientific contribution.