about
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings (arxiv.org)
6 points by warship on Feb 4, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Co-cited papers provide the similarity signal for training better vectors of biomedical text, while each layer gains several specialist sub-networks, one per field. One model covers many domains where standard versions handle just one, and specializing one layer captures 85% of the full gain.

Abstract

The advancement of transformer neural networks has significantly elevated the capabilities of sentence similarity models, but they still struggle with highly discriminative tasks and may produce sub-optimal representations of important documents like scientific literature. With the increased reliance on retrieval augmentation and search, representing diverse documents as concise and descriptive vectors is crucial. This paper improves upon the vectors embeddings of scientific text by assembling niche datasets using co-citations as a similarity metric, focusing on biomedical domains. We apply a novel Mixture of Experts (MoE) extension pipeline to pretrained BERT models, where every multi-layer perceptron section is enlarged and copied into multiple distinct experts. Our MoE variants perform well over $N$ scientific domains with $N$ dedicated experts, whereas standard BERT models excel in only one domain at a time. Notably, extending just a single transformer block to MoE captures 85% of the benefit seen from full MoE extension at every layer. This holds promise for versatile and efficient One-Size-Fits-All transformer networks for numerically representing diverse inputs. Our methodology marks advancements in representation learning and holds promise for enhancing vector database search and compilation.

Logan Hallee, Rohan Kapur, Arjun Patel, Jason P. Gleghorn, Bohdan Khomtchouk
arXiv:2401.15713 · cs.LG, cs.AI, cs.CL · submitted Jan 28, 2024 · updated Dec 17, 2024
abstract · pdf · html

add comment on HN

Vector embeddings are key in tech and science. Our research shows that Mixture of Experts (MoE) transformers beat traditional sentence similarity methods. A huge result. Why?

LLMs often misread nuances in scientific texts. Our scalable solution upgrades pre-trained language models to MoE versions, with expert groups matching multiple model performances across various benchmarks.

Focusing on the cardiovascular disease and chronic obstructive pulmonary disease subfields, we generated training datasets using co-citations (publications citing each other) as a similarity metric.

These scalable datasets required no manual labeling, yet took advantage of the expert collective intelligence of the scientific community (in terms of citation networks), allowing our models to outperform all other tested models.

We believe this new approach marks significant and timely advancements in scientific text classification approaches and holds promise for enhancing vector database tasks, which is an active area of research (and teaching) in our lab.

Congratulations to co-authors Rohan Kapur, Logan Hallee, and Arjun Patel for innovating with this tour-de-force framework in the post-GPT era we all live in today. We hope you may find our work useful in the context of your own AI/machine learning research.