about
Jina Embeddings: A Novel Set of High-Performance Sentence Embedding Models (arxiv.org)
1 point by PaulHoule on Jul 27, 2023 | hide | past | pdf | discuss on HN

In plain words: These models turn sentences into numbers that capture meaning, for search and similarity, trained on carefully cleaned pairs and triplets of text. A new set of statements with and without 'not' teaches them to notice negation; they score well on a standard embedding test.

Abstract

Jina Embeddings constitutes a set of high-performance sentence embedding models adept at translating textual inputs into numerical representations, capturing the semantics of the text. These models excel in applications like dense retrieval and semantic textual similarity. This paper details the development of Jina Embeddings, starting with the creation of high-quality pairwise and triplet datasets. It underlines the crucial role of data cleaning in dataset preparation, offers in-depth insights into the model training process, and concludes with a comprehensive performance evaluation using the Massive Text Embedding Benchmark (MTEB). Furthermore, to increase the model's awareness of grammatical negation, we construct a novel training and evaluation dataset of negated and non-negated statements, which we make publicly available to the community.

Michael Günther, Louis Milliken, Jonathan Geuter, Georgios Mastrapas, Bo Wang, Han Xiao
arXiv:2307.11224 · cs.CL, cs.AI, cs.IR, cs.LG · submitted Jul 20, 2023 · updated Oct 20, 2023
abstract · pdf · html · 9 pages, 2 page appendix

add comment on HN