In plain words: A small multilingual model turns text into compact number lists, so near-duplicate texts can be spotted even when lightly disguised or rewritten. It beat the usual fingerprint-based duplicate detector and other neural text matchers on deduplication, disguised near-duplicate retrieval, and spam grouping.
Abstract
This paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplicate text retrieval, clustering, and dataset deduplication tasks. We demonstrate that RETSim is significantly more robust and accurate than MinHash and neural text embeddings, achieving new state-of-the-art performance on dataset deduplication, adversarial text retrieval benchmarks, and spam clustering tasks. We also introduce the W4NT3D benchmark (Wiki-40B 4dversarial Near-T3xt Dataset) for evaluating multilingual, near-duplicate text retrieval capabilities under adversarial settings. RETSim and the W4NT3D benchmark are open-sourced under the MIT License at https://github.com/google/unisim.
Marina Zhang, Owen Vallis, Aysegul Bumin, Tanay Vakharia, Elie Bursztein
arXiv:2311.17264 · cs.CL · submitted Nov 28, 2023
abstract · pdf · html