In plain words: A language model trained on carefully filtered Dark Web text, so it understands the slang and messy structure of those sites. It beat general-purpose language models at Dark Web text tasks, making it a better tool for studying that hidden web.
Abstract
Recent research has suggested that there are clear differences in the language used in the Dark Web compared to that of the Surface Web. As studies on the Dark Web commonly require textual analysis of the domain, language models specific to the Dark Web may provide valuable insights to researchers. In this work, we introduce DarkBERT, a language model pretrained on Dark Web data. We describe the steps taken to filter and compile the text data used to train DarkBERT to combat the extreme lexical and structural diversity of the Dark Web that may be detrimental to building a proper representation of the domain. We evaluate DarkBERT and its vanilla counterpart along with other widely used language models to validate the benefits that a Dark Web domain specific model offers in various use cases. Our evaluations show that DarkBERT outperforms current language models and may serve as a valuable resource for future research on the Dark Web.
Youngjin Jin, Eugene Jang, Jian Cui, Jin-Woo Chung, Yongjae Lee, Seungwon Shin
arXiv:2305.08596 · cs.CL · submitted May 15, 2023 · updated May 18, 2023
abstract · pdf · html · 9 pages (main paper), 17 pages (including bibliography and appendix), to appear at the ACL 2023 Main Conference
1. The whole rationale of classifying a category of a site is dubious. More like an usual computational exercise.
2. The ransomware leak site detection is decent, but I'd expect HUMINT to perform much better for years to come.
3. The threat detection performance is close to random and hence useless (F1 around 53-54). The sample is also based on RaidForums alone and hence very small.
4. As they note, a restriction to English is a further limitation.
While I would have accepted the paper, I am not sure whether there is anything novel here.