about
DeepStyle: User Style Embedding for Authorship Attribution of Short Texts (arxiv.org)
1 point by alimov on Apr 29, 2022 | hide | past | pdf | 3 comments on HN

In plain words: It learns a compact picture of each writer's distinctive habits from their posts, instead of sorting texts by a single feature like word pairs. On Twitter and Weibo posts, it identified authors more accurately than the best earlier tools.

Abstract

Authorship attribution (AA), which is the task of finding the owner of a given text, is an important and widely studied research topic with many applications. Recent works have shown that deep learning methods could achieve significant accuracy improvement for the AA task. Nevertheless, most of these proposed methods represent user posts using a single type of feature (e.g., word bi-grams) and adopt a text classification approach to address the task. Furthermore, these methods offer very limited explainability of the AA results. In this paper, we address these limitations by proposing DeepStyle, a novel embedding-based framework that learns the representations of users' salient writing styles. We conduct extensive experiments on two real-world datasets from Twitter and Weibo. Our experiment results show that DeepStyle outperforms the state-of-the-art baselines on the AA task.

Zhiqiang Hu, Roy Ka-Wei Lee, Lei Wang, Ee-Peng Lim, Bo Dai
arXiv:2103.11798 · cs.CL, cs.SI · submitted Mar 14, 2021
abstract · pdf · html · Paper accepted for 4th APWeb-WAIM Joint Conference on Web and Big Data

add comment on HN

Some years ago J.K Rowlings pseudonym was discovered via “Forensic Linguistics” (see: Smithsonian Magazine - “How Did Computers Uncover J.K. Rowling’s Pseudonym?”). I have never met anyone doing this kind of work as their day job or hobby, but would really appreciate any additional related reading material that covers Authorship Attribution and the use of NLP.
Some years ago J.K Rowlings pseudonym was discovered via “Forensic Linguistics”

It was discovered because Rowlings' lawyer's wife told a friend, who told a journalist, who leaked it. The forensic linguistics was only used after the fact to try to confirm the story.

Ah, that’s certainly less exciting. I hadn’t known that bit.