about
Bad characters: imperceptible Natural Language Processing attacks [pdf] (arxiv.org)
35 points by giuliomagnifico on Dec 30, 2021 | hide | past | pdf | 6 comments on HN

In plain words: Attackers hide invisible characters, look-alike letters, or tiny letter swaps in text so it looks normal to people but confuses text-processing systems. One hidden tweak already hurt performance noticeably, and three were enough to break most tested systems, including commercial search and translation tools.

Abstract · Bad Characters: Imperceptible NLP Attacks

Several years of research have shown that machine-learning systems are vulnerable to adversarial examples, both in theory and in practice. Until now, such attacks have primarily targeted visual models, exploiting the gap between human and machine perception. Although text-based models have also been attacked with adversarial examples, such attacks struggled to preserve semantic meaning and indistinguishability. In this paper, we explore a large class of adversarial examples that can be used to attack text-based models in a black-box setting without making any human-perceptible visual modification to inputs. We use encoding-specific perturbations that are imperceptible to the human eye to manipulate the outputs of a wide range of Natural Language Processing (NLP) systems from neural machine-translation pipelines to web search engines. We find that with a single imperceptible encoding injection -- representing one invisible character, homoglyph, reordering, or deletion -- an attacker can significantly reduce the performance of vulnerable models, and with three injections most models can be functionally broken. Our attacks work against currently-deployed commercial systems, including those produced by Microsoft and Google, in addition to open source models published by Facebook, IBM, and HuggingFace. This novel series of attacks presents a significant threat to many language processing systems: an attacker can affect systems in a targeted manner without any assumptions about the underlying model. We conclude that text-based NLP systems require careful input sanitization, just like conventional applications, and that given such systems are now being deployed rapidly at scale, the urgent attention of architects and operators is required.

Nicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas Papernot
arXiv:2106.09898 · cs.CL, cs.CR, cs.LG · submitted Jun 18, 2021 · updated Dec 11, 2021
abstract · pdf · html · To appear in the 43rd IEEE Symposium on Security and Privacy. Revisions: NER & sentiment analysis experiments, previous work comparison, defense evaluation

add comment on HN

I think they should write their next paper about adversarial attacks on ‘eval()’.

Models aren’t designed to understand Unicode. It’s the tokenizers job to chop that up and feed it to the model properly. What this person found is a step missing from the preprocessing pipelines of these models, or, more likely, something that should be part of the steps that come way before this data gets anywhere near an ML model. So, I think saying this is an adversarial attack on the machine learning model it’s self is a bit disingenuous, because for the input they were given, they did pretty well.

A similar example to this for a computer vision model would be something like inserting some kind of malformed payload in the headers of an image. The data is already bad long before it gets to the model.

A computer vision model could solve the homoglyph issue for the parser. Even humans are subject to this attack vector, though.
I was kind of excited reading the headline, but this seems mostly to be a pretty straight forward application of existing homoglyph attacks that have been common for decades at this point (albeit maybe not common against these types of programs previously).

I thought it was going to be adding imperceptible changes that allow you to choose the interpretation (sort of like the visual adversarial examples) not just substitute a homoglyph and the binary classifier no longer works.

Don't get me wrong, this is still important work, i just got a bit too excited.

It seems like most of these imperceptible changes could be addressed by something like ascii folding (https://www.elastic.co/guide/en/elasticsearch/reference/curr...) but this might not apply for non-english use cases.

If you're interested in adversarial NLP, I also recommend reading this blog post on adversarial attacks on GPT2 with universal triggers (e.g. adding "nobody" as prefix for all inputs causes all entailments to be predicted as contradiction).

Which browser add-ons and bookmarklets exist that help me conveniently obfuscate messages so as to avoid automated surveillance?
I guess you could use one of the utils here to make your text weird: (the type of text that causes issues of the type in the article)

https://lingojam.com/