about
Semantic Tokenizer for Enhanced Natural Language Processing (arxiv.org)
70 points by PaulHoule on Apr 27, 2023 | hide | past | pdf | 28 comments on HN

In plain words: A new tokenizer builds its vocabulary from word meanings and stems, so related word forms share pieces instead of being split blindly. It covers more than twice as many word forms, so models train faster and beat far larger ones on standard language tests.

Abstract

Traditionally, NLP performance improvement has been focused on improving models and increasing the number of model parameters. NLP vocabulary construction has remained focused on maximizing the number of words represented through subword regularization. We present a novel tokenizer that uses semantics to drive vocabulary construction. The tokenizer includes a trainer that uses stemming to enhance subword formation. Further optimizations and adaptations are implemented to minimize the number of words that cannot be encoded. The encoder is updated to integrate with the trainer. The tokenizer is implemented as a drop-in replacement for the SentencePiece tokenizer. The new tokenizer more than doubles the number of wordforms represented in the vocabulary. The enhanced vocabulary significantly improves NLP model convergence, and improves quality of word and sentence embeddings. Our experimental results show top performance on two Glue tasks using BERT-base, improving on models more than 50X in size.

Sandeep Mehta, Darpan Shah, Ravindra Kulkarni, Cornelia Caragea
arXiv:2304.12404 · cs.CL · submitted Apr 24, 2023
abstract · pdf · html

add comment on HN

I'm going to add a contrarian take here: this preprint is not a research paper. While it's nice to see that there is an improvement here on their one task, this is not "semantically" driven tokenization. It's morphologically driven. To be semantically driven, it would be reasonable to expect that synonyms would have similar representations. I got really excited from the title, and the content is a let-down.

The line of research here has been going on for 30+ years, from Michael Brent's work, to Linguistica, to Morfessor, and now several approaches to incorporate morphology into tokenizers. The stand-out example is [0]. This paper doesn't seem to acknowledge any of that intellectual legacy. It's not a _research_ paper.

I'm getting a bit tired of people putting their class projects or quick engineering projects on arXiv. I don't know why they're surfacing so high on HN either.

[0]: https://aclanthology.org/2021.acl-long.279/

If a transformer has a good "place" to assign meanings to I think it does a pretty good job of (1) discovering similar meanings in synonyms, (2) representing words differently based on context. That later one is a huge advance over word embeddings which I thought were holding progress back instead of advancing it.

You're right that what they are doing is morphological, not semantic, but it helps a lot. I would say that

   日本語
"Japanese Language" is a good token to apply embedding, attention, etc. to because it has a definite meaning to which the transformer can attach whatever syntax and semantics it learns in terms of activations. If BPE gives up and processes it as UTF-8 bytes

  e6 97 a5 e6 9c ac e8 aa 9e
there is no clear meaning for any one of those tokens, and the model is going to have to work a lot harder.
By your first paragraph's argument, the semantics are in the Transformer, not the tokenizer.

And yes, what they do helps on their two test tasks. I'm not disputing that. It's the fact that there's no scholarship here.

There are so many thousands of knobs to twiddle with in a model these days, and they went after one that's commonly regarded in the NLP community as the 'defect'—the only part of the model that's not end-to-end trained along with the rest. Which would be great, if they acknowledged it! But there's no citation to any tokenization literature beyond BPE or SentencePiece. The literature review is as superficial as what you could find in a blog.

There are certainly byte-level or character-level tokenizers (think about CANINE or ByT5), and we can argue back and forth about their data-hungriness or slow inference. It would be nice to give more helpful units to a Transformer, so it doesn't have to learn syllables (or even characters) all on its own. Rebracketing/incorrect segmentation is a problem! And these authors have clued into that, but so have several hundred (or thousand?) researchers they don't cite.

What I'm having trouble with is the notion that this paper uncovered some exciting, revelatory fact about tokenization. Yes, "Japanese Language" would be a reasonable semantic unit! But these authors didn't discover that fact. Nobody's questioning whether 'good tokenization is better than bad tokenization'. Tokenization has seen ongoing attention in NLP forever.

These authors tried one variant, compared it against a library default option (and nothing else), evaluated on one task, put a bit of marketing around it, and called it a day. In the NLP course I used to TA, this wouldn't even qualify as a complete final project for the course.

> I'm getting a bit tired of people putting their class projects or quick engineering projects on arXiv.

Whenever something becomes a status symbol there will be people willing to exploit it. Perhaps ArXiv should hire some volunteers to check for a minimum of quality before acceptance? (/s, in case it's not clear).

Anecdotally, the second worst paper I've ever read was hosted on ArXiv and presented in an NLP group as a possible breakthrough. Tearing it apart in front of the person presenting it was no fun.

> To be semantically driven, it would be reasonable to expect that synonyms would have similar representations.

How could a tokenizer do anything about that unless the synonyms actually share substrings? The vector embedding is learned, not part of the tokenizer.

It couldn't, which is why it's a good idea to avoid the word, "semantic".

The same problem also exists in the name, "Large Language Model". Sure, the content being modeled contains language, but the model itself is not specific or limited to language patterns. We ought to call them "Large Text Models"; or better yet, "Text Inference Models".

The words we use to describe software are very important: they inform goals and expectations. They define the context that software exists in.

I see our biggest mistake as calling these tools, "Artificial Intelligence". That title began as a goal and a category of work: it doesn't belong in the title or description of software unless that software has actually met the goal.

A morpheme is the smallest *meaningful* unit in a language though.
I was being generous - stemming is poor man's morphology. Empirically useful (ask the IR folks) but incredibly heuristic.
> I'm getting a bit tired of people putting their class projects or quick engineering projects on arXiv.

I got downvoted when I expressed similar opinion with regards to MiniGPT4. I guess HN crowd value usefulness more than real contribution.

This paper makes the case that we can get a big improvement in transformer performance on text by using smarter tokenizing. Byte pair encoding and similar techniques tend to find words and meaningful word fragments, but really words should be tokenized in a way that represents word morphology, so “magically” should be “magic”-“al”-“ly” or something like that. BPE can do that sometimes by accident but doing it consistently improves performance a lot.

One more small improvement that will boost future models.

Notably these guys have found semantic tokenization helps with embedding-based search

https://www.youtube.com/watch?v=BczDZ59seII

And what about phonetics? Wouldn't it be easier for the system if it doesn't have to figure it out by itself?
Exactly. But practically you have to trade one thing or another.

Before BPE I bailed on a project because the sponsor insisted on using word vectors and I thought "Look, the most important words in our documents will be out-of-dictionary and that's like playing chess down a queen, a rook and two pawns."

Once BPE and similar tokenizers came out now you could say that the model has a chance when it confronts out-of-dictionary situations which will always be important. This was critical to the success of transformers for text.

On the other hand there are many things wrong with tokenization for particular applications. If you want to handle Japanese text you'd think a word like 日本語 "Japanese" should be tokenized as a word or as 日本 + 語 ("japan" + "language")

A multilingual model however is very likely to tokenize those at the unicode character level so you don't even get 日 + 本 + 語 ("sun" + "origin" + "language") but might get underlying UTF-8 bytes like e6 + 97 + a5 + e6 + 9c + ac + e8 + aa + 9e which is just awful.

The trouble is an English language model doesn't want to waste a limited supply of tokens on other languages even though it should be able to handle a few foreign characters. A Japanese language model would clearly make different decisions, a model that supports a large number of languages is going to struggle to allocate tokens between them.

Why is the supply of tokens limited? Are they currently represented as 16 bit unsigned ints (I hear vocab size of about 50k for GPT3)? If so, is there a performance penalty for going to u32 beyond the extra memory?
Tokenization use one-hot encodings, so that matrix will always be n^2 the number of tokens. This has an impact on all the subsequent layers and final number of parameters. You want to use as information dense tokens as possible, while being able to represent weird or unseen tokens, but discrete enough to allow differentiation of concepts.
I would actually be less worried about a sequence of raw bytes than the tokens generated by BPE. If "Japan" is 01 and "language" is 02, then "Japanese" will probably be 03, which has no connection at all to 01 or 02. But raw, verbose encoding slows down convergence at the beginning. (Well, at least in English it does.)
I’m glad someone’s tackling this. BPE seemed like such a simplistic approach for something with such a big effect on how your model "sees" the data. For larger models with tons of data it might matter less, since the model will eventually see that similar words are used similarly even if they are represented with different tokens. But this should help models converge faster at first - and if you have a small amount of data, it will affect your final result more.

I always wondered if languages with more regular spelling and conjugation rules would converge faster, or if languages like Chinese might be more efficient since they can get more semantic meaning into a pair of bytes than English.

Also, the technique in the paper could be extended to irregular forms with a custom decoder. E.g. encode "mouse" + "##plural", then decode that sequence to "mice".

Yup. There are so many reasons why BPE is an issue:

1. It really hurts the ability to generate creative writing/poetry (e.g. impossible for even ChatGPT-4 to fully understand syllable counts, leading to incorrect haikus and even poor rhyming), see https://gwern.net/gpt-3 and https://paperswithcode.com/paper/most-language-models-can-be...

2. It means that silly stuff like "Glitch Tokens" are a huge issue. i.e. whole tokens dedicated to weird usernames caused by people counting on a counting subreddit so much that their names got their own token. See https://www.youtube.com/watch?v=WO2X3oZEJOA

3. BPE has a lot of terrible vocabulary choices independent of Glitch Tokens. Lots of massive punctuation, garbage sequences, etc.

1 could be fixed by using a tokenizer that only splits at syllable boundaries, but 2 and 3 are independent of the specifics of the tokenizer, and switching out BPE for something else won't help much.

Glitch tokens are glitchy because they were not present in the training data for the embeddings, so they're essentially uninitialized. The real issue was not the use of BPE, but that they didn't restrict the tokenizer to only output tokens that the model was trained on.

Similarly, garbage sequences are present in the tokenizer vocabulary because those sequences were common in the data that the tokenizer was trained on. The solution is to not include garbage in your training data, though admittedly that's a tad tough when you're starting out with an uncurated dump of random internet content.

> 1 could be fixed by using a tokenizer that only splits at syllable boundaries

Maybe? Wouldn't that require the tokenizer to understand the word's usage in context of the surrounding text, or even worse, dialects?

What do you do with "fire" or "laboratory" or, god help us, "nuclear"?

Yes, you will need to write the tokenizer to behave in the dialect you are trying to write poetry in. An alternate approach is to filter on the output side by writing a sampler that only allows acceptable outputs. I think it would be easier to handle the ambiguous ones that way, because instead of encoding the ambiguity directly in the input or output stream, the sampler can just allow either pronunciation. https://www.youtube.com/watch?v=92hyQJRcFpM https://github.com/Hellisotherpeople/Constrained-Text-Genera...
I know CTGS is a few months old, but I learned about it literally yesterday which is why custom decoders are on my mind.
That's a lot of words to not explain what the thing is.

Any example tokens?

See Table 4
Table 4's BERT/Semantic tokenizer examples are full of word splittings that only a chemist could love. The BERT/WordPiece splittings make more semantic sense to me.
The problem is that each prefix gets a separate token, so the LLM has to learn from context that they refer to the same thing. In the semantic encoding, all of the words start with the same token, so the LLM can learn that the words are related with a lot less data.
Your paragraph told me so much more than the abstract in the OP.
Next gen Transformer to me could be: Imagination is all you need.