about
Neural Decipherment via Minimum-Cost Flow: From Ugaritic to Linear B (arxiv.org)
30 points by godelmachine on Oct 21, 2020 | hide | past | pdf | 13 comments on HN

In plain words: A system deciphers lost languages by learning how letters changed between related languages, then pairing words at lowest total cost instead of guessing matches. It beat the best earlier Ugaritic results by 5.5% and translated 67.3% of Linear B words related to ancient Greek.

Abstract · Neural Decipherment via Minimum-Cost Flow: from Ugaritic to Linear B

In this paper we propose a novel neural approach for automatic decipherment of lost languages. To compensate for the lack of strong supervision signal, our model design is informed by patterns in language change documented in historical linguistics. The model utilizes an expressive sequence-to-sequence model to capture character-level correspondences between cognates. To effectively train the model in an unsupervised manner, we innovate the training procedure by formalizing it as a minimum-cost flow problem. When applied to the decipherment of Ugaritic, we achieve a 5.5% absolute improvement over state-of-the-art results. We also report the first automatic results in deciphering Linear B, a syllabic language related to ancient Greek, where our model correctly translates 67.3% of cognates.

Jiaming Luo, Yuan Cao, Regina Barzilay
arXiv:1906.06718 · cs.CL · submitted Jun 16, 2019
abstract · pdf · html · Accepted by ACL 2019

add comment on HN
Also discussed: Jul 2019 (76 points, 7 comments) · Jul 2019 (2 points, 0 comments)

But no Linear-A? That would be the real test

https://en.wikipedia.org/wiki/Linear_A

I think it must have to be genetically related to known languages to predict cognates, etc. AFAIK, no one has too much of an idea what language family the Minoan language could be a part of.
They could run their method with different language families and see which one fits best.

AFAIK that is how the existing attempts to match Linear A to a language worked too. But perhaps the machine could do it better.

Wow, this is very cool. Any fellow readers have suggestions for looking at SOTA with respect to computational approaches to lost languages? Ever since I found etymonline.com and I've started tracing word etymologies back to PIE, I've been thinking that I can't get enough of this stuff.
As to whether this is SOTA or not, I will not say. It is certainly a controversial method. https://en.wikipedia.org/wiki/Mass_comparison

I will say that as a teenager, mass lexical comparison, in Merrit Ruhlen's book 'The Origin of Language: Tracing the Evolution of the Mother Tongue', was very compelling. Ruhlen worked closely with Greenberg, who was a famous historical/comparative linguist of African languages.

Ten or more years later, with the benefit of experience, I feel like aligning hundreds of languages in a common vector space (multilingual word embeddings) will probably answer these questions better. But I am not a comparative linguist.

Here is a candy store for etymology enthusiasts: https://starling.rinet.ru
Wow! This is super cool -- another sibling commenter mentioned a Starostin paper as well. Thanks a ton.
Warning: etymonline.org goes to a sketchy black hole. The correct URL is etymonline.com
Thank you! Edited.
It really is so cool to see how cognates in different languages evolved :)
This needs to applied to the Indus script, stat!
I was thinking Voynich.