about
Deep Learning for Classical Japanese Literature (arxiv.org)
1 point by hardmaru on Dec 8, 2018 | hide | past | pdf | discuss on HN

In plain words: A collection of scanned cursive Japanese characters, split into three sets from simple letters to rare complex ones, so machines can learn to read old handwriting. The sets grow from an easy starter set to harder ones, opening a path into classical Japanese texts.

Abstract

Much of machine learning research focuses on producing models which perform well on benchmark tasks, in turn improving our understanding of the challenges associated with those tasks. From the perspective of ML researchers, the content of the task itself is largely irrelevant, and thus there have increasingly been calls for benchmark tasks to more heavily focus on problems which are of social or cultural relevance. In this work, we introduce Kuzushiji-MNIST, a dataset which focuses on Kuzushiji (cursive Japanese), as well as two larger, more challenging datasets, Kuzushiji-49 and Kuzushiji-Kanji. Through these datasets, we wish to engage the machine learning community into the world of classical Japanese literature. Dataset available at https://github.com/rois-codh/kmnist

Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, David Ha
arXiv:1812.01718 · cs.CV, cs.LG, stat.ML · submitted Dec 3, 2018
abstract · pdf · html · To appear at Neural Information Processing Systems 2018 Workshop on Machine Learning for Creativity and Design

add comment on HN
Also discussed: Jul 2019 (1 point, 0 comments)