about
R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging (arxiv.org)
4 points by PaulHoule 336 days ago | hide | past | pdf | discuss on HN

In plain words: Rule-to-Tag turns hand-written grammar rules into the training signal itself, letting a tagger learn from rules and unlabeled text with no labeled examples at all. On Zarma part-of-speech tagging it scored 0.968, nearly matching a supervised model trained on 300 labeled sentences (0.975).

Abstract · R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages

We introduce Rule-to-Tag (R2T), a framework that turns linguistic rules into the training signal for neural sequence taggers in low-resource languages. R2T encodes lexical, morphological, and syntactic rules as differentiable loss terms, so that a tagger learns from rules and unlabeled text, with no labeled training data. It also adds an out-of-vocabulary (OOV) loss term that discourages confident predictions on words that no rule covers. R2T is a first instance of a broader paradigm we call principled learning (PrL): using explicit principles as a learning signal, rather than relying on example-based supervision alone. We evaluate R2T on part-of-speech (POS) tagging for Zarma (Songhay), Bambara (Mande), and French (Romance), and on named entity recognition (NER) for Zarma. On Zarma POS tagging, R2T-BiLSTM uses no labeled training data, yet it reaches 0.968 Macro F1. This score comes within 0.007 of a supervised BiLSTM-CRF trained on 300 labeled sentences (0.975) and exceeds AfriBERTa fine-tuned on the same sentences (0.941). On NER, R2T works well as pre-training: after R2T pre-training, a model fine-tuned on 50 labeled sentences outperforms AfriBERTa fine-tuned on 300 (0.83 vs.\ 0.79 span F1). On Bambara, R2T reaches 0.91 Macro F1 with rules written in about 2.75 hours, while a supervised Masakhane tagger reaches 0.78. We release ZarmaPOS-Bench, a silver-standard Zarma POS corpus, together with ZarmaNER-600 and all trained models, to support future work on under-resourced languages.

Mamadou K. Keita, Christopher Homan, Sebastien Diarra
arXiv:2510.13854 · cs.CL, cs.LG · submitted Oct 12, 2025 · updated Sep 27, 2026
abstract · pdf · html

add comment on HN