In plain words: A small language model is trained to rewrite scientific sentences as nested JSON outlines, then another model rebuilds the sentence from the outline to see what survives. The rebuilt sentences matched the originals in meaning and wording, showing the nested format holds information well.
Abstract
This paper investigates whether structured representations can preserve the meaning of scientific sentences. To test this, a lightweight LLM is fine-tuned using a novel structural loss function to generate hierarchical JSON structures from sentences collected from scientific articles. These JSONs are then used by a generative model to reconstruct the original text. Comparing the original and reconstructed sentences using semantic and lexical similarity we show that hierarchical formats are capable of retaining information of scientific texts effectively.
Satya Sri Rajiteswari Nimmagadda, Ethan Young, Niladri Sengupta, Ananya Jana, Aniruddha Maiti
arXiv:2603.23532 · cs.CL, cs.AI · submitted Mar 8, 2026
abstract · pdf · html · accepted to 21th International Conference on Semantic Computing (IEEE ICSC 2026)