In plain words: Feeding a language model formal math definitions pulled from a structured knowledge base makes its answers more reliable. On math problems it scored better when the retrieved definitions fit the question, but worse when they did not.
Abstract · Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge
Language models exhibit fundamental limitations -- hallucination, brittleness, and lack of formal grounding -- that are particularly problematic in high-stakes specialist fields requiring verifiable reasoning. I investigate whether formal domain ontologies can enhance language model reliability through retrieval-augmented generation. Using mathematics as proof of concept, I implement a neuro-symbolic pipeline leveraging the OpenMath ontology with hybrid retrieval and cross-encoder reranking to inject relevant definitions into model prompts. Evaluation on the MATH benchmark with three open-source models reveals that ontology-guided context improves performance when retrieval quality is high, but irrelevant context actively degrades it -- highlighting both the promise and challenges of neuro-symbolic approaches.
Marcelo Labre
arXiv:2602.17826 · cs.AI, cs.LG, cs.SC · submitted Feb 19, 2026 · updated Aug 31, 2026
abstract · pdf · html · Supplementary materials and code: https://doi.org/10.5281/zenodo.18665030
Standard RAG retrieves semantic noise when confronted the logic of specialist domains like maths. I wanted to see if we could treat language models more like compilers by anchoring them to a structural ground truth.
I built a neuro-symbolic pipeline that grounds smaller models (<9B parameters, like Gemma2 and Qwen2.5-Math) in the OpenMath ontology using hybrid retrieval and cross-encoder reranking.
Evaluating on the MATH 500 benchmark revealed a severe bottleneck. When retrieval succeeds, reasoning and convergence improve. But the semantic gap between natural language and formal definitions is massive. When retrieval fails, the injected irrelevant ontological context actively degrades performance, hitting a hard "context utilization ceiling" in smaller models.
The paper and pipeline code are open. I ran these experiments locally without hyperscaler compute. I would love the community's technical feedback.
I’m continuing with the research now toward solving the retrieval quality bottleneck.