about
Modeling and Solving Operations Research Problems with Tool Augmented LLMs (arxiv.org)
2 points by PaulHoule 342 days ago | hide | past | pdf | discuss on HN

In plain words: A small open model is trained on generated planning puzzles and taught to hand the math to a standard solver, so private data stays off closed services. It answered up to 80.1% of standard test problems correctly, beating equally sized models.

Abstract · OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models

Large language models (LLMs) demonstrate strong mathematical reasoning, but reliance on closed-source APIs for OR tasks raises privacy concerns, and training open-source models from scratch incurs high compute costs. We introduce OR-Toolformer, which fine-tunes Llama-3.1-8B-Instruct with a semi-automatic data synthesis pipeline that generates diverse OR problem-answer pairs and augments the model with external solvers to produce API calls. On three of four standard benchmarks, OR-Toolformer achieves up to 80.1% execution accuracy, exceeding size-matched baselines by over 4.3%. In zero-shot evaluation on two unseen OR problem types, it attains 54% average accuracy, a 21 percentage-point improvement over the strongest baseline. These findings validate the efficacy of tool-augmented fine-tuning LLMs for accurate and generalizable OR problem modeling and solving.

Jianzhang Zhang, Jialong Zhou, Chuang Liu
arXiv:2510.01253 · cs.AI, cs.LG · submitted Sep 24, 2025
abstract · pdf · html

add comment on HN