about
How Many Instruction Can LLMs Follow at Once? (arxiv.org)
11 points by djdistyl on Jul 16, 2025 | hide | past | pdf | discuss on HN

In plain words: A new test asks models to write a business report while including up to 500 specified keywords, checking how many they actually follow as the list grows. Even the best models hit only 68% at 500, and they favor instructions listed earlier.

Abstract · How Many Instructions Can LLMs Follow at Once?

Production-grade LLM systems require robust adherence to dozens or even hundreds of instructions simultaneously. However, the instruction-following capabilities of LLMs at high instruction densities have not yet been characterized, as existing benchmarks only evaluate models on tasks with a single or few instructions. We introduce IFScale, a simple benchmark of 500 keyword-inclusion instructions for a business report writing task to measure how instruction-following performance degrades as instruction density increases. We evaluate 20 state-of-the-art models across seven major providers and find that even the best frontier models only achieve 68% accuracy at the max density of 500 instructions. Our analysis reveals model size and reasoning capability to correlate with 3 distinct performance degradation patterns, bias towards earlier instructions, and distinct categories of instruction-following errors. Our insights can help inform design of instruction-dense prompts in real-world applications and highlight important performance-latency tradeoffs. We open-source the benchmark and all results for further analysis at https://distylai.github.io/IFScale.

Daniel Jaroslawicz, Brendan Whiting, Parth Shah, Karime Maamari
arXiv:2507.11538 · cs.AI · submitted Jul 15, 2025
abstract · pdf · html

add comment on HN