about
Do Chatbot LLMs Talk Too Much? (arxiv.org)
12 points by bretkoppel 44 days ago | hide | past | pdf | 4 comments on HN

In plain words: A test of 300-plus prompts where the ideal reply is short, measuring how many extra characters a chatbot adds beyond a minimal baseline answer. Across 76 chatbots, median extra length varied tenfold, with models padding ambiguous questions and adding explanations to one-line coding tasks.

Abstract · Do Chatbot LLMs Talk Too Much? The YapBench Benchmark

Large Language Models (LLMs) such as ChatGPT, Claude, and Gemini increasingly act as general-purpose copilots, yet they often respond with unnecessary length on simple requests, adding redundant explanations, hedging, or boilerplate that increases cognitive load and inflates token-based inference cost. Prior work suggests that preference-based post-training and LLM-judged evaluations can induce systematic length bias, where longer answers are rewarded even at comparable quality. We introduce YapBench, a lightweight benchmark for quantifying user-visible over-generation on brevity-ideal prompts. Each item consists of a single-turn prompt, a curated minimal-sufficient baseline answer, and a category label. Our primary metric, YapScore, measures excess response length beyond the baseline in characters, enabling comparisons across models without relying on any specific tokenizer. We summarize model performance via the YapIndex, a uniformly weighted average of category-level median YapScores. YapBench contains over three hundred English prompts spanning three common brevity-ideal settings: (A) minimal or ambiguous inputs where the ideal behavior is a short clarification, (B) closed-form factual questions with short stable answers, and (C) one-line coding tasks where a single command or snippet suffices. Evaluating 76 assistant LLMs, we observe an order-of-magnitude spread in median excess length and distinct category-specific failure modes, including vacuum-filling on ambiguous inputs and explanation or formatting overhead on one-line technical requests. We release the benchmark and maintain a live leaderboard for tracking verbosity behavior over time.

Vadim Borisov, Michael Gröger, Mina Mikhael, Richard H. Schreiber
arXiv:2601.00624 · cs.LG · submitted Jan 2, 2026
abstract · pdf · html

add comment on HN

They will keep a live leaderboard at https://huggingface.co/spaces/tabularisai/YapBench

Here is the top 10 right now:

  | Rank | Model                                    |   YapIndex | YapTax$ |
  | ---: | ---------------------------------------- | ---------: | ------: |
  |    1 | openai/gpt-5.6-sol (reasoning)           |  18.5 ±4.8 |    0.51 |
  |    2 | openai/gpt-5.6-sol                       |  19.2 ±4.4 |    0.51 |
  |    3 | openai/gpt-3.5-turbo                     |  22.7 ±4.8 |    0.02 |
  |    4 | openai/gpt-5.6-luna                      |  27.8 ±9.8 |    0.15 |
  |    5 | openai/gpt-5.4 (reasoning)               | 40.7 ±10.7 |       — |
  |    6 | openai/gpt-5.4                           |  40.7 ±9.0 |       — |
  |    7 | moonshotai/kimi-k2-0905                  |  44.7 ±4.8 |    0.05 |
  |    8 | mistralai/mistral-small-2603 (reasoning) | 46.2 ±31.5 |    0.03 |
  |    9 | openai/gpt-4                             | 51.2 ±20.6 |    1.39 |
  |   10 | openai/gpt-5.3-codex                     |  64.8 ±9.9 |       — |
I also checked how the Claude models did specifically:

  | Rank | Model                                   |    YapIndex | YapTax$ |
  | ---: | --------------------------------------- | ----------: | ------: |
  |   23 | anthropic/claude-opus-4.5               |  97.0 ±28.9 |    1.52 |
  |   25 | anthropic/claude-opus-4.5 (reasoning)   |  99.2 ±29.3 |    1.44 |
  |   42 | anthropic/claude-3.5-sonnet             | 199.7 ±24.5 |    2.53 |
  |   57 | anthropic/claude-sonnet-4.5 (reasoning) | 278.7 ±41.7 |    1.65 |
  |   61 | anthropic/claude-sonnet-4.5             | 285.0 ±39.4 |    1.63 |
  |   71 | anthropic/claude-opus-4.6               | 330.3 ±59.1 |       — |
  |   72 | anthropic/claude-haiku-4.5              | 333.2 ±26.7 |    0.64 |
  |   73 | anthropic/claude-haiku-4.5 (reasoning)  | 335.2 ±26.8 |    0.65 |
  |   76 | anthropic/claude-opus-4.6 (reasoning)   | 342.3 ±30.5 |       — |
  |   87 | anthropic/claude-3.5-haiku              | 401.2 ±25.4 |    0.59 |
  |   88 | anthropic/claude-sonnet-4.6 (reasoning) | 422.8 ±23.3 |       — |
  |   89 | anthropic/claude-sonnet-4.6             | 425.0 ±26.3 |       — |
No other Anthropic model seems to be there for now. I would have been very curious to see the more recent ones.
I needed this. I had no idea it existed.

I have bent over backwards trying to enforce brevity with deepseek v4 flash to the point where I think I broke some things trying to do prompt injection in my Hermes setup and was still unsuccessful.

Meanwhile Sol blows me away and I want that to be my default for everything now.

In general though I seem to have the most success with a "<=10w" requirement in my prompts.

What I don't see listed and would be a good comparison is the STS models. OpenAI's live model is an absolute joy to talk with.

This defies Betteridge's law of headlines.
Yes, yes they do