about
Stories from August 19, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025) (arxiv.org)
316 points by nunodonato 46 days ago | hide | past | pdf | 281 comments
2. Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025) (arxiv.org)
66 points by florianherrengt 46 days ago | hide | past | pdf | 38 comments
3. Improving the Matrix Multiplication Exponent (arxiv.org)
6 points by aaraujo002 46 days ago | hide | past | pdf | discuss
4. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arxiv.org)
3 points by Anon84 46 days ago | hide | past | pdf | discuss
5. Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows (arxiv.org)
3 points by tcp_handshaker 46 days ago | hide | past | pdf | discuss
6. Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification (arxiv.org)
3 points by tcp_handshaker 46 days ago | hide | past | pdf | discuss
7. Model Hypnosis: Strong control of AI via additive subliminal effects (arxiv.org)
3 points by sbulaev 46 days ago | hide | past | pdf | discuss
8. Stealing Reasoning Traces from Proprietary LLM APIs (arxiv.org)
2 points by amai 46 days ago | hide | past | pdf | 1 comment
9. Sampling More, Getting Less: Calibration Is the Diversity Bottleneck in LLMs (arxiv.org)
2 points by clukic 46 days ago | hide | past | pdf | discuss
10. AutoResearch: Insight In, Hallucination Out (arxiv.org)
2 points by johnbarron 46 days ago | hide | past | pdf | discuss