ML News
new
|
past
|
best
|
rss
|
submit
about
Stories from August 19, 2026 (UTC)
Go back a
day
,
month
, or
year
. Go forward a
day
.
1.
Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
(
arxiv.org
)
316 points
by
nunodonato
46 days ago
|
hide
|
past
|
pdf
|
281 comments
2.
Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)
(
arxiv.org
)
66 points
by
florianherrengt
46 days ago
|
hide
|
past
|
pdf
|
38 comments
3.
Improving the Matrix Multiplication Exponent
(
arxiv.org
)
6 points
by
aaraujo002
46 days ago
|
hide
|
past
|
pdf
|
discuss
4.
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
(
arxiv.org
)
3 points
by
Anon84
46 days ago
|
hide
|
past
|
pdf
|
discuss
5.
Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
(
arxiv.org
)
3 points
by
tcp_handshaker
46 days ago
|
hide
|
past
|
pdf
|
discuss
6.
Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
(
arxiv.org
)
3 points
by
tcp_handshaker
46 days ago
|
hide
|
past
|
pdf
|
discuss
7.
Model Hypnosis: Strong control of AI via additive subliminal effects
(
arxiv.org
)
3 points
by
sbulaev
46 days ago
|
hide
|
past
|
pdf
|
discuss
8.
Stealing Reasoning Traces from Proprietary LLM APIs
(
arxiv.org
)
2 points
by
amai
46 days ago
|
hide
|
past
|
pdf
|
1 comment
9.
Sampling More, Getting Less: Calibration Is the Diversity Bottleneck in LLMs
(
arxiv.org
)
2 points
by
clukic
46 days ago
|
hide
|
past
|
pdf
|
discuss
10.
AutoResearch: Insight In, Hallucination Out
(
arxiv.org
)
2 points
by
johnbarron
46 days ago
|
hide
|
past
|
pdf
|
discuss
About
|
RSS
|
RSS (all)
|
HN arXiv