ML News
new
|
past
|
best
|
rss
|
submit
about
721.
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
(
arxiv.org
)
2 points
by
wek
99 days ago
|
hide
|
past
|
pdf
|
discuss
722.
Tapered Language Models
(
arxiv.org
)
2 points
by
sonabinu
99 days ago
|
hide
|
past
|
pdf
|
discuss
723.
Red-Teaming the Agentic Red-Team
(
arxiv.org
)
3 points
by
infwhispers
100 days ago
|
hide
|
past
|
pdf
|
discuss
724.
Combining LLMs Rarely Beats the Best Single Model, I tested 67 frontier models
(
arxiv.org
)
1 point
by
josefchen
100 days ago
|
hide
|
past
|
pdf
|
discuss
725.
Mapping Networks: CVPR 2026 Best Paper Award Nominee
(
arxiv.org
)
4 points
by
aurenvale
100 days ago
|
hide
|
past
|
pdf
|
1 comment
726.
A Structured Generation Framework for Transforming Scientific Papers into Patent
(
arxiv.org
)
2 points
by
teleforce
100 days ago
|
hide
|
past
|
pdf
|
discuss
727.
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Trac
(
arxiv.org
)
4 points
by
xiaoyu2006
100 days ago
|
hide
|
past
|
pdf
|
discuss
728.
Reading AI Model Compilation in MLIR Through the Lens of Formal Theories
(
arxiv.org
)
2 points
by
matt_d
100 days ago
|
hide
|
past
|
pdf
|
discuss
729.
David vs. Goliath in Next Activity Prediction: Argmax vs. LSTM, Transformer, LLM
(
arxiv.org
)
2 points
by
hramezani
101 days ago
|
hide
|
past
|
pdf
|
discuss
730.
Autodata: An agentic data scientist to create high quality synthetic data
(
arxiv.org
)
4 points
by
root-parent
101 days ago
|
hide
|
past
|
pdf
|
1 comment
731.
Wikipedia advocacy shapes LLM values
(
arxiv.org
)
3 points
by
50kIters
101 days ago
|
hide
|
past
|
pdf
|
discuss
732.
The False Promise of Imitating Proprietary LLMs
(
arxiv.org
)
2 points
by
handfuloflight
101 days ago
|
hide
|
past
|
pdf
|
discuss
733.
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
(
arxiv.org
)
3 points
by
NavinF
101 days ago
|
hide
|
past
|
pdf
|
discuss
734.
World Action Models: A Survey
(
arxiv.org
)
4 points
by
simonpure
101 days ago
|
hide
|
past
|
pdf
|
discuss
735.
Code as Agent Harness
(
arxiv.org
)
5 points
by
matt_d
101 days ago
|
hide
|
past
|
pdf
|
1 comment
736.
LLMs use "safety" specific neuron layers to identify vulnerabilities in code
(
arxiv.org
)
5 points
by
summarity
101 days ago
|
hide
|
past
|
pdf
|
3 comments
737.
Submodular Context Selection as a Pluggable Engine for LLM Agents
(
arxiv.org
)
2 points
by
Elof
102 days ago
|
hide
|
past
|
pdf
|
discuss
738.
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
(
arxiv.org
)
2 points
by
yogthos
102 days ago
|
hide
|
past
|
pdf
|
discuss
739.
The Promptware Kill Chain
(
arxiv.org
)
3 points
by
wslh
102 days ago
|
hide
|
past
|
pdf
|
discuss
740.
Qwen-AgentWorld: Language World Models for General Agents
(
arxiv.org
)
199 points
by
ilreb
102 days ago
|
hide
|
past
|
pdf
|
55 comments
741.
Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
(
arxiv.org
)
3 points
by
MediaSquirrel
103 days ago
|
hide
|
past
|
pdf
|
discuss
742.
Inference Compute Shapes Frontier LLM Evaluation
(
arxiv.org
)
2 points
by
matt_d
103 days ago
|
hide
|
past
|
pdf
|
discuss
743.
Concordia: JIT-Compiled Persistent-Kernel Checkpt for Fault-Tolerant Inference
(
arxiv.org
)
2 points
by
matt_d
103 days ago
|
hide
|
past
|
pdf
|
discuss
744.
Confidence estimation is a better metric than agreement for LLM judges
(
arxiv.org
)
3 points
by
rapiddev
103 days ago
|
hide
|
past
|
pdf
|
discuss
745.
PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations
(
arxiv.org
)
3 points
by
Anon84
103 days ago
|
hide
|
past
|
pdf
|
discuss
746.
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (2023)
(
arxiv.org
)
25 points
by
Anon84
103 days ago
|
hide
|
past
|
pdf
|
46 comments
747.
Unlimited OCR Works
(
arxiv.org
)
3 points
by
ilreb
103 days ago
|
hide
|
past
|
pdf
|
discuss
748.
Tapered Language Models
(
arxiv.org
)
3 points
by
E-Reverance
103 days ago
|
hide
|
past
|
pdf
|
discuss
749.
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
(
arxiv.org
)
55 points
by
teleforce
103 days ago
|
hide
|
past
|
pdf
|
9 comments
750.
VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
(
arxiv.org
)
398 points
by
timhigins
103 days ago
|
hide
|
past
|
pdf
|
205 comments
More
About
|
RSS
|
RSS (all)
|
HN arXiv