ML News
new
|
past
|
best
|
rss
|
submit
about
1471.
Agents of Chaos
(
arxiv.org
)
3 points
by
wslh
222 days ago
|
hide
|
past
|
pdf
|
discuss
1472.
PersonaLive Expressive Portrait Image Animation for Live Streaming
(
arxiv.org
)
2 points
by
tamnd
223 days ago
|
hide
|
past
|
pdf
|
discuss
1473.
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
(
arxiv.org
)
2 points
by
yorwba
223 days ago
|
hide
|
past
|
pdf
|
discuss
1474.
Towards a Science of AI Agent Reliability
(
arxiv.org
)
2 points
by
smartmic
223 days ago
|
hide
|
past
|
pdf
|
discuss
1475.
Improving Chain-of-Thought Monitorability Through Information Theory
(
arxiv.org
)
1 point
by
simonpure
223 days ago
|
hide
|
past
|
pdf
|
discuss
1476.
The Landscape of Non-Equilibrium Memories with Neural Cellular Automata
(
arxiv.org
)
1 point
by
PaulHoule
223 days ago
|
hide
|
past
|
pdf
|
discuss
1477.
Fast and Optimal Mapping for Accelerator Modeling and Evaluation
(
arxiv.org
)
2 points
by
PaulHoule
223 days ago
|
hide
|
past
|
pdf
|
discuss
1478.
Agents of Chaos: Breaches of trust in autonomous LLM agents
(
arxiv.org
)
4 points
by
cool-RR
223 days ago
|
hide
|
past
|
pdf
|
1 comment
1479.
Computer-Using World Model
(
arxiv.org
)
1 point
by
steamboatwillie
223 days ago
|
hide
|
past
|
pdf
|
discuss
1480.
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
(
arxiv.org
)
1 point
by
PaulHoule
223 days ago
|
hide
|
past
|
pdf
|
discuss
1481.
The Statistical Signature of LLMs
(
arxiv.org
)
1 point
by
bikenaga
224 days ago
|
hide
|
past
|
pdf
|
discuss
1482.
Towards Compressive and Scalable Recurrent Memory
(
arxiv.org
)
2 points
by
PaulHoule
224 days ago
|
hide
|
past
|
pdf
|
discuss
1483.
A Visual Document Benchmark for Scientific Retrieval and Question Answering
(
arxiv.org
)
2 points
by
bobvanluijt
224 days ago
|
hide
|
past
|
pdf
|
discuss
1484.
Unified Latents (UL): How to train your latents
(
arxiv.org
)
2 points
by
pama
224 days ago
|
hide
|
past
|
pdf
|
discuss
1485.
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
(
arxiv.org
)
2 points
by
Gillesray
224 days ago
|
hide
|
past
|
pdf
|
1 comment
1486.
The Principles of Deep Learning Theory (2021)
(
arxiv.org
)
2 points
by
vinhnx
224 days ago
|
hide
|
past
|
pdf
|
discuss
1487.
HLE-Verified: A Verification and Revision of Humanity's Last Exam
(
arxiv.org
)
6 points
by
ravenical
224 days ago
|
hide
|
past
|
pdf
|
discuss
1488.
Realistic Adversarial Testing of Computer-Use Agents in Web-OS Environments
(
arxiv.org
)
2 points
by
yakkomajuri
224 days ago
|
hide
|
past
|
pdf
|
discuss
1489.
Large-scale online deanonymization with LLMs (including HN users)
(
arxiv.org
)
3 points
by
salkahfi
225 days ago
|
hide
|
past
|
pdf
|
2 comments
1490.
Stress Tests Reveal Fragile Grounding in Video-Language Models
(
arxiv.org
)
2 points
by
PaulHoule
225 days ago
|
hide
|
past
|
pdf
|
discuss
1491.
Surprising Effectiveness of Masking Updates in Adaptive Optimizers
(
arxiv.org
)
2 points
by
energy123
225 days ago
|
hide
|
past
|
pdf
|
discuss
1492.
Ferret-UI Lite – Apple on-device end-to-end GUI agent
(
arxiv.org
)
1 point
by
twalichiewicz
225 days ago
|
hide
|
past
|
pdf
|
discuss
1493.
What Language Is This? Ask Your Tokenizer
(
arxiv.org
)
2 points
by
tigoo
225 days ago
|
hide
|
past
|
pdf
|
discuss
1494.
Interactive Tools for Gaussian Splat Selection with AI and Human in the Loop
(
arxiv.org
)
2 points
by
PaulHoule
225 days ago
|
hide
|
past
|
pdf
|
discuss
1495.
End-to-End Test-Time Training for Long Context
(
arxiv.org
)
3 points
by
handfuloflight
225 days ago
|
hide
|
past
|
pdf
|
discuss
1496.
Large-scale online deanonymization with LLMs
(
arxiv.org
)
1 point
by
paulpauper
226 days ago
|
hide
|
past
|
pdf
|
discuss
1497.
Task Specific Knowledge Graphs
(
arxiv.org
)
1 point
by
alansaber
226 days ago
|
hide
|
past
|
pdf
|
discuss
1498.
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
(
arxiv.org
)
3 points
by
mpweiher
226 days ago
|
hide
|
past
|
pdf
|
discuss
1499.
Large Language Model Reasoning Failures
(
arxiv.org
)
40 points
by
T-A
226 days ago
|
hide
|
past
|
pdf
|
82 comments
1500.
The Fundamental Limits of LLMs at Scale
(
arxiv.org
)
4 points
by
o4c
226 days ago
|
hide
|
past
|
pdf
|
discuss
More
About
|
RSS
|
RSS (all)
|
HN arXiv