ML News
new
|
past
|
best
|
rss
|
submit
about
511.
Every Time I Hire a Linguist, Inference Costs Go Down
(
arxiv.org
)
3 points
by
cwbuilds
66 days ago
|
hide
|
past
|
pdf
|
discuss
512.
CryptanalysisBench: Can LLMs Do Cryptanalysis?
(
arxiv.org
)
1 point
by
zdw
66 days ago
|
hide
|
past
|
pdf
|
discuss
513.
Teaching agents to predict and pre-execute their next tool call
(
arxiv.org
)
6 points
by
rotariuvladimir
67 days ago
|
hide
|
past
|
pdf
|
discuss
514.
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
(
arxiv.org
)
3 points
by
root-parent
67 days ago
|
hide
|
past
|
pdf
|
discuss
515.
Visual prompt engineering for video models
(
arxiv.org
)
4 points
by
root-parent
67 days ago
|
hide
|
past
|
pdf
|
discuss
516.
Handbook.md shows that long policy documents do not reliably govern agents
(
arxiv.org
)
325 points
by
spIrr
67 days ago
|
hide
|
past
|
pdf
|
209 comments
517.
Detecting CSAM Text-to-Image LoRAs from Weights
(
arxiv.org
)
6 points
by
sbulaev
67 days ago
|
hide
|
past
|
pdf
|
discuss
518.
Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
(
arxiv.org
)
4 points
by
tcp_handshaker
67 days ago
|
hide
|
past
|
pdf
|
discuss
519.
Adaptive Agentic Attacks on LLM Vulnerability Detectors via Adversarial Comments
(
arxiv.org
)
3 points
by
tcp_handshaker
67 days ago
|
hide
|
past
|
pdf
|
discuss
520.
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
(
arxiv.org
)
3 points
by
frozenseven
67 days ago
|
hide
|
past
|
pdf
|
1 comment
521.
Mapping CVEs to Mitre ATT&CK Techniques
(
arxiv.org
)
3 points
by
adulau
67 days ago
|
hide
|
past
|
pdf
|
1 comment
522.
Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model
(
arxiv.org
)
3 points
by
sbulaev
67 days ago
|
hide
|
past
|
pdf
|
discuss
523.
The SpiNNaker2 chip: a many-core platform for brain-inspired computing
(
arxiv.org
)
1 point
by
Jimmc414
67 days ago
|
hide
|
past
|
pdf
|
discuss
524.
Distilling proprietary model reasoning into open-source search agents
(
arxiv.org
)
5 points
by
cpard
67 days ago
|
hide
|
past
|
pdf
|
discuss
525.
CryptanalysisBench: Can LLMs Do Cryptanalysis?
(
arxiv.org
)
2 points
by
rvz
67 days ago
|
hide
|
past
|
pdf
|
discuss
526.
"Uncensored" open LLMs are measurably more optimistic than their base models
(
arxiv.org
)
43 points
by
oleczek
67 days ago
|
hide
|
past
|
pdf
|
22 comments
527.
Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
(
arxiv.org
)
2 points
by
zhinit
67 days ago
|
hide
|
past
|
pdf
|
discuss
528.
Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
(
arxiv.org
)
306 points
by
ronfriedhaber
68 days ago
|
hide
|
past
|
pdf
|
133 comments
529.
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
(
arxiv.org
)
2 points
by
binyu
68 days ago
|
hide
|
past
|
pdf
|
discuss
530.
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language
(
arxiv.org
)
2 points
by
sbulaev
68 days ago
|
hide
|
past
|
pdf
|
discuss
531.
Neural Representation of Minimal Surfaces
(
arxiv.org
)
2 points
by
E-Reverance
68 days ago
|
hide
|
past
|
pdf
|
discuss
532.
Show HN: A 6M-token movable window on a single 46GB GPU
(
arxiv.org
)
7 points
by
Wetime
68 days ago
|
hide
|
past
|
pdf
|
16 comments
533.
Protocol-Level Attacks on Agentic Commerce Platforms: Taxonomy and Defense
(
arxiv.org
)
2 points
by
sbulaev
68 days ago
|
hide
|
past
|
pdf
|
discuss
534.
Lost in Context: Addressing Context Anxiety in Large Language Models
(
arxiv.org
)
2 points
by
StatsAreFun
68 days ago
|
hide
|
past
|
pdf
|
discuss
535.
The Polynomial-Time Low-Degree Conjecture Is False
(
arxiv.org
)
4 points
by
MarcoDewey
68 days ago
|
hide
|
past
|
pdf
|
discuss
536.
Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core
(
arxiv.org
)
2 points
by
Jimmc414
68 days ago
|
hide
|
past
|
pdf
|
discuss
537.
Paper – Agent Memory
(
arxiv.org
)
2 points
by
qspencer
69 days ago
|
hide
|
past
|
pdf
|
discuss
538.
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
(
arxiv.org
)
1 point
by
sbulaev
69 days ago
|
hide
|
past
|
pdf
|
discuss
539.
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic
(
arxiv.org
)
2 points
by
sbulaev
69 days ago
|
hide
|
past
|
pdf
|
discuss
540.
Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains
(
arxiv.org
)
2 points
by
MarcoDewey
69 days ago
|
hide
|
past
|
pdf
|
1 comment
More
About
|
RSS
|
RSS (all)
|
HN arXiv