ML News
new
|
past
|
best
|
rss
|
submit
about
1531.
Prompt Repetition Improves Non Reasoning LLM
(
arxiv.org
)
2 points
by
jdthedisciple
229 days ago
|
hide
|
past
|
pdf
|
discuss
1532.
GLM-5 Technical Report
(
arxiv.org
)
12 points
by
meetpateltech
229 days ago
|
hide
|
past
|
pdf
|
discuss
1533.
Training-Free Group Relative Policy Optimization
(
arxiv.org
)
1 point
by
readitalready
229 days ago
|
hide
|
past
|
pdf
|
discuss
1534.
Composition-RL: Compose Verifiable Prompts for Reinforcement Learning of LLMs
(
arxiv.org
)
3 points
by
gmays
229 days ago
|
hide
|
past
|
pdf
|
discuss
1535.
Randomness in Agentic Evals
(
arxiv.org
)
1 point
by
andre15silva
230 days ago
|
hide
|
past
|
pdf
|
discuss
1536.
Hunt Globally
(
arxiv.org
)
1 point
by
salkahfi
230 days ago
|
hide
|
past
|
pdf
|
discuss
1537.
Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises
(
arxiv.org
)
1 point
by
salkahfi
230 days ago
|
hide
|
past
|
pdf
|
discuss
1538.
Learning State-Tracking from Code Using Linear RNNs
(
arxiv.org
)
2 points
by
jul8234
230 days ago
|
hide
|
past
|
pdf
|
1 comment
1539.
A Survey of In-Context Reinforcement Learning
(
arxiv.org
)
2 points
by
handfuloflight
230 days ago
|
hide
|
past
|
pdf
|
discuss
1540.
Soft Contamination Means Benchmarks Test Shallow Generalization
(
arxiv.org
)
2 points
by
cjbarber
230 days ago
|
hide
|
past
|
pdf
|
1 comment
1541.
SkillsBench: Benchmarking how well agent skills work across diverse tasks
(
arxiv.org
)
364 points
by
mustaphah
230 days ago
|
hide
|
past
|
pdf
|
171 comments
1542.
Virtual Width Networks (VWN)
(
arxiv.org
)
9 points
by
tesserato
230 days ago
|
hide
|
past
|
pdf
|
discuss
1543.
CodeLogician: Neuro-symbolic reasoning for precise software analysis
(
arxiv.org
)
2 points
by
NTCTech
231 days ago
|
hide
|
past
|
pdf
|
1 comment
1544.
Delegated Agent Authorization Constrained to Semantic Task-to-Scope Matching
(
arxiv.org
)
1 point
by
mooreds
231 days ago
|
hide
|
past
|
pdf
|
discuss
1545.
Evaluating AGENTS.md: are they helpful for coding agents?
(
arxiv.org
)
232 points
by
mustaphah
231 days ago
|
hide
|
past
|
pdf
|
161 comments
1546.
Multi-Agent Teams Hold Experts Back
(
arxiv.org
)
1 point
by
fauigerzigerk
231 days ago
|
hide
|
past
|
pdf
|
discuss
1547.
Large Language Model Reasoning Failures
(
arxiv.org
)
1 point
by
kawera
231 days ago
|
hide
|
past
|
pdf
|
discuss
1548.
Towards Autonomous Mathematics Research
(
arxiv.org
)
107 points
by
gmays
232 days ago
|
hide
|
past
|
pdf
|
53 comments
1549.
Retrieval-Aware Distillation for Transformer-SSM Hybrids
(
arxiv.org
)
2 points
by
readitalready
232 days ago
|
hide
|
past
|
pdf
|
discuss
1550.
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
(
arxiv.org
)
2 points
by
mpweiher
233 days ago
|
hide
|
past
|
pdf
|
discuss
1551.
Towards Autonomous Mathematics Research (Google DeepMind)
(
arxiv.org
)
1 point
by
u1hcw9nx
233 days ago
|
hide
|
past
|
pdf
|
discuss
1552.
Remote Labor Index: Measuring AI Automation of Remote Work
(
arxiv.org
)
2 points
by
Leynos
233 days ago
|
hide
|
past
|
pdf
|
discuss
1553.
Generalized on-policy distillation with reward extrapolation
(
arxiv.org
)
3 points
by
fzliu
233 days ago
|
hide
|
past
|
pdf
|
discuss
1554.
Adversarial Patch: images that make classifiers ignore other items in a scene
(
arxiv.org
)
1 point
by
felineflock
234 days ago
|
hide
|
past
|
pdf
|
discuss
1555.
Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators
(
arxiv.org
)
1 point
by
PaulHoule
234 days ago
|
hide
|
past
|
pdf
|
discuss
1556.
Fine-Tuning GPT-5 for GPU Kernel Generation
(
arxiv.org
)
4 points
by
matt_d
234 days ago
|
hide
|
past
|
pdf
|
discuss
1557.
SWE-ContextBench: context learning benchmark in coding
(
arxiv.org
)
1 point
by
mustaphah
234 days ago
|
hide
|
past
|
pdf
|
discuss
1558.
LLMs exceed physicians on complex text-based differential diagnosis
(
arxiv.org
)
3 points
by
rippeltippel
234 days ago
|
hide
|
past
|
pdf
|
2 comments
1559.
Learning to Reason in 13 Parameters
(
arxiv.org
)
2 points
by
stared
234 days ago
|
hide
|
past
|
pdf
|
discuss
1560.
LLM Reasoning Failures
(
arxiv.org
)
1 point
by
gradus_ad
234 days ago
|
hide
|
past
|
pdf
|
discuss
More
About
|
RSS
|
RSS (all)
|
HN arXiv