about
1531. Prompt Repetition Improves Non Reasoning LLM (arxiv.org)
2 points by jdthedisciple 229 days ago | hide | past | pdf | discuss
1532. GLM-5 Technical Report (arxiv.org)
12 points by meetpateltech 229 days ago | hide | past | pdf | discuss
1533. Training-Free Group Relative Policy Optimization (arxiv.org)
1 point by readitalready 229 days ago | hide | past | pdf | discuss
1534. Composition-RL: Compose Verifiable Prompts for Reinforcement Learning of LLMs (arxiv.org)
3 points by gmays 229 days ago | hide | past | pdf | discuss
1535. Randomness in Agentic Evals (arxiv.org)
1 point by andre15silva 230 days ago | hide | past | pdf | discuss
1536. Hunt Globally (arxiv.org)
1 point by salkahfi 230 days ago | hide | past | pdf | discuss
1537. Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises (arxiv.org)
1 point by salkahfi 230 days ago | hide | past | pdf | discuss
1538. Learning State-Tracking from Code Using Linear RNNs (arxiv.org)
2 points by jul8234 230 days ago | hide | past | pdf | 1 comment
1539. A Survey of In-Context Reinforcement Learning (arxiv.org)
2 points by handfuloflight 230 days ago | hide | past | pdf | discuss
1540. Soft Contamination Means Benchmarks Test Shallow Generalization (arxiv.org)
2 points by cjbarber 230 days ago | hide | past | pdf | 1 comment
1541. SkillsBench: Benchmarking how well agent skills work across diverse tasks (arxiv.org)
364 points by mustaphah 230 days ago | hide | past | pdf | 171 comments
1542. Virtual Width Networks (VWN) (arxiv.org)
9 points by tesserato 230 days ago | hide | past | pdf | discuss
1543. CodeLogician: Neuro-symbolic reasoning for precise software analysis (arxiv.org)
2 points by NTCTech 231 days ago | hide | past | pdf | 1 comment
1544. Delegated Agent Authorization Constrained to Semantic Task-to-Scope Matching (arxiv.org)
1 point by mooreds 231 days ago | hide | past | pdf | discuss
1545. Evaluating AGENTS.md: are they helpful for coding agents? (arxiv.org)
232 points by mustaphah 231 days ago | hide | past | pdf | 161 comments
1546. Multi-Agent Teams Hold Experts Back (arxiv.org)
1 point by fauigerzigerk 231 days ago | hide | past | pdf | discuss
1547. Large Language Model Reasoning Failures (arxiv.org)
1 point by kawera 231 days ago | hide | past | pdf | discuss
1548. Towards Autonomous Mathematics Research (arxiv.org)
107 points by gmays 232 days ago | hide | past | pdf | 53 comments
1549. Retrieval-Aware Distillation for Transformer-SSM Hybrids (arxiv.org)
2 points by readitalready 232 days ago | hide | past | pdf | discuss
1550. Biases in the Blind Spot: Detecting What LLMs Fail to Mention (arxiv.org)
2 points by mpweiher 233 days ago | hide | past | pdf | discuss
1551. Towards Autonomous Mathematics Research (Google DeepMind) (arxiv.org)
1 point by u1hcw9nx 233 days ago | hide | past | pdf | discuss
1552. Remote Labor Index: Measuring AI Automation of Remote Work (arxiv.org)
2 points by Leynos 233 days ago | hide | past | pdf | discuss
1553. Generalized on-policy distillation with reward extrapolation (arxiv.org)
3 points by fzliu 233 days ago | hide | past | pdf | discuss
1554. Adversarial Patch: images that make classifiers ignore other items in a scene (arxiv.org)
1 point by felineflock 234 days ago | hide | past | pdf | discuss
1555. Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators (arxiv.org)
1 point by PaulHoule 234 days ago | hide | past | pdf | discuss
1556. Fine-Tuning GPT-5 for GPU Kernel Generation (arxiv.org)
4 points by matt_d 234 days ago | hide | past | pdf | discuss
1557. SWE-ContextBench: context learning benchmark in coding (arxiv.org)
1 point by mustaphah 234 days ago | hide | past | pdf | discuss
1558. LLMs exceed physicians on complex text-based differential diagnosis (arxiv.org)
3 points by rippeltippel 234 days ago | hide | past | pdf | 2 comments
1559. Learning to Reason in 13 Parameters (arxiv.org)
2 points by stared 234 days ago | hide | past | pdf | discuss
1560. LLM Reasoning Failures (arxiv.org)
1 point by gradus_ad 234 days ago | hide | past | pdf | discuss