about
241. Post-Training Language Models for Gold-Medal Performance in Coding Competitions (arxiv.org)
2 points by sbulaev 30 days ago | hide | past | pdf | discuss
242. Redwood: A Frontier AI Accelerator Designed from Scratch in 2 Weeks by AI (arxiv.org)
2 points by imakwana 30 days ago | hide | past | pdf | discuss
243. Coral: An LLM-Native Harness for Production Recommender Systems (arxiv.org)
2 points by Anon84 30 days ago | hide | past | pdf | discuss
244. Bandits in Prod: Hyperparameter Optimization at Inference Time (arxiv.org)
2 points by Labo333 30 days ago | hide | past | pdf | discuss
245. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents (arxiv.org)
2 points by gmays 30 days ago | hide | past | pdf | discuss
246. Flavourbench: LLM Eval with Executable Culinary Ground Truth (arxiv.org)
2 points by josefchen 30 days ago | hide | past | pdf | discuss
247. Qwen-Drive-1.0 (arxiv.org)
2 points by amancina79 30 days ago | hide | past | pdf | discuss
248. Understanding, Mitigating Numerical Sources of Nondeterminism in LLM Inference (arxiv.org)
1 point by shakna 30 days ago | hide | past | pdf | discuss
249. Post-Training Language Models for Gold-Medal Performance in Coding Competitions (arxiv.org)
1 point by oliviayii 31 days ago | hide | past | pdf | discuss
250. Language Models Can Control Their Own Attention (arxiv.org)
2 points by E-Reverance 31 days ago | hide | past | pdf | discuss
251. Proper Scoring Rules Shape LLM Forecasting (arxiv.org)
3 points by bturtel 31 days ago | hide | past | pdf | discuss
252. LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes (arxiv.org)
20 points by sbulaev 31 days ago | hide | past | pdf | 14 comments
253. What's in Your Agent's Context? Context Privilege Escalation Attacks Against AI (arxiv.org)
1 point by sbulaev 31 days ago | hide | past | pdf | discuss
254. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arxiv.org)
3 points by mdp2021 31 days ago | hide | past | pdf | 1 comment
255. The Emergent Symbolic Structure of Artificial Neural Networks (arxiv.org)
297 points by schmuhblaster 32 days ago | hide | past | pdf | 110 comments
256. One in three AI scribe notes carries a verified clinical error (arxiv.org)
2 points by sbulaev 32 days ago | hide | past | pdf | 1 comment
257. SKILL.state: Scalable Long-Horizon Agent Skills (arxiv.org)
1 point by Hoefner 32 days ago | hide | past | pdf | discuss
258. Ghost Couple:Correlated LLM Name Priors Haunting the Web and Academic Publishing (arxiv.org)
1 point by bookofjoe 32 days ago | hide | past | pdf | discuss
259. Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 (arxiv.org)
3 points by simonpure 32 days ago | hide | past | pdf | discuss
260. Can escalation channels redirect reward hacking toward defect disclosure? (arxiv.org)
2 points by sbulaev 32 days ago | hide | past | pdf | discuss
261. Temporal Analysis of NetFlow Datasets for Network Intrusion Detection Systems (arxiv.org)
1 point by ENOMEM 32 days ago | hide | past | pdf | discuss
262. Breaking Execution Continuity of Agent Systems via Rollback (arxiv.org)
1 point by sbulaev 32 days ago | hide | past | pdf | discuss
263. Program Learning with Verifiable Rewards: Symbolic Backpropagation (arxiv.org)
2 points by sbulaev 33 days ago | hide | past | pdf | discuss
264. Sliding-window beats linear attention (arxiv.org)
2 points by simonpure 33 days ago | hide | past | pdf | 1 comment
265. Logic and the 2-Simplicial Transformer (2019) (arxiv.org)
2 points by wslh 33 days ago | hide | past | pdf | discuss
266. A Policy Algebra for Trust-Preserving Agentic AI Execution (arxiv.org)
1 point by gmays 33 days ago | hide | past | pdf | discuss
267. Sliding-window beats linear attention (arxiv.org)
4 points by sbulaev 33 days ago | hide | past | pdf | discuss
268. AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI (arxiv.org)
2 points by Jimmc414 34 days ago | hide | past | pdf | discuss
269. Benchmarking Confidential Computing Performance on Nvidia Blackwell GPUs (arxiv.org)
1 point by Jimmc414 34 days ago | hide | past | pdf | discuss
270. Static Evaluation of Model Switching in LLM Agents Scores the Wrong World (arxiv.org)
4 points by danbitengo 34 days ago | hide | past | pdf | discuss