about
451. MatrAIx: Simulating the world with 8.3B persona agents (arxiv.org)
3 points by anigbrowl 58 days ago | hide | past | pdf | discuss
452. SciCode-Verified: How Benchmark Defects Underestimated LLM Scientific-Coding (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss
453. Beyond the Library: An Agentic Framework for Autoformalizing Research Math (arxiv.org)
3 points by matt_d 58 days ago | hide | past | pdf | discuss
454. ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence (arxiv.org)
2 points by Anon84 59 days ago | hide | past | pdf | discuss
455. Token-Budget-Aware LLM Reasoning (arxiv.org)
1 point by ankitg12 59 days ago | hide | past | pdf | discuss
456. Evading Chain-of-Thought Monitoring Through Model Poisoning (arxiv.org)
1 point by sbulaev 59 days ago | hide | past | pdf | discuss
457. AOHP: OS-Level Agent Harness for Personalized, Efficient and Secure Interaction (arxiv.org)
1 point by zhaoshanhui 59 days ago | hide | past | pdf | discuss
458. Unified Representation for Continuous-Latent Diffusion Language Modeling (arxiv.org)
1 point by E-Reverance 59 days ago | hide | past | pdf | 1 comment
459. Can Agents Deceive? Evaluating Reasoning+Deception Using a Social Deduction Game (arxiv.org)
2 points by theanonymousone 59 days ago | hide | past | pdf | discuss
460. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025) (arxiv.org)
175 points by robin_reala 59 days ago | hide | past | pdf | 107 comments
461. AI Security Leaderboard: Methodology, Results and Minimal Standard (arxiv.org)
3 points by sbulaev 59 days ago | hide | past | pdf | discuss
462. A Survey on LLM-as-a-Judge (arxiv.org)
3 points by Anon84 59 days ago | hide | past | pdf | discuss
463. MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025 (arxiv.org)
3 points by janandonly 59 days ago | hide | past | pdf | 1 comment
464. The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest (arxiv.org)
2 points by Jimmc414 59 days ago | hide | past | pdf | discuss
465. An AI Approach to Verified Production Cryptographic Libraries (arxiv.org)
2 points by sbulaev 60 days ago | hide | past | pdf | discuss
466. Zero-Mem: Zero-Token Memory Operations for LLM Agents (arxiv.org)
101 points by theanonymousone 60 days ago | hide | past | pdf | 12 comments
467. The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest (arxiv.org)
3 points by tosh 60 days ago | hide | past | pdf | discuss
468. AAFlow: Scalable Patterns for Agentic AI Workflows (arxiv.org)
2 points by wslh 60 days ago | hide | past | pdf | discuss
469. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production (arxiv.org)
2 points by matt_d 60 days ago | hide | past | pdf | discuss
470. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation (arxiv.org)
104 points by doppp 60 days ago | hide | past | pdf | 131 comments
471. Beyond the Final Prompt: How Conversation Context Changes AI Answers (arxiv.org)
1 point by bentannenbaum 61 days ago | hide | past | pdf | discuss
472. Can AI agents conduct open-ended AI research? (arxiv.org)
4 points by galsapir 61 days ago | hide | past | pdf | discuss
473. Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa (arxiv.org)
3 points by mil22 61 days ago | hide | past | pdf | 2 comments
474. Why Large Language Models Fail at Tabular Prediction (arxiv.org)
115 points by sbulaev 61 days ago | hide | past | pdf | 33 comments
475. Walking to the Car Wash: The Salience Bias of LLMs in Commonsense Reasoning (arxiv.org)
2 points by theanonymousone 61 days ago | hide | past | pdf | 1 comment
476. Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arxiv.org)
2 points by sbulaev 61 days ago | hide | past | pdf | discuss
477. Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repos (arxiv.org)
1 point by jamesblonde 61 days ago | hide | past | pdf | 1 comment
478. Context Compaction Theory (arxiv.org)
1 point by jadidbourbaki 61 days ago | hide | past | pdf | discuss
479. Hollow-LLM Attack: Ghost Weights That Fool Zero-Knowledge LLM Verification (arxiv.org)
1 point by sbulaev 61 days ago | hide | past | pdf | discuss
480. Compiler-Grounded Hierarchical Diagnosis for LLM Triton Kernel Optimization (arxiv.org)
1 point by matt_d 61 days ago | hide | past | pdf | discuss