about
Stories from July 31, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Orca-Bench: How Ready Are Language Model Agents for Oncall? (arxiv.org)
30 points by yruzin 64 days ago | hide | past | pdf | 11 comments
2. Neuro-Inspired Inverse Learning for Planning and Control (arxiv.org)
3 points by kensai 64 days ago | hide | past | pdf | 1 comment
3. CircuitProver: Agentic Lean 4 Theorem Proving for Hardware Verification (arxiv.org)
2 points by Jimmc414 64 days ago | hide | past | pdf | discuss
4. Metis: Memory Foundation Model (arxiv.org)
2 points by mfiguiere 64 days ago | hide | past | pdf | discuss
5. Stealthy Concurrent Audio Prompt Injections Against Multimodal LLM Agents (arxiv.org)
2 points by zhinit 64 days ago | hide | past | pdf | discuss
6. Study: Coding agents rarely retrieve open-source contribution rules (arxiv.org)
2 points by wek 64 days ago | hide | past | pdf | discuss
7. Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling (arxiv.org)
2 points by sbulaev 64 days ago | hide | past | pdf | discuss