about
Stories from May 5, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents (arxiv.org)
163 points by gmays 151 days ago | hide | past | pdf | 32 comments
2. DeepSeek V4's indexer OOMs at 65K context. We got it to 1M in 6G (arxiv.org)
8 points by OsamaJaber 152 days ago | hide | past | pdf | discuss
3. CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion (arxiv.org)
5 points by matt_d 151 days ago | hide | past | pdf | discuss
4. A Theory of Generalization in Deep Learning (arxiv.org)
4 points by E-Reverance 151 days ago | hide | past | pdf | discuss
5. Process-Level Reward Modeling for Agentic Data Analysis (arxiv.org)
4 points by gmays 152 days ago | hide | past | pdf | discuss
6. Shared Lexical Task Representations Explain Behavioral Variability in LLMs (arxiv.org)
2 points by PaulHoule 151 days ago | hide | past | pdf | discuss
7. Models hallucinate more than you think (arxiv.org)
2 points by axelriet 151 days ago | hide | past | pdf | 1 comment
8. Unlocking Long-Context LLM Training via Compiler-Based Sequence Parallelism (arxiv.org)
2 points by PaulHoule 151 days ago | hide | past | pdf | discuss
9. When innocent tools form dangerous chains to jailbreak LLM agents (arxiv.org)
2 points by leecoursey 151 days ago | hide | past | pdf | discuss
10. The Last Human-Written Paper: Agent-Native Research Artifacts (arxiv.org)
2 points by amberjcjj 151 days ago | hide | past | pdf | discuss
11. Faster RL Post-Training Rollouts via System-Integrated Speculative Decoding (arxiv.org)
1 point by gmays 152 days ago | hide | past | pdf | discuss