about
1621. Pulse: Decentralized RL training centralized speed (100x weight sync reduction) (arxiv.org)
2 points by synapz_org 243 days ago | hide | past | pdf | discuss
1622. Who's in Charge? Disempowerment Patterns in Real-World LLM Usage (arxiv.org)
3 points by 7777777phil 243 days ago | hide | past | pdf | discuss
1623. Are Pre-Trained Convolutions Better Than Pre-Trained Transformers? (2021) (arxiv.org)
2 points by fzliu 243 days ago | hide | past | pdf | discuss
1624. Pendulum: A Benchmark for Assessing Sycophancy in MLLM's (arxiv.org)
1 point by onestay42 243 days ago | hide | past | pdf | discuss
1625. Exploration Posteriors for Generative Modeling Using Only Negative Rewards (arxiv.org)
1 point by numeri 243 days ago | hide | past | pdf | discuss
1626. Revisiting Disaggregated LLM Serving for Performance and Energy Implications (arxiv.org)
1 point by PaulHoule 244 days ago | hide | past | pdf | discuss
1627. Linear representations in LLMs can change dramatically over a conversation (arxiv.org)
5 points by gmays 244 days ago | hide | past | pdf | discuss
1628. Language-Related Ideological Divergence in LLM Analysis of Political Documents (arxiv.org)
1 point by PaulHoule 244 days ago | hide | past | pdf | discuss
1629. Reasoning Models Generate Societies of Thought (arxiv.org)
2 points by PaulHoule 244 days ago | hide | past | pdf | discuss
1630. Strategies of cooperation and defection in five large language models (arxiv.org)
1 point by PaulHoule 244 days ago | hide | past | pdf | discuss
1631. A Pragmatic VLA Foundation Model (arxiv.org)
1 point by mountainview 244 days ago | hide | past | pdf | discuss
1632. PaperBanana: Automating Academic Illustration for AI Scientists (arxiv.org)
1 point by fzliu 244 days ago | hide | past | pdf | discuss
1633. Power Aware Dynamic Reallocation for Inference (arxiv.org)
3 points by PaulHoule 245 days ago | hide | past | pdf | discuss
1634. Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs (arxiv.org)
3 points by PaulHoule 245 days ago | hide | past | pdf | discuss
1635. The SWE-Bench Illusion: When LLMs Remember Instead of Reason (arxiv.org)
2 points by cadabrabra 245 days ago | hide | past | pdf | discuss
1636. YuriiFormer: A Suite of Nesterov-Accelerated Transformers (arxiv.org)
2 points by kelseyfrog 245 days ago | hide | past | pdf | discuss
1637. 3DGS-Drag: Dragging Gaussians for Intuitive Point-Based 3D Editing (arxiv.org)
3 points by PaulHoule 245 days ago | hide | past | pdf | discuss
1638. Semi-Autonomous Mathematics Discovery with Gemini: Erdős Problems Case Study (arxiv.org)
1 point by tzury 245 days ago | hide | past | pdf | discuss
1639. Forcing and Diagnosing Failure Modes of Fourier Neural Operators (arxiv.org)
3 points by TimorousBestie 246 days ago | hide | past | pdf | 1 comment
1640. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability (arxiv.org)
4 points by gmays 246 days ago | hide | past | pdf | discuss
1641. Demystifying ARM SME to Optimize General Matrix Multiplications (arxiv.org)
88 points by matt_d 247 days ago | hide | past | pdf | 19 comments
1642. VaultGemma: A Differentially Private LLM (arxiv.org)
1 point by PaulHoule 247 days ago | hide | past | pdf | discuss
1643. Magellan: Autonomous Discovery of Compiler Optimization Heuristics w/AlphaEvolve (arxiv.org)
4 points by matt_d 247 days ago | hide | past | pdf | discuss
1644. Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with LLMs (arxiv.org)
5 points by PaulHoule 247 days ago | hide | past | pdf | discuss
1645. Shaping capabilities with token-level data filtering (arxiv.org)
2 points by brandonb 247 days ago | hide | past | pdf | 1 comment
1646. Self-Distillation Enables Continual Learning (arxiv.org)
2 points by simonpure 247 days ago | hide | past | pdf | discuss
1647. The End of Transformers (2025) (arxiv.org)
1 point by teleforce 247 days ago | hide | past | pdf | discuss
1648. Qwen3-ASR Technical Report (arxiv.org)
7 points by _____k 247 days ago | hide | past | pdf | discuss
1649. Lost in the Middle: How Language Models Use Long Contexts (2023) (arxiv.org)
2 points by wslh 248 days ago | hide | past | pdf | discuss
1650. Scaling Embeddings Outperforms Scaling Experts in Language Models (arxiv.org)
1 point by simonpure 248 days ago | hide | past | pdf | discuss