about
4411. Grokking at the edge of linear separability (arxiv.org)
89 points by marojejian on Oct 11, 2024 | hide | past | pdf | 26 comments
4412. The Role of Anchor Tokens in Self-Attention Networks (arxiv.org)
18 points by smooke on Oct 11, 2024 | hide | past | pdf | 5 comments
4413. Scaling Laws for Diffusion Transformers (arxiv.org)
1 point by georgehill on Oct 11, 2024 | hide | past | pdf | discuss
4414. Cross-Domain Content Generation with Domain-Specific Small Language Models (arxiv.org)
1 point by PaulHoule on Oct 11, 2024 | hide | past | pdf | discuss
4415. Understanding the Limitations of Mathematical Reasoning in LLMs (arxiv.org)
282 points by hnhn34 on Oct 11, 2024 | hide | past | pdf | 266 comments
4416. ARIA: An Open Multimodal Native Mixture-of-Experts Model (arxiv.org)
97 points by jinqueeny on Oct 11, 2024 | hide | past | pdf | 21 comments
4417. Intelligence at the Edge of Chaos (arxiv.org)
3 points by p1esk on Oct 10, 2024 | hide | past | pdf | discuss
4418. Everything Everywhere All at Once: LLMs Can In-Context Learn Multiple Tasks (arxiv.org)
1 point by jasondavies on Oct 10, 2024 | hide | past | pdf | discuss
4419. Commissioning of the First-Gen BrainScaleS Wafer-Scale Neuromorphic System (arxiv.org)
1 point by rbanffy on Oct 10, 2024 | hide | past | pdf | discuss
4420. Pixtral 12B Technical Report (arxiv.org)
1 point by ahiknsr on Oct 10, 2024 | hide | past | pdf | discuss
4421. Comprehensive Survey of Mamba Architectures for Medical Image Analysis,Beyond (arxiv.org)
2 points by maddyawesome on Oct 10, 2024 | hide | past | pdf | discuss
4422. INT8 FlashAttention (arxiv.org)
2 points by ebalit on Oct 10, 2024 | hide | past | pdf | discuss
4423. Multilingual Jailbreak Challenges in Large Language Models (arxiv.org)
1 point by rntn on Oct 10, 2024 | hide | past | pdf | discuss
4424. ε -VAE: Denoising as Visual Decoding (arxiv.org)
5 points by lnyan on Oct 10, 2024 | hide | past | pdf | 1 comment
4425. Intelligence at the Edge of Chaos (arxiv.org)
2 points by hardmaru on Oct 10, 2024 | hide | past | pdf | discuss
4426. Intelligence at the Edge of Chaos (arxiv.org)
3 points by jasondavies on Oct 9, 2024 | hide | past | pdf | 1 comment
4427. Follow-Up Attention: A Study of Developer and Neural Model Code Exploration (arxiv.org)
1 point by azhenley on Oct 9, 2024 | hide | past | pdf | discuss
4428. Mamba for Scalable and Efficient Personalized Recommendations (arxiv.org)
2 points by PaulHoule on Oct 9, 2024 | hide | past | pdf | discuss
4429. Learned Indexes for a Google-Scale Disk-Based Database (arxiv.org)
1 point by hamilyon2 on Oct 9, 2024 | hide | past | pdf | discuss
4430. Addition is all you need for energy-efficient language models (arxiv.org)
334 points by InvisibleUp on Oct 9, 2024 | hide | past | pdf | 126 comments
4431. Language Models Are Multilingual Chain-of-Thought Reasoners (2022) (arxiv.org)
1 point by rntn on Oct 9, 2024 | hide | past | pdf | discuss
4432. INT-FlashAttention: Enabling Flash Attention for INT8 Quantization (arxiv.org)
6 points by PaulHoule on Oct 9, 2024 | hide | past | pdf | discuss
4433. LLMs as Markov Chains (arxiv.org)
5 points by akrymski on Oct 8, 2024 | hide | past | pdf | discuss
4434. TableRAG: Million-Token Table Understanding with Language Models (arxiv.org)
2 points by fzliu on Oct 8, 2024 | hide | past | pdf | discuss
4435. How Chinese Are Chinese Language Models? Lack of Language Policy in China's LLMs (arxiv.org)
4 points by rntn on Oct 8, 2024 | hide | past | pdf | discuss
4436. Eliciting Better Multilingual Structured Reasoning from LLMs Through Code (arxiv.org)
2 points by rntn on Oct 8, 2024 | hide | past | pdf | discuss
4437. Large Language Models as Markov Chains (arxiv.org)
1 point by tosh on Oct 8, 2024 | hide | past | pdf | discuss
4438. You Only Use Reactive Attention Slice for Long Context Retrieval (arxiv.org)
2 points by PaulHoule on Oct 8, 2024 | hide | past | pdf | discuss
4439. Evil Geniuses: Delving into the Safety of LLM-Based Agents [pdf] (arxiv.org)
1 point by squircle on Oct 8, 2024 | hide | past | pdf | discuss
4440. MobileLLM: Optimizing Subbillion Parameter Language Models for OnDevice UseCases (arxiv.org)
2 points by rkwz on Oct 8, 2024 | hide | past | pdf | discuss