| 4411. |
Grokking at the edge of linear separability (arxiv.org) |
|
89 points by marojejian on Oct 11, 2024 | hide | past | pdf | 26 comments
|
| 4412. |
The Role of Anchor Tokens in Self-Attention Networks (arxiv.org) |
|
18 points by smooke on Oct 11, 2024 | hide | past | pdf | 5 comments
|
| 4413. |
Scaling Laws for Diffusion Transformers (arxiv.org) |
|
1 point by georgehill on Oct 11, 2024 | hide | past | pdf | discuss
|
| 4414. |
Cross-Domain Content Generation with Domain-Specific Small Language Models (arxiv.org) |
|
1 point by PaulHoule on Oct 11, 2024 | hide | past | pdf | discuss
|
| 4415. |
Understanding the Limitations of Mathematical Reasoning in LLMs (arxiv.org) |
|
282 points by hnhn34 on Oct 11, 2024 | hide | past | pdf | 266 comments
|
| 4416. |
ARIA: An Open Multimodal Native Mixture-of-Experts Model (arxiv.org) |
|
97 points by jinqueeny on Oct 11, 2024 | hide | past | pdf | 21 comments
|
| 4417. |
Intelligence at the Edge of Chaos (arxiv.org) |
|
3 points by p1esk on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4418. |
Everything Everywhere All at Once: LLMs Can In-Context Learn Multiple Tasks (arxiv.org) |
|
1 point by jasondavies on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4419. |
Commissioning of the First-Gen BrainScaleS Wafer-Scale Neuromorphic System (arxiv.org) |
|
1 point by rbanffy on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4420. |
Pixtral 12B Technical Report (arxiv.org) |
|
1 point by ahiknsr on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4421. |
Comprehensive Survey of Mamba Architectures for Medical Image Analysis,Beyond (arxiv.org) |
|
2 points by maddyawesome on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4422. |
INT8 FlashAttention (arxiv.org) |
|
2 points by ebalit on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4423. |
Multilingual Jailbreak Challenges in Large Language Models (arxiv.org) |
|
1 point by rntn on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4424. |
ε -VAE: Denoising as Visual Decoding (arxiv.org) |
|
5 points by lnyan on Oct 10, 2024 | hide | past | pdf | 1 comment
|
| 4425. |
Intelligence at the Edge of Chaos (arxiv.org) |
|
2 points by hardmaru on Oct 10, 2024 | hide | past | pdf | discuss
|
| 4426. |
Intelligence at the Edge of Chaos (arxiv.org) |
|
3 points by jasondavies on Oct 9, 2024 | hide | past | pdf | 1 comment
|
| 4427. |
Follow-Up Attention: A Study of Developer and Neural Model Code Exploration (arxiv.org) |
|
1 point by azhenley on Oct 9, 2024 | hide | past | pdf | discuss
|
| 4428. |
Mamba for Scalable and Efficient Personalized Recommendations (arxiv.org) |
|
2 points by PaulHoule on Oct 9, 2024 | hide | past | pdf | discuss
|
| 4429. |
Learned Indexes for a Google-Scale Disk-Based Database (arxiv.org) |
|
1 point by hamilyon2 on Oct 9, 2024 | hide | past | pdf | discuss
|
| 4430. |
Addition is all you need for energy-efficient language models (arxiv.org) |
|
334 points by InvisibleUp on Oct 9, 2024 | hide | past | pdf | 126 comments
|
| 4431. |
Language Models Are Multilingual Chain-of-Thought Reasoners (2022) (arxiv.org) |
|
1 point by rntn on Oct 9, 2024 | hide | past | pdf | discuss
|
| 4432. |
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization (arxiv.org) |
|
6 points by PaulHoule on Oct 9, 2024 | hide | past | pdf | discuss
|
| 4433. |
LLMs as Markov Chains (arxiv.org) |
|
5 points by akrymski on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4434. |
TableRAG: Million-Token Table Understanding with Language Models (arxiv.org) |
|
2 points by fzliu on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4435. |
How Chinese Are Chinese Language Models? Lack of Language Policy in China's LLMs (arxiv.org) |
|
4 points by rntn on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4436. |
Eliciting Better Multilingual Structured Reasoning from LLMs Through Code (arxiv.org) |
|
2 points by rntn on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4437. |
Large Language Models as Markov Chains (arxiv.org) |
|
1 point by tosh on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4438. |
You Only Use Reactive Attention Slice for Long Context Retrieval (arxiv.org) |
|
2 points by PaulHoule on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4439. |
Evil Geniuses: Delving into the Safety of LLM-Based Agents [pdf] (arxiv.org) |
|
1 point by squircle on Oct 8, 2024 | hide | past | pdf | discuss
|
| 4440. |
MobileLLM: Optimizing Subbillion Parameter Language Models for OnDevice UseCases (arxiv.org) |
|
2 points by rkwz on Oct 8, 2024 | hide | past | pdf | discuss
|
| More |