| 4321. |
Solving Global Lyapunov functions: open problem in mathematics with transformers (arxiv.org) |
|
3 points by famouswaffles on Oct 27, 2024 | hide | past | pdf | discuss
|
| 4322. |
Decomposing the Dark Matter of Sparse Autoencoders (arxiv.org) |
|
3 points by cgadski on Oct 27, 2024 | hide | past | pdf | discuss
|
| 4323. |
Rethinking Softmax: Self-Attention with Polynomial Activations (arxiv.org) |
|
2 points by veryluckyxyz on Oct 27, 2024 | hide | past | pdf | discuss
|
| 4324. |
Moonshine speech to text model 1.7x faster than OpenAI's Whisper, as accurate (arxiv.org) |
|
2 points by Curiositry on Oct 26, 2024 | hide | past | pdf | 1 comment
|
| 4325. |
State-space models can learn in-context by gradient descent (arxiv.org) |
|
86 points by dsalaj on Oct 26, 2024 | hide | past | pdf | 58 comments
|
| 4326. |
Post-Training Layer Scaling Prevents Forgetting and Enhances Model Merging (arxiv.org) |
|
1 point by veryluckyxyz on Oct 26, 2024 | hide | past | pdf | discuss
|
| 4327. |
Leopard: A Vision Language Model for Text-Rich Multi-Image Tasks (arxiv.org) |
|
6 points by PaulHoule on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4328. |
Breaking Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss (arxiv.org) |
|
1 point by belter on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4329. |
Large Language Models for Mathematicians (arxiv.org) |
|
1 point by belter on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4330. |
Dualformer: Controllable Fast/Slow Thinking,Learning with Rand. Reasoning Traces (arxiv.org) |
|
2 points by gadilif on Oct 25, 2024 | hide | past | pdf | 1 comment
|
| 4331. |
Improving Generalization Performance by Switching from Adam to SGD (2017) (arxiv.org) |
|
1 point by fzliu on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4332. |
Point Cloud Compression with Bits-Back Coding (arxiv.org) |
|
1 point by sandwichsphinx on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4333. |
What Matters in Transformers? Not All Attention Is Needed (arxiv.org) |
|
4 points by Anon84 on Oct 25, 2024 | hide | past | pdf | discuss
|
| 4334. |
Perils and Promises of Synthetic Data in a Self-Generating World (arxiv.org) |
|
1 point by Wheatman on Oct 24, 2024 | hide | past | pdf | 1 comment
|
| 4335. |
ALTA: Compiler-Based Analysis of Transformers (arxiv.org) |
|
2 points by belter on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4336. |
SymGen: Towards verifiable text generation with symbolic references (arxiv.org) |
|
1 point by hhs on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4337. |
Unsupervised Human Preference Learning (arxiv.org) |
|
3 points by PaulHoule on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4338. |
Torch.manual_seed(3407) is all you need: On the influence of random seeds (arxiv.org) |
|
2 points by reqo on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4339. |
Audio-Driven Emotional 3D Talking-Head Generation (arxiv.org) |
|
2 points by sandwichsphinx on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4340. |
The Fair Language Model Paradox (arxiv.org) |
|
1 point by PaulHoule on Oct 24, 2024 | hide | past | pdf | discuss
|
| 4341. |
From Tokens to Words: On the Inner Lexicon of LLMs (arxiv.org) |
|
2 points by PaulHoule on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4342. |
We discovered a way to measure LLM bias while building a recruitment tool (arxiv.org) |
|
1 point by dreamfactored on Oct 23, 2024 | hide | past | pdf | 1 comment
|
| 4343. |
What makes your model a low-empathy or warmth person? (arxiv.org) |
|
1 point by PaulHoule on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4344. |
Remote Timing Attacks on Efficient Language Model Inference (arxiv.org) |
|
1 point by belter on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4345. |
Neural Network Quine (Attempt) (arxiv.org) |
|
1 point by diginova on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4346. |
Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces (arxiv.org) |
|
44 points by hislaziness on Oct 23, 2024 | hide | past | pdf | 15 comments
|
| 4347. |
What Matters in Transformers? Not All Attention Is Needed (arxiv.org) |
|
3 points by jonbaer on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4348. |
Assessing the Performance of Human-Capable LLMs – Are LLMs Coming for Your Job? (arxiv.org) |
|
1 point by sandwichsphinx on Oct 23, 2024 | hide | past | pdf | discuss
|
| 4349. |
RepoGraph: Enhancing AI Software Engineering with Repository-Level Code Graph (arxiv.org) |
|
1 point by sandwichsphinx on Oct 22, 2024 | hide | past | pdf | discuss
|
| 4350. |
Computational Copyright: Towards a Royalty Model for Music Generative AI (arxiv.org) |
|
1 point by samanbb on Oct 22, 2024 | hide | past | pdf | discuss
|
| More |