| 3601. |
The FFT Strikes Back: An Efficient Alternative to Self-Attention (arxiv.org) |
|
456 points by iNic on Feb 26, 2025 | hide | past | pdf | 168 comments
|
| 3602. |
Demonstrating specification gaming in reasoning models (arxiv.org) |
|
1 point by wluk on Feb 26, 2025 | hide | past | pdf | 1 comment
|
| 3603. |
Comply: Learning Sentences with Complex Weights Inspired by Fruit Fly Olfaction (arxiv.org) |
|
2 points by PaulHoule on Feb 25, 2025 | hide | past | pdf | discuss
|
| 3604. |
Discovering Chunks in Neural Embeddings for Interpretability (arxiv.org) |
|
2 points by PaulHoule on Feb 25, 2025 | hide | past | pdf | discuss
|
| 3605. |
The Influence of Prompt Politeness on LLM Performance (arxiv.org) |
|
17 points by blululu on Feb 25, 2025 | hide | past | pdf | discuss
|
| 3606. |
Robust Ladder Climbing with a Quadrupedal Robot (arxiv.org) |
|
3 points by nill0 on Feb 25, 2025 | hide | past | pdf | 1 comment
|
| 3607. |
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing (arxiv.org) |
|
1 point by wilson090 on Feb 25, 2025 | hide | past | pdf | discuss
|
| 3608. |
Hijacking Chain-of-Thought Safety Reasoning to Jailbreak Large Reasoning Models (arxiv.org) |
|
1 point by rntn on Feb 25, 2025 | hide | past | pdf | discuss
|
| 3609. |
Combining Base and Instruction-Tuned LMs for Better Synthetic Data Generation (arxiv.org) |
|
1 point by PaulHoule on Feb 24, 2025 | hide | past | pdf | discuss
|
| 3610. |
VLMaterial: Procedural Material Generation with Large Vision-Language Models (arxiv.org) |
|
26 points by eamag on Feb 24, 2025 | hide | past | pdf | discuss
|
| 3611. |
Sift: Grounding LLM Reasoning in Contexts via Stickers (arxiv.org) |
|
2 points by zengzihao on Feb 24, 2025 | hide | past | pdf | 1 comment
|
| 3612. |
Diffusion Models Learn Low-Dimensional Distributions via Subspace Clustering (arxiv.org) |
|
1 point by pizza on Feb 24, 2025 | hide | past | pdf | discuss
|
| 3613. |
Agentic Deep Graph Reasoning Yields Self-Organizing Knowledge Networks (arxiv.org) |
|
5 points by Dezash on Feb 23, 2025 | hide | past | pdf | discuss
|
| 3614. |
Chartist: Task-Driven Eye Movement Control for Chart Reading (arxiv.org) |
|
14 points by PaulHoule on Feb 23, 2025 | hide | past | pdf | 2 comments
|
| 3615. |
Large Language Diffusion Models (arxiv.org) |
|
7 points by kadushka on Feb 23, 2025 | hide | past | pdf | 3 comments
|
| 3616. |
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
1 point by azhenley on Feb 23, 2025 | hide | past | pdf | discuss
|
| 3617. |
Intuitive physics understanding emerges from self-supervised pretraining (arxiv.org) |
|
1 point by asah on Feb 23, 2025 | hide | past | pdf | 1 comment
|
| 3618. |
AI and ML Accelerator Survey and Trends (2022) (arxiv.org) |
|
1 point by teleforce on Feb 23, 2025 | hide | past | pdf | discuss
|
| 3619. |
Resource-Efficient and Effective Code Summarization (arxiv.org) |
|
1 point by PaulHoule on Feb 23, 2025 | hide | past | pdf | discuss
|
| 3620. |
None of the Others: General Technique to Distinguish Reasoning from Memorization (arxiv.org) |
|
2 points by amunozo on Feb 23, 2025 | hide | past | pdf | discuss
|
| 3621. |
Evaluating LLMs Capabilities Towards Understanding Social Dynamics (arxiv.org) |
|
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
|
| 3622. |
BaxBench: Can LLMs Generate Correct and Secure Back Ends? (arxiv.org) |
|
3 points by elashri on Feb 22, 2025 | hide | past | pdf | discuss
|
| 3623. |
Experiments in News Bias Detection with Pre-Trained Neural Transformers (arxiv.org) |
|
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
|
| 3624. |
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
1 point by hexhowells on Feb 22, 2025 | hide | past | pdf | discuss
|
| 3625. |
Too Noisy to Learn: Enhancing Data Quality for Code Review (arxiv.org) |
|
1 point by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
|
| 3626. |
Some critical issues with the SWE-bench dataset (arxiv.org) |
|
350 points by joshwa on Feb 21, 2025 | hide | past | pdf | 116 comments
|
| 3627. |
MLGym: A New Framework and Benchmark for Advancing AI Research Agents (arxiv.org) |
|
1 point by jonbaer on Feb 21, 2025 | hide | past | pdf | discuss
|
| 3628. |
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
2 points by azhenley on Feb 21, 2025 | hide | past | pdf | discuss
|
| 3629. |
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging (arxiv.org) |
|
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
|
| 3630. |
The Case for Cognitive-Dissonance-Aware Knowledge Updates in LLMs (arxiv.org) |
|
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
|
| More |