about
3601. The FFT Strikes Back: An Efficient Alternative to Self-Attention (arxiv.org)
456 points by iNic on Feb 26, 2025 | hide | past | pdf | 168 comments
3602. Demonstrating specification gaming in reasoning models (arxiv.org)
1 point by wluk on Feb 26, 2025 | hide | past | pdf | 1 comment
3603. Comply: Learning Sentences with Complex Weights Inspired by Fruit Fly Olfaction (arxiv.org)
2 points by PaulHoule on Feb 25, 2025 | hide | past | pdf | discuss
3604. Discovering Chunks in Neural Embeddings for Interpretability (arxiv.org)
2 points by PaulHoule on Feb 25, 2025 | hide | past | pdf | discuss
3605. The Influence of Prompt Politeness on LLM Performance (arxiv.org)
17 points by blululu on Feb 25, 2025 | hide | past | pdf | discuss
3606. Robust Ladder Climbing with a Quadrupedal Robot (arxiv.org)
3 points by nill0 on Feb 25, 2025 | hide | past | pdf | 1 comment
3607. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing (arxiv.org)
1 point by wilson090 on Feb 25, 2025 | hide | past | pdf | discuss
3608. Hijacking Chain-of-Thought Safety Reasoning to Jailbreak Large Reasoning Models (arxiv.org)
1 point by rntn on Feb 25, 2025 | hide | past | pdf | discuss
3609. Combining Base and Instruction-Tuned LMs for Better Synthetic Data Generation (arxiv.org)
1 point by PaulHoule on Feb 24, 2025 | hide | past | pdf | discuss
3610. VLMaterial: Procedural Material Generation with Large Vision-Language Models (arxiv.org)
26 points by eamag on Feb 24, 2025 | hide | past | pdf | discuss
3611. Sift: Grounding LLM Reasoning in Contexts via Stickers (arxiv.org)
2 points by zengzihao on Feb 24, 2025 | hide | past | pdf | 1 comment
3612. Diffusion Models Learn Low-Dimensional Distributions via Subspace Clustering (arxiv.org)
1 point by pizza on Feb 24, 2025 | hide | past | pdf | discuss
3613. Agentic Deep Graph Reasoning Yields Self-Organizing Knowledge Networks (arxiv.org)
5 points by Dezash on Feb 23, 2025 | hide | past | pdf | discuss
3614. Chartist: Task-Driven Eye Movement Control for Chart Reading (arxiv.org)
14 points by PaulHoule on Feb 23, 2025 | hide | past | pdf | 2 comments
3615. Large Language Diffusion Models (arxiv.org)
7 points by kadushka on Feb 23, 2025 | hide | past | pdf | 3 comments
3616. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
1 point by azhenley on Feb 23, 2025 | hide | past | pdf | discuss
3617. Intuitive physics understanding emerges from self-supervised pretraining (arxiv.org)
1 point by asah on Feb 23, 2025 | hide | past | pdf | 1 comment
3618. AI and ML Accelerator Survey and Trends (2022) (arxiv.org)
1 point by teleforce on Feb 23, 2025 | hide | past | pdf | discuss
3619. Resource-Efficient and Effective Code Summarization (arxiv.org)
1 point by PaulHoule on Feb 23, 2025 | hide | past | pdf | discuss
3620. None of the Others: General Technique to Distinguish Reasoning from Memorization (arxiv.org)
2 points by amunozo on Feb 23, 2025 | hide | past | pdf | discuss
3621. Evaluating LLMs Capabilities Towards Understanding Social Dynamics (arxiv.org)
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
3622. BaxBench: Can LLMs Generate Correct and Secure Back Ends? (arxiv.org)
3 points by elashri on Feb 22, 2025 | hide | past | pdf | discuss
3623. Experiments in News Bias Detection with Pre-Trained Neural Transformers (arxiv.org)
1 point by rntn on Feb 22, 2025 | hide | past | pdf | discuss
3624. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
1 point by hexhowells on Feb 22, 2025 | hide | past | pdf | discuss
3625. Too Noisy to Learn: Enhancing Data Quality for Code Review (arxiv.org)
1 point by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
3626. Some critical issues with the SWE-bench dataset (arxiv.org)
350 points by joshwa on Feb 21, 2025 | hide | past | pdf | 116 comments
3627. MLGym: A New Framework and Benchmark for Advancing AI Research Agents (arxiv.org)
1 point by jonbaer on Feb 21, 2025 | hide | past | pdf | discuss
3628. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
2 points by azhenley on Feb 21, 2025 | hide | past | pdf | discuss
3629. Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging (arxiv.org)
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
3630. The Case for Cognitive-Dissonance-Aware Knowledge Updates in LLMs (arxiv.org)
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss