about
3331. Reasoning Models Can Be Effective Without Thinking (arxiv.org)
21 points by mfiguiere on Apr 16, 2025 | hide | past | pdf | 2 comments
3332. Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2 (arxiv.org)
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
3333. A Survey on Structured State Space Sequence (S4) Models (arxiv.org)
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
3334. Scaling Laws of Synthetic Data for Language Models (arxiv.org)
2 points by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
3335. M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models (arxiv.org)
33 points by dpstart01 on Apr 15, 2025 | hide | past | pdf | 3 comments
3336. A Block-Wise Pruning Algorithm for Efficient Large Language Model Compression (arxiv.org)
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
3337. Differentially Private Synthetic Data via Foundation Model APIs (arxiv.org)
3 points by jsenn on Apr 15, 2025 | hide | past | pdf | 1 comment
3338. Memory and Bandwidth Are All You Need for Sharded Data Parallel (arxiv.org)
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
3339. What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks (arxiv.org)
1 point by ziptron on Apr 15, 2025 | hide | past | pdf | discuss
3340. Teuken-7B-Base and Teuken-7B-Instruct: Towards European LLMs (2024) (arxiv.org)
248 points by doener on Apr 15, 2025 | hide | past | pdf | 95 comments
3341. Mimrs: A Survey on Masked Image Modeling in Remote Sensing (arxiv.org)
1 point by PaulHoule on Apr 14, 2025 | hide | past | pdf | discuss
3342. Scraping the Shadows: Deep Learning Breakthroughs in Dark Web Intelligence (arxiv.org)
2 points by PaulHoule on Apr 14, 2025 | hide | past | pdf | discuss
3343. Wanting to Be Understood (arxiv.org)
2 points by fzliu on Apr 14, 2025 | hide | past | pdf | discuss
3344. Automatic Functional Differentiation in Jax (arxiv.org)
1 point by teleforce on Apr 14, 2025 | hide | past | pdf | discuss
3345. NoProp: Training neural networks without back-propagation or forward-propagation (arxiv.org)
161 points by belleville on Apr 14, 2025 | hide | past | pdf | 49 comments
3346. Robustly identifying concepts introduced during chat fine-tuning with crosscoder (arxiv.org)
6 points by veryluckyxyz on Apr 13, 2025 | hide | past | pdf | discuss
3347. Transfer between Modalities with MetaQueries (arxiv.org)
25 points by Xiaozaa on Apr 12, 2025 | hide | past | pdf | 12 comments
3348. Defeating Prompt Injections by Design (arxiv.org)
2 points by xnx on Apr 11, 2025 | hide | past | pdf | discuss
3349. KIMI-VL (Efficient Open-Source Moe VLM) Techical Report (arxiv.org)
1 point by wertyk on Apr 11, 2025 | hide | past | pdf | discuss
3350. Π-NeSy: A Possibilistic Neuro-Symbolic Approach (arxiv.org)
1 point by _jh5l on Apr 10, 2025 | hide | past | pdf | discuss
3351. A Comprehensive Survey on Long Context Language Modeling (arxiv.org)
3 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
3352. ProtoGS: Efficient and High-Quality Rendering with 3D Gaussian Prototypes (arxiv.org)
22 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
3353. Contextualize-Then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B (arxiv.org)
1 point by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
3354. Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs (arxiv.org)
2 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
3355. NNN: Next-Generation Neural Networks for Marketing Mix Modeling (arxiv.org)
25 points by tmulc on Apr 9, 2025 | hide | past | pdf | 3 comments
3356. Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best? (arxiv.org)
3 points by s1l3nt on Apr 8, 2025 | hide | past | pdf | 1 comment
3357. Rethinking Reflection in Pre-Training (arxiv.org)
1 point by swyx on Apr 8, 2025 | hide | past | pdf | discuss
3358. Are Domain-Specific Trade-Offs Undermining On-Device Language Models? (arxiv.org)
1 point by PaulHoule on Apr 8, 2025 | hide | past | pdf | discuss
3359. Can reinforcement learning for LLMs scale beyond math and coding tasks? Probably (arxiv.org)
6 points by GabrielBianconi on Apr 8, 2025 | hide | past | pdf | 4 comments
3360. SmolVLM: Redefining small and efficient multimodal models (arxiv.org)
2 points by wertyk on Apr 8, 2025 | hide | past | pdf | discuss