| 3331. |
Reasoning Models Can Be Effective Without Thinking (arxiv.org) |
|
21 points by mfiguiere on Apr 16, 2025 | hide | past | pdf | 2 comments
|
| 3332. |
Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2 (arxiv.org) |
|
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3333. |
A Survey on Structured State Space Sequence (S4) Models (arxiv.org) |
|
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3334. |
Scaling Laws of Synthetic Data for Language Models (arxiv.org) |
|
2 points by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3335. |
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models (arxiv.org) |
|
33 points by dpstart01 on Apr 15, 2025 | hide | past | pdf | 3 comments
|
| 3336. |
A Block-Wise Pruning Algorithm for Efficient Large Language Model Compression (arxiv.org) |
|
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3337. |
Differentially Private Synthetic Data via Foundation Model APIs (arxiv.org) |
|
3 points by jsenn on Apr 15, 2025 | hide | past | pdf | 1 comment
|
| 3338. |
Memory and Bandwidth Are All You Need for Sharded Data Parallel (arxiv.org) |
|
1 point by PaulHoule on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3339. |
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks (arxiv.org) |
|
1 point by ziptron on Apr 15, 2025 | hide | past | pdf | discuss
|
| 3340. |
Teuken-7B-Base and Teuken-7B-Instruct: Towards European LLMs (2024) (arxiv.org) |
|
248 points by doener on Apr 15, 2025 | hide | past | pdf | 95 comments
|
| 3341. |
Mimrs: A Survey on Masked Image Modeling in Remote Sensing (arxiv.org) |
|
1 point by PaulHoule on Apr 14, 2025 | hide | past | pdf | discuss
|
| 3342. |
Scraping the Shadows: Deep Learning Breakthroughs in Dark Web Intelligence (arxiv.org) |
|
2 points by PaulHoule on Apr 14, 2025 | hide | past | pdf | discuss
|
| 3343. |
Wanting to Be Understood (arxiv.org) |
|
2 points by fzliu on Apr 14, 2025 | hide | past | pdf | discuss
|
| 3344. |
Automatic Functional Differentiation in Jax (arxiv.org) |
|
1 point by teleforce on Apr 14, 2025 | hide | past | pdf | discuss
|
| 3345. |
NoProp: Training neural networks without back-propagation or forward-propagation (arxiv.org) |
|
161 points by belleville on Apr 14, 2025 | hide | past | pdf | 49 comments
|
| 3346. |
Robustly identifying concepts introduced during chat fine-tuning with crosscoder (arxiv.org) |
|
6 points by veryluckyxyz on Apr 13, 2025 | hide | past | pdf | discuss
|
| 3347. |
Transfer between Modalities with MetaQueries (arxiv.org) |
|
25 points by Xiaozaa on Apr 12, 2025 | hide | past | pdf | 12 comments
|
| 3348. |
Defeating Prompt Injections by Design (arxiv.org) |
|
2 points by xnx on Apr 11, 2025 | hide | past | pdf | discuss
|
| 3349. |
KIMI-VL (Efficient Open-Source Moe VLM) Techical Report (arxiv.org) |
|
1 point by wertyk on Apr 11, 2025 | hide | past | pdf | discuss
|
| 3350. |
Π-NeSy: A Possibilistic Neuro-Symbolic Approach (arxiv.org) |
|
1 point by _jh5l on Apr 10, 2025 | hide | past | pdf | discuss
|
| 3351. |
A Comprehensive Survey on Long Context Language Modeling (arxiv.org) |
|
3 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
|
| 3352. |
ProtoGS: Efficient and High-Quality Rendering with 3D Gaussian Prototypes (arxiv.org) |
|
22 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
|
| 3353. |
Contextualize-Then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B (arxiv.org) |
|
1 point by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
|
| 3354. |
Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs (arxiv.org) |
|
2 points by PaulHoule on Apr 9, 2025 | hide | past | pdf | discuss
|
| 3355. |
NNN: Next-Generation Neural Networks for Marketing Mix Modeling (arxiv.org) |
|
25 points by tmulc on Apr 9, 2025 | hide | past | pdf | 3 comments
|
| 3356. |
Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best? (arxiv.org) |
|
3 points by s1l3nt on Apr 8, 2025 | hide | past | pdf | 1 comment
|
| 3357. |
Rethinking Reflection in Pre-Training (arxiv.org) |
|
1 point by swyx on Apr 8, 2025 | hide | past | pdf | discuss
|
| 3358. |
Are Domain-Specific Trade-Offs Undermining On-Device Language Models? (arxiv.org) |
|
1 point by PaulHoule on Apr 8, 2025 | hide | past | pdf | discuss
|
| 3359. |
Can reinforcement learning for LLMs scale beyond math and coding tasks? Probably (arxiv.org) |
|
6 points by GabrielBianconi on Apr 8, 2025 | hide | past | pdf | 4 comments
|
| 3360. |
SmolVLM: Redefining small and efficient multimodal models (arxiv.org) |
|
2 points by wertyk on Apr 8, 2025 | hide | past | pdf | discuss
|
| More |