about
3511. Transformers Without Normalization (arxiv.org)
4 points by qianli_cs on Mar 14, 2025 | hide | past | pdf | discuss
3512. Medical Hallucinations in Foundation Models and Their Impact on Healthcare (arxiv.org)
1 point by rntn on Mar 14, 2025 | hide | past | pdf | discuss
3513. A Survey of Long Chain-of-Thought for Reasoning Large Language Models (arxiv.org)
2 points by belter on Mar 14, 2025 | hide | past | pdf | discuss
3514. Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (arxiv.org)
2 points by flyingpumba on Mar 14, 2025 | hide | past | pdf | discuss
3515. Introduction to Sequence Modeling with Transformers (arxiv.org)
4 points by PaulHoule on Mar 13, 2025 | hide | past | pdf | discuss
3516. Robust agents learn causal world models (arxiv.org)
2 points by sebg on Mar 13, 2025 | hide | past | pdf | discuss
3517. FlexControl: Dynamic Block Activation for Diffusion Models (arxiv.org)
1 point by jinqueeny on Mar 13, 2025 | hide | past | pdf | discuss
3518. Slim attention: cut your context memory in half without loss of accuracy (arxiv.org)
7 points by DrewWas on Mar 12, 2025 | hide | past | pdf | 3 comments
3519. A Survey on Post-Training of Large Language Models (arxiv.org)
2 points by Anon84 on Mar 12, 2025 | hide | past | pdf | discuss
3520. Balancing Content Size in RAG-Text2SQL System (arxiv.org)
1 point by PaulHoule on Mar 12, 2025 | hide | past | pdf | discuss
3521. Sarcasm Detection: Improving Stance Detection with Cross-Target Capabilities (arxiv.org)
1 point by PaulHoule on Mar 12, 2025 | hide | past | pdf | discuss
3522. Traveling Waves Integrate Spatial Information Through Time (arxiv.org)
1 point by jv22222 on Mar 11, 2025 | hide | past | pdf | discuss
3523. Generalized Interpolating Discrete Diffusion (arxiv.org)
2 points by wigl on Mar 11, 2025 | hide | past | pdf | discuss
3524. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
2 points by fovc on Mar 10, 2025 | hide | past | pdf | discuss
3525. Content Rating Violations in Android Applications: A Vision-Language Approach (arxiv.org)
1 point by PaulHoule on Mar 10, 2025 | hide | past | pdf | discuss
3526. Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs (arxiv.org)
1 point by PaulHoule on Mar 10, 2025 | hide | past | pdf | discuss
3527. How to Efficiently Serve Trillions of Parameters for Online Ads Recommendation (arxiv.org)
2 points by PaulHoule on Mar 10, 2025 | hide | past | pdf | discuss
3528. Introduction to Online Control (arxiv.org)
2 points by rnjailamba on Mar 10, 2025 | hide | past | pdf | discuss
3529. Probabilistic Artificial Intelligence (arxiv.org)
352 points by pavanto on Mar 10, 2025 | hide | past | pdf | 97 comments
3530. A GS-Cache Inference Framework for Large-Scale Gaussian Splatting Models (arxiv.org)
19 points by PaulHoule on Mar 9, 2025 | hide | past | pdf | 1 comment
3531. Linguistic Generalizations Are Not Rules: Impacts on Evaluation of LMs (arxiv.org)
1 point by PaulHoule on Mar 9, 2025 | hide | past | pdf | discuss
3532. Natural Language Queries for NoSQL Databases Through Text-to-NoSQL Translation (arxiv.org)
1 point by PaulHoule on Mar 9, 2025 | hide | past | pdf | discuss
3533. Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression (arxiv.org)
4 points by PaulHoule on Mar 9, 2025 | hide | past | pdf | discuss
3534. FlexControl: Dynamic Gating for Efficient Diffusion Control (arxiv.org)
1 point by jinqueeny on Mar 8, 2025 | hide | past | pdf | discuss
3535. All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning (arxiv.org)
3 points by gkswamy98 on Mar 8, 2025 | hide | past | pdf | discuss
3536. Karatsuba Matrix Multiplication and Its Efficient Hardware Implementations (arxiv.org)
5 points by emacs28 on Mar 8, 2025 | hide | past | pdf | discuss
3537. Smaller but Better: Unifying Layout Generation with Smaller LLMs (arxiv.org)
24 points by PaulHoule on Mar 8, 2025 | hide | past | pdf | 3 comments
3538. RingFormer: Rethinking Recurrent Transformer with Adaptive Level Signals (arxiv.org)
3 points by PaulHoule on Mar 7, 2025 | hide | past | pdf | discuss
3539. Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence (arxiv.org)
1 point by PaulHoule on Mar 7, 2025 | hide | past | pdf | discuss
3540. Deep Learning Is Not So Mysterious or Different (arxiv.org)
1 point by thoughtpeddler on Mar 7, 2025 | hide | past | pdf | discuss