about
3451. Accurate INT8 Training Through Dynamic Block-Level Fallback (arxiv.org)
3 points by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
3452. EvoBPE: Evolutionary Protein Sequence Tokenization (arxiv.org)
1 point by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
3453. Bold: Boolean Logic Deep Learning (arxiv.org)
2 points by cubefox on Mar 24, 2025 | hide | past | pdf | 1 comment
3454. Sentence-Level Reward Model Can Better Aligning LLM from Human Preference (arxiv.org)
1 point by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
3455. Every Flop Counts: Scaling a 300B LLM Without Premium GPUs (arxiv.org)
117 points by bretpiatt on Mar 24, 2025 | hide | past | pdf | 9 comments
3456. Can Large Vision Language Models Read Maps Like a Human? (arxiv.org)
9 points by thinkingemote on Mar 24, 2025 | hide | past | pdf | 1 comment
3457. Self-Optimization of Hopfield Networks: The Creativity of Unsupervised Learning (arxiv.org)
1 point by pizza on Mar 24, 2025 | hide | past | pdf | discuss
3458. Natural Quantization of Neural Networks (arxiv.org)
2 points by xenophonf on Mar 23, 2025 | hide | past | pdf | discuss
3459. Stick to Facts: Towards Fidelity-Oriented Product Description Generation (arxiv.org)
1 point by PaulHoule on Mar 23, 2025 | hide | past | pdf | discuss
3460. Exploring Hidden Reasoning Process of Large Language Models by Misleading Them (arxiv.org)
8 points by belter on Mar 23, 2025 | hide | past | pdf | discuss
3461. Can Large Vision Language Models Read Maps Like a Human? (arxiv.org)
3 points by eamag on Mar 23, 2025 | hide | past | pdf | discuss
3462. Stop using the elbow criterion for k-means (arxiv.org)
79 points by Anon84 on Mar 23, 2025 | hide | past | pdf | 28 comments
3463. Can AI Compress Like a Genius? (arxiv.org)
2 points by handfuloflight on Mar 23, 2025 | hide | past | pdf | discuss
3464. Revisiting semi-supervised learning in the era of foundation models (arxiv.org)
2 points by PaulHoule on Mar 22, 2025 | hide | past | pdf | discuss
3465. Quantitative Finance: Kronecker-Factored Approximate Curvature Deep Hedging (arxiv.org)
5 points by walterbell on Mar 22, 2025 | hide | past | pdf | discuss
3466. Interpreting the Repeated Token Phenomenon in Large Language Models (arxiv.org)
2 points by PaulHoule on Mar 21, 2025 | hide | past | pdf | discuss
3467. ϕ -Decoding: Adaptive Foresight Sampling for Balanced Infer-Time Explore/Exploit (arxiv.org)
2 points by cma on Mar 21, 2025 | hide | past | pdf | discuss
3468. Pen and Paper Exercises in Machine Learning (2022) (arxiv.org)
413 points by ibobev on Mar 21, 2025 | hide | past | pdf | 58 comments
3469. Fin-R1: A Large Language Model for Financial Reasoning Through RL (arxiv.org)
2 points by pama on Mar 21, 2025 | hide | past | pdf | 1 comment
3470. Compute Optimal Scaling of Skills: Knowledge vs. Reasoning (arxiv.org)
2 points by pama on Mar 21, 2025 | hide | past | pdf | discuss
3471. Measuring AI Ability to Complete Long Tasks (arxiv.org)
4 points by mefengl on Mar 21, 2025 | hide | past | pdf | discuss
3472. The Curse of Depth in Large Language Models (arxiv.org)
1 point by veryluckyxyz on Mar 21, 2025 | hide | past | pdf | discuss
3473. SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversion (arxiv.org)
66 points by prats226 on Mar 21, 2025 | hide | past | pdf | 12 comments
3474. Empowering LLMs for Time Series Forecasting with Temporal Patterns and Semantics (arxiv.org)
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss
3475. Why Do Multi-Agent LLM Systems Fail? (arxiv.org)
1 point by amrrs on Mar 20, 2025 | hide | past | pdf | discuss
3476. WinClick: GUI Grounding with Multimodal Large Language Models (arxiv.org)
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss
3477. Aardvark weather: end-to-end data-driven weather forecasting (arxiv.org)
2 points by cyberlimerence on Mar 20, 2025 | hide | past | pdf | discuss
3478. Tapered Off-Policy Reinforce: Stable and Efficient RL for LLMs (arxiv.org)
2 points by pama on Mar 20, 2025 | hide | past | pdf | discuss
3479. Measuring AI Ability to Complete Long Tasks (arxiv.org)
5 points by mellosouls on Mar 20, 2025 | hide | past | pdf | discuss
3480. An Inference and Training Framework for WiFi-Based Human Activity Recognition (arxiv.org)
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss