| 3451. |
Accurate INT8 Training Through Dynamic Block-Level Fallback (arxiv.org) |
|
3 points by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
|
| 3452. |
EvoBPE: Evolutionary Protein Sequence Tokenization (arxiv.org) |
|
1 point by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
|
| 3453. |
Bold: Boolean Logic Deep Learning (arxiv.org) |
|
2 points by cubefox on Mar 24, 2025 | hide | past | pdf | 1 comment
|
| 3454. |
Sentence-Level Reward Model Can Better Aligning LLM from Human Preference (arxiv.org) |
|
1 point by PaulHoule on Mar 24, 2025 | hide | past | pdf | discuss
|
| 3455. |
Every Flop Counts: Scaling a 300B LLM Without Premium GPUs (arxiv.org) |
|
117 points by bretpiatt on Mar 24, 2025 | hide | past | pdf | 9 comments
|
| 3456. |
Can Large Vision Language Models Read Maps Like a Human? (arxiv.org) |
|
9 points by thinkingemote on Mar 24, 2025 | hide | past | pdf | 1 comment
|
| 3457. |
Self-Optimization of Hopfield Networks: The Creativity of Unsupervised Learning (arxiv.org) |
|
1 point by pizza on Mar 24, 2025 | hide | past | pdf | discuss
|
| 3458. |
Natural Quantization of Neural Networks (arxiv.org) |
|
2 points by xenophonf on Mar 23, 2025 | hide | past | pdf | discuss
|
| 3459. |
Stick to Facts: Towards Fidelity-Oriented Product Description Generation (arxiv.org) |
|
1 point by PaulHoule on Mar 23, 2025 | hide | past | pdf | discuss
|
| 3460. |
Exploring Hidden Reasoning Process of Large Language Models by Misleading Them (arxiv.org) |
|
8 points by belter on Mar 23, 2025 | hide | past | pdf | discuss
|
| 3461. |
Can Large Vision Language Models Read Maps Like a Human? (arxiv.org) |
|
3 points by eamag on Mar 23, 2025 | hide | past | pdf | discuss
|
| 3462. |
Stop using the elbow criterion for k-means (arxiv.org) |
|
79 points by Anon84 on Mar 23, 2025 | hide | past | pdf | 28 comments
|
| 3463. |
Can AI Compress Like a Genius? (arxiv.org) |
|
2 points by handfuloflight on Mar 23, 2025 | hide | past | pdf | discuss
|
| 3464. |
Revisiting semi-supervised learning in the era of foundation models (arxiv.org) |
|
2 points by PaulHoule on Mar 22, 2025 | hide | past | pdf | discuss
|
| 3465. |
Quantitative Finance: Kronecker-Factored Approximate Curvature Deep Hedging (arxiv.org) |
|
5 points by walterbell on Mar 22, 2025 | hide | past | pdf | discuss
|
| 3466. |
Interpreting the Repeated Token Phenomenon in Large Language Models (arxiv.org) |
|
2 points by PaulHoule on Mar 21, 2025 | hide | past | pdf | discuss
|
| 3467. |
ϕ -Decoding: Adaptive Foresight Sampling for Balanced Infer-Time Explore/Exploit (arxiv.org) |
|
2 points by cma on Mar 21, 2025 | hide | past | pdf | discuss
|
| 3468. |
Pen and Paper Exercises in Machine Learning (2022) (arxiv.org) |
|
413 points by ibobev on Mar 21, 2025 | hide | past | pdf | 58 comments
|
| 3469. |
Fin-R1: A Large Language Model for Financial Reasoning Through RL (arxiv.org) |
|
2 points by pama on Mar 21, 2025 | hide | past | pdf | 1 comment
|
| 3470. |
Compute Optimal Scaling of Skills: Knowledge vs. Reasoning (arxiv.org) |
|
2 points by pama on Mar 21, 2025 | hide | past | pdf | discuss
|
| 3471. |
Measuring AI Ability to Complete Long Tasks (arxiv.org) |
|
4 points by mefengl on Mar 21, 2025 | hide | past | pdf | discuss
|
| 3472. |
The Curse of Depth in Large Language Models (arxiv.org) |
|
1 point by veryluckyxyz on Mar 21, 2025 | hide | past | pdf | discuss
|
| 3473. |
SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversion (arxiv.org) |
|
66 points by prats226 on Mar 21, 2025 | hide | past | pdf | 12 comments
|
| 3474. |
Empowering LLMs for Time Series Forecasting with Temporal Patterns and Semantics (arxiv.org) |
|
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3475. |
Why Do Multi-Agent LLM Systems Fail? (arxiv.org) |
|
1 point by amrrs on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3476. |
WinClick: GUI Grounding with Multimodal Large Language Models (arxiv.org) |
|
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3477. |
Aardvark weather: end-to-end data-driven weather forecasting (arxiv.org) |
|
2 points by cyberlimerence on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3478. |
Tapered Off-Policy Reinforce: Stable and Efficient RL for LLMs (arxiv.org) |
|
2 points by pama on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3479. |
Measuring AI Ability to Complete Long Tasks (arxiv.org) |
|
5 points by mellosouls on Mar 20, 2025 | hide | past | pdf | discuss
|
| 3480. |
An Inference and Training Framework for WiFi-Based Human Activity Recognition (arxiv.org) |
|
1 point by PaulHoule on Mar 20, 2025 | hide | past | pdf | discuss
|
| More |