about
Stories from October 21, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. Why can't transformers learn multiplication? (arxiv.org)
161 points by PaulHoule 348 days ago | hide | past | pdf | 107 comments
2. Binary Retrieval-Augmented Reward Mitigates Hallucinations (arxiv.org)
44 points by MarlonPro 348 days ago | hide | past | pdf | 3 comments
3. OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training (arxiv.org)
11 points by PaulHoule 348 days ago | hide | past | pdf | discuss
4. Prompt Baking (arxiv.org)
8 points by jxmorris12 348 days ago | hide | past | pdf | 1 comment
5. Tensor Logic: The Language of AI (arxiv.org)
5 points by fofoz 348 days ago | hide | past | pdf | discuss
6. Reasoning with Sampling: Your Base Model Is Smarter Than You Think (arxiv.org)
3 points by Anon84 348 days ago | hide | past | pdf | discuss
7. Evaluating Agentic Cybersecurity in Attack/Defense CTFs: Offensive Is Not Better (arxiv.org)
2 points by vmayoral 348 days ago | hide | past | pdf | 1 comment
8. Qwen Language Confusion Gate (arxiv.org)
2 points by CollinZ 348 days ago | hide | past | pdf | discuss
9. Knowledge Transfer from High-Resource to Low-Resource Languages for Code LLMs (2023) (arxiv.org)
1 point by peatmoss 347 days ago | hide | past | pdf | discuss
10. Explaining Why Hallucinate Large Language Models (arxiv.org)
1 point by gagan30 348 days ago | hide | past | pdf | 1 comment
11. Verifiable ML Without Determinism: Tolerance-Aware Optimistic Verification (arxiv.org)
1 point by cnyaojz 348 days ago | hide | past | pdf | 1 comment