about
Stories from February 17, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Composition-RL: Compose Verifiable Prompts for Reinforcement Learning of LLMs (arxiv.org)
3 points by gmays 229 days ago | hide | past | pdf | discuss
2. Learning State-Tracking from Code Using Linear RNNs (arxiv.org)
2 points by jul8234 229 days ago | hide | past | pdf | 1 comment
3. A Survey of In-Context Reinforcement Learning (arxiv.org)
2 points by handfuloflight 229 days ago | hide | past | pdf | discuss
4. Training-Free Group Relative Policy Optimization (arxiv.org)
1 point by readitalready 229 days ago | hide | past | pdf | discuss
5. Randomness in Agentic Evals (arxiv.org)
1 point by andre15silva 229 days ago | hide | past | pdf | discuss
6. Hunt Globally (arxiv.org)
1 point by salkahfi 229 days ago | hide | past | pdf | discuss
7. Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises (arxiv.org)
1 point by salkahfi 229 days ago | hide | past | pdf | discuss