about
Stories from April 19, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. CaMeL: Defeating Prompt Injections by Design (arxiv.org)
71 points by tomrod on Apr 19, 2025 | hide | past | pdf | 16 comments
2. Inferring the Phylogeny of Large Language Models (arxiv.org)
69 points by weinzierl on Apr 19, 2025 | hide | past | pdf | 6 comments
3. Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arxiv.org)
14 points by distalx on Apr 19, 2025 | hide | past | pdf | discuss
4. Do Reasoning Models Show Better Verbalized Calibration? (arxiv.org)
2 points by veryluckyxyz on Apr 19, 2025 | hide | past | pdf | discuss
5. Defeating Prompt Injections by Design (arxiv.org)
2 points by theptip on Apr 19, 2025 | hide | past | pdf | discuss
6. How to evaluate control measures for LLM agents? (arxiv.org)
2 points by handfuloflight on Apr 19, 2025 | hide | past | pdf | discuss
7. Using Small Models to Compare Language Learning and Tokenizer Performance (arxiv.org)
1 point by PaulHoule on Apr 19, 2025 | hide | past | pdf | discuss