about
Stories from May 6, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. DoomArena: A Framework for Testing AI Agents Against Evolving Security Threats (arxiv.org)
10 points by PaulHoule on May 6, 2025 | hide | past | pdf | 2 comments
2. Don't be lazy: CompleteP enables compute-efficient deep transformers (arxiv.org)
4 points by nsdey on May 6, 2025 | hide | past | pdf | discuss
3. Towards Dataset Copyright Evasion Attack Against Personalized Diffusion Models (arxiv.org)
3 points by badmonster on May 6, 2025 | hide | past | pdf | discuss
4. Unveiling the Hidden: Movie Genre and User Bias in Spoiler Detection (arxiv.org)
3 points by PaulHoule on May 6, 2025 | hide | past | pdf | discuss
5. (How) Do reasoning models reason? (arxiv.org)
3 points by YeGoblynQueenne on May 6, 2025 | hide | past | pdf | discuss
6. Language Representations Can Be What Recommenders Need: Findings and Potentials (arxiv.org)
2 points by PaulHoule on May 6, 2025 | hide | past | pdf | discuss
7. HalluMix Benchmark: Detecting Hallucinations in Real-World Scenarios (arxiv.org)
2 points by rancar2 on May 6, 2025 | hide | past | pdf | discuss
8. Evaluating Frontier Models for Stealth and Situational Awareness (arxiv.org)
2 points by badmonster on May 6, 2025 | hide | past | pdf | discuss
9. Cost-of-Pass: An Economic Framework for Evaluating Language Models (arxiv.org)
1 point by JumpCrisscross on May 6, 2025 | hide | past | pdf | discuss