about
Stories from July 29, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Handbook.md shows that long policy documents do not reliably govern agents (arxiv.org)
325 points by spIrr 66 days ago | hide | past | pdf | 209 comments
2. Teaching agents to predict and pre-execute their next tool call (arxiv.org)
6 points by rotariuvladimir 66 days ago | hide | past | pdf | discuss
3. Detecting CSAM Text-to-Image LoRAs from Weights (arxiv.org)
6 points by sbulaev 66 days ago | hide | past | pdf | discuss
4. Visual prompt engineering for video models (arxiv.org)
4 points by root-parent 66 days ago | hide | past | pdf | discuss
5. Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation (arxiv.org)
4 points by tcp_handshaker 66 days ago | hide | past | pdf | discuss
6. Every Time I Hire a Linguist, Inference Costs Go Down (arxiv.org)
3 points by cwbuilds 66 days ago | hide | past | pdf | discuss
7. Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation (arxiv.org)
3 points by root-parent 66 days ago | hide | past | pdf | discuss
8. Adaptive Agentic Attacks on LLM Vulnerability Detectors via Adversarial Comments (arxiv.org)
3 points by tcp_handshaker 66 days ago | hide | past | pdf | discuss
9. ProofCouncil: An LLM Agent for Solving Open Mathematical Problems (arxiv.org)
3 points by frozenseven 66 days ago | hide | past | pdf | 1 comment
10. Mapping CVEs to Mitre ATT&CK Techniques (arxiv.org)
3 points by adulau 66 days ago | hide | past | pdf | 1 comment
11. Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model (arxiv.org)
3 points by sbulaev 66 days ago | hide | past | pdf | discuss
12. CryptanalysisBench: Can LLMs Do Cryptanalysis? (arxiv.org)
1 point by zdw 66 days ago | hide | past | pdf | discuss