about
511. Every Time I Hire a Linguist, Inference Costs Go Down (arxiv.org)
3 points by cwbuilds 66 days ago | hide | past | pdf | discuss
512. CryptanalysisBench: Can LLMs Do Cryptanalysis? (arxiv.org)
1 point by zdw 66 days ago | hide | past | pdf | discuss
513. Teaching agents to predict and pre-execute their next tool call (arxiv.org)
6 points by rotariuvladimir 67 days ago | hide | past | pdf | discuss
514. Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation (arxiv.org)
3 points by root-parent 67 days ago | hide | past | pdf | discuss
515. Visual prompt engineering for video models (arxiv.org)
4 points by root-parent 67 days ago | hide | past | pdf | discuss
516. Handbook.md shows that long policy documents do not reliably govern agents (arxiv.org)
325 points by spIrr 67 days ago | hide | past | pdf | 209 comments
517. Detecting CSAM Text-to-Image LoRAs from Weights (arxiv.org)
6 points by sbulaev 67 days ago | hide | past | pdf | discuss
518. Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation (arxiv.org)
4 points by tcp_handshaker 67 days ago | hide | past | pdf | discuss
519. Adaptive Agentic Attacks on LLM Vulnerability Detectors via Adversarial Comments (arxiv.org)
3 points by tcp_handshaker 67 days ago | hide | past | pdf | discuss
520. ProofCouncil: An LLM Agent for Solving Open Mathematical Problems (arxiv.org)
3 points by frozenseven 67 days ago | hide | past | pdf | 1 comment
521. Mapping CVEs to Mitre ATT&CK Techniques (arxiv.org)
3 points by adulau 67 days ago | hide | past | pdf | 1 comment
522. Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model (arxiv.org)
3 points by sbulaev 67 days ago | hide | past | pdf | discuss
523. The SpiNNaker2 chip: a many-core platform for brain-inspired computing (arxiv.org)
1 point by Jimmc414 67 days ago | hide | past | pdf | discuss
524. Distilling proprietary model reasoning into open-source search agents (arxiv.org)
5 points by cpard 67 days ago | hide | past | pdf | discuss
525. CryptanalysisBench: Can LLMs Do Cryptanalysis? (arxiv.org)
2 points by rvz 67 days ago | hide | past | pdf | discuss
526. "Uncensored" open LLMs are measurably more optimistic than their base models (arxiv.org)
43 points by oleczek 67 days ago | hide | past | pdf | 22 comments
527. Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents (arxiv.org)
2 points by zhinit 67 days ago | hide | past | pdf | discuss
528. Kimi Linear: An Expressive, Efficient Attention Architecture (2025) (arxiv.org)
306 points by ronfriedhaber 68 days ago | hide | past | pdf | 133 comments
529. LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference (arxiv.org)
2 points by binyu 68 days ago | hide | past | pdf | discuss
530. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language (arxiv.org)
2 points by sbulaev 68 days ago | hide | past | pdf | discuss
531. Neural Representation of Minimal Surfaces (arxiv.org)
2 points by E-Reverance 68 days ago | hide | past | pdf | discuss
532. Show HN: A 6M-token movable window on a single 46GB GPU (arxiv.org)
7 points by Wetime 68 days ago | hide | past | pdf | 16 comments
533. Protocol-Level Attacks on Agentic Commerce Platforms: Taxonomy and Defense (arxiv.org)
2 points by sbulaev 68 days ago | hide | past | pdf | discuss
534. Lost in Context: Addressing Context Anxiety in Large Language Models (arxiv.org)
2 points by StatsAreFun 68 days ago | hide | past | pdf | discuss
535. The Polynomial-Time Low-Degree Conjecture Is False (arxiv.org)
4 points by MarcoDewey 68 days ago | hide | past | pdf | discuss
536. Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core (arxiv.org)
2 points by Jimmc414 68 days ago | hide | past | pdf | discuss
537. Paper – Agent Memory (arxiv.org)
2 points by qspencer 69 days ago | hide | past | pdf | discuss
538. Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks (arxiv.org)
1 point by sbulaev 69 days ago | hide | past | pdf | discuss
539. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic (arxiv.org)
2 points by sbulaev 69 days ago | hide | past | pdf | discuss
540. Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains (arxiv.org)
2 points by MarcoDewey 69 days ago | hide | past | pdf | 1 comment