about
931. Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction (arxiv.org)
57 points by root-parent 130 days ago | hide | past | pdf | 9 comments
932. Spreadsheet-RL: Advancing LLM Agents on Realistic Spreadsheet Tasks (arxiv.org)
3 points by ankitg12 130 days ago | hide | past | pdf | discuss
933. Stateful Inference for Low-Latency Multi-Agent Tool Calling (arxiv.org)
2 points by logotype 130 days ago | hide | past | pdf | 1 comment
934. All of human cooking compressed into 2 megabytes (arxiv.org)
444 points by josefchen 130 days ago | hide | past | pdf | 174 comments
935. Tool-schema compression enables agentic RAG under constrained context budgets (arxiv.org)
2 points by Sakizli 130 days ago | hide | past | pdf | 1 comment
936. How sure is the activation oracle? (arxiv.org)
1 point by evilscript 130 days ago | hide | past | pdf | discuss
937. Advancing Mathematics Research with AI-Driven Formal Proof Search (arxiv.org)
1 point by mrkn1 130 days ago | hide | past | pdf | discuss
938. FML-Bench: A Controlled Study of AI Research Agent Strategies (arxiv.org)
1 point by matt_d 130 days ago | hide | past | pdf | discuss
939. Barriers to Complexity-Theoretic Proofs That "AGI" Using ML Is Impossible (arxiv.org)
4 points by mike_uoftdcs 131 days ago | hide | past | pdf | discuss
940. Agentic Harness Engineering (arxiv.org)
3 points by cobblr_mosaic 131 days ago | hide | past | pdf | discuss
941. SkillOpt: Executive Strategy for Self-Evolving Agent Skills (arxiv.org)
4 points by theaniketmaurya 131 days ago | hide | past | pdf | discuss
942. Polar: Agentic RL on Any Harness at Scale (arxiv.org)
3 points by Brajeshwar 131 days ago | hide | past | pdf | discuss
943. A sleep-like consolidation mechanism for LLMs (arxiv.org)
212 points by juxtapose 131 days ago | hide | past | pdf | 140 comments
944. Quest: Training Frontier Deep Research Agents with Synthetic Tasks (arxiv.org)
2 points by Brajeshwar 131 days ago | hide | past | pdf | discuss
945. Investigating how prompt politeness affects LLM accuracy (2025) (arxiv.org)
156 points by KnuthIsGod 131 days ago | hide | past | pdf | 208 comments
946. Contrastive Decoding Diffing: Recovering Finetuning Data Without Weight Access (arxiv.org)
2 points by Timofeibu 131 days ago | hide | past | pdf | discuss
947. ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention (arxiv.org)
4 points by joesharratt29 131 days ago | hide | past | pdf | 1 comment
948. Continual Speaker Identity Unlearning with Minimal Interference (arxiv.org)
2 points by berlianta 131 days ago | hide | past | pdf | discuss
949. MileStone: A Multi-Objective Compiler Phase Ordering Framework (arxiv.org)
1 point by matt_d 131 days ago | hide | past | pdf | discuss
950. LLMs require curated context for reliable political fact-checking (arxiv.org)
3 points by teleforce 132 days ago | hide | past | pdf | discuss
951. Advancing mathematics research with AI-driven formal proof search (arxiv.org)
2 points by azhenley 132 days ago | hide | past | pdf | discuss
952. Evaluating Large Language Models in a Complex Hidden Role Game (arxiv.org)
1 point by Brajeshwar 132 days ago | hide | past | pdf | discuss
953. SkillOpt: Executive Strategy for Self-Evolving Agent Skills (arxiv.org)
4 points by simonpure 132 days ago | hide | past | pdf | discuss
954. A Language for Describing Agentic LLM Contexts (arxiv.org)
4 points by mpweiher 133 days ago | hide | past | pdf | discuss
955. Advancing Mathematics Research with AI-Driven Formal Proof Search (arxiv.org)
3 points by tamnd 133 days ago | hide | past | pdf | discuss
956. Constraint Decay: The Fragility of LLM Agents in Back End Code Generation (arxiv.org)
287 points by wek 133 days ago | hide | past | pdf | 197 comments
957. SSV: Sparse Speculative Verification for Efficient LLM Inference (arxiv.org)
4 points by matt_d 133 days ago | hide | past | pdf | discuss
958. Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection (arxiv.org)
2 points by matt_d 134 days ago | hide | past | pdf | discuss
959. Customizing an LLM for Enterprise Software Engineering (arxiv.org)
4 points by daureg 134 days ago | hide | past | pdf | discuss
960. Agentic Compilation: Reducing LLM Rerun Costs (arxiv.org)
3 points by rebekkamikkoa 134 days ago | hide | past | pdf | discuss