about
721. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering (arxiv.org)
2 points by wek 99 days ago | hide | past | pdf | discuss
722. Tapered Language Models (arxiv.org)
2 points by sonabinu 99 days ago | hide | past | pdf | discuss
723. Red-Teaming the Agentic Red-Team (arxiv.org)
3 points by infwhispers 100 days ago | hide | past | pdf | discuss
724. Combining LLMs Rarely Beats the Best Single Model, I tested 67 frontier models (arxiv.org)
1 point by josefchen 100 days ago | hide | past | pdf | discuss
725. Mapping Networks: CVPR 2026 Best Paper Award Nominee (arxiv.org)
4 points by aurenvale 100 days ago | hide | past | pdf | 1 comment
726. A Structured Generation Framework for Transforming Scientific Papers into Patent (arxiv.org)
2 points by teleforce 100 days ago | hide | past | pdf | discuss
727. Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Trac (arxiv.org)
4 points by xiaoyu2006 100 days ago | hide | past | pdf | discuss
728. Reading AI Model Compilation in MLIR Through the Lens of Formal Theories (arxiv.org)
2 points by matt_d 100 days ago | hide | past | pdf | discuss
729. David vs. Goliath in Next Activity Prediction: Argmax vs. LSTM, Transformer, LLM (arxiv.org)
2 points by hramezani 101 days ago | hide | past | pdf | discuss
730. Autodata: An agentic data scientist to create high quality synthetic data (arxiv.org)
4 points by root-parent 101 days ago | hide | past | pdf | 1 comment
731. Wikipedia advocacy shapes LLM values (arxiv.org)
3 points by 50kIters 101 days ago | hide | past | pdf | discuss
732. The False Promise of Imitating Proprietary LLMs (arxiv.org)
2 points by handfuloflight 101 days ago | hide | past | pdf | discuss
733. IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures (arxiv.org)
3 points by NavinF 101 days ago | hide | past | pdf | discuss
734. World Action Models: A Survey (arxiv.org)
4 points by simonpure 101 days ago | hide | past | pdf | discuss
735. Code as Agent Harness (arxiv.org)
5 points by matt_d 101 days ago | hide | past | pdf | 1 comment
736. LLMs use "safety" specific neuron layers to identify vulnerabilities in code (arxiv.org)
5 points by summarity 101 days ago | hide | past | pdf | 3 comments
737. Submodular Context Selection as a Pluggable Engine for LLM Agents (arxiv.org)
2 points by Elof 102 days ago | hide | past | pdf | discuss
738. DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference (arxiv.org)
2 points by yogthos 102 days ago | hide | past | pdf | discuss
739. The Promptware Kill Chain (arxiv.org)
3 points by wslh 102 days ago | hide | past | pdf | discuss
740. Qwen-AgentWorld: Language World Models for General Agents (arxiv.org)
199 points by ilreb 102 days ago | hide | past | pdf | 55 comments
741. Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild (arxiv.org)
3 points by MediaSquirrel 103 days ago | hide | past | pdf | discuss
742. Inference Compute Shapes Frontier LLM Evaluation (arxiv.org)
2 points by matt_d 103 days ago | hide | past | pdf | discuss
743. Concordia: JIT-Compiled Persistent-Kernel Checkpt for Fault-Tolerant Inference (arxiv.org)
2 points by matt_d 103 days ago | hide | past | pdf | discuss
744. Confidence estimation is a better metric than agreement for LLM judges (arxiv.org)
3 points by rapiddev 103 days ago | hide | past | pdf | discuss
745. PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations (arxiv.org)
3 points by Anon84 103 days ago | hide | past | pdf | discuss
746. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (2023) (arxiv.org)
25 points by Anon84 103 days ago | hide | past | pdf | 46 comments
747. Unlimited OCR Works (arxiv.org)
3 points by ilreb 103 days ago | hide | past | pdf | discuss
748. Tapered Language Models (arxiv.org)
3 points by E-Reverance 103 days ago | hide | past | pdf | discuss
749. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models (arxiv.org)
55 points by teleforce 103 days ago | hide | past | pdf | 9 comments
750. VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO (arxiv.org)
398 points by timhigins 103 days ago | hide | past | pdf | 205 comments