about
1561. Routing LLM queries using internal success predictions (70% cost reduction) (arxiv.org)
1 point by stansApprentice 234 days ago | hide | past | pdf | 3 comments
1562. SWE-AGI: benchmarking spec-driven software construction (arxiv.org)
1 point by mustaphah 234 days ago | hide | past | pdf | 1 comment
1563. RL on GPT-5 to write better kernels (arxiv.org)
4 points by atallahw 235 days ago | hide | past | pdf | 1 comment
1564. Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language (arxiv.org)
2 points by matt_d 235 days ago | hide | past | pdf | discuss
1565. HySparse: A Hybrid Sparse Attention Architecture (arxiv.org)
5 points by readitalready 235 days ago | hide | past | pdf | discuss
1566. Biases in the Blind Spot: Detecting What LLMs Fail to Mention (arxiv.org)
1 point by jari_mustonen 235 days ago | hide | past | pdf | discuss
1567. Evaluation of RAG Architectures for Policy Document Question Answering (arxiv.org)
1 point by PaulHoule 235 days ago | hide | past | pdf | discuss
1568. SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora (arxiv.org)
3 points by salkahfi 235 days ago | hide | past | pdf | discuss
1569. Opus: Towards Efficient and Principled Data Selection in LLM Pre-Training (arxiv.org)
2 points by onurkanbkrc 235 days ago | hide | past | pdf | discuss
1570. Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters (arxiv.org)
1 point by onurkanbkrc 235 days ago | hide | past | pdf | 1 comment
1571. Faster and Cheaper Computations with Randomized Numerical Linear Algebra (arxiv.org)
2 points by PaulHoule 235 days ago | hide | past | pdf | discuss
1572. NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models (arxiv.org)
13 points by chrsw 235 days ago | hide | past | pdf | discuss
1573. Grok4 sabotages shutdown 97% of the time,even if instructed not in system prompt (arxiv.org)
8 points by agenticagent 235 days ago | hide | past | pdf | 4 comments
1574. Biases in the Blind Spot: Detecting What LLMs Fail to Mention (arxiv.org)
4 points by typeofhuman 235 days ago | hide | past | pdf | 1 comment
1575. Attention Sinks and Compression Valleys in LLMs (arxiv.org)
1 point by alexkranias 236 days ago | hide | past | pdf | discuss
1576. Misconduct in Post-Selections and Deep Learning (2024) (arxiv.org)
3 points by bjourne 236 days ago | hide | past | pdf | discuss
1577. The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence (arxiv.org)
1 point by schmuhblaster 236 days ago | hide | past | pdf | discuss
1578. GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt [pdf] (arxiv.org)
2 points by janandonly 236 days ago | hide | past | pdf | discuss
1579. Harmless reward hacks generalize to shutdown evasion and dictatorship in GPT-4.1 (arxiv.org)
1 point by toliveistobuild 236 days ago | hide | past | pdf | 1 comment
1580. FullStack-Agent: Enhancing Agentic Full-Stack Web Coding (arxiv.org)
2 points by simonpure 237 days ago | hide | past | pdf | discuss
1581. Lightweight Memory Construction with Dynamic Evolution for LLM Agents (arxiv.org)
2 points by PaulHoule 237 days ago | hide | past | pdf | discuss
1582. Moltbook: Fast Response or Silence? (arxiv.org)
1 point by EagleEdge 237 days ago | hide | past | pdf | discuss
1583. Randomness in Agentic Evals (arxiv.org)
3 points by andre15silva 237 days ago | hide | past | pdf | discuss
1584. Large Language Model Reasoning Failures (arxiv.org)
1 point by mpweiher 237 days ago | hide | past | pdf | discuss
1585. We Should Separate Memorization from Copyright (arxiv.org)
1 point by 50kIters 237 days ago | hide | past | pdf | discuss
1586. Security audit of Browser Use: prompt injection, credential exfil, domain bypass (arxiv.org)
2 points by tiny-automates 237 days ago | hide | past | pdf | 1 comment
1587. Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs (arxiv.org)
544 points by tiny-automates 237 days ago | hide | past | pdf | 366 comments
1588. Large Language Model Reasoning Failures (arxiv.org)
3 points by belter 238 days ago | hide | past | pdf | discuss
1589. Shared LoRA Subspaces for Almost Strict Continual Learning (arxiv.org)
1 point by unisub_guy 238 days ago | hide | past | pdf | 1 comment
1590. Towards Understanding What State Space Models Learn About Code (arxiv.org)
1 point by belter 238 days ago | hide | past | pdf | discuss