about
Stories from August 7, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. HarnessOpt-Bench: Evaluating LLMs at Harness Optimization (arxiv.org)
3 points by wslh 57 days ago | hide | past | pdf | discuss
2. INT2 KV-cache quantization is getting surprisingly good (arxiv.org)
3 points by yunuyean 58 days ago | hide | past | pdf | discuss
3. MatrAIx: Simulating the world with 8.3B persona agents (arxiv.org)
3 points by anigbrowl 58 days ago | hide | past | pdf | discuss
4. AI Meet CAD: Beat Opus/Mythos 5, GPT 5.6 Sol on BenchCAD (arxiv.org)
2 points by SUPERustam 58 days ago | hide | past | pdf | 1 comment
5. TensorLift: Auto Extraction of ISA Semantics from Accelerator RTL via MLIR (arxiv.org)
1 point by matt_d 57 days ago | hide | past | pdf | discuss
6. The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss
7. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss
8. Benchmarking the Residual: What Long-Horizon Evals Add Beyond Short-Task Perf (arxiv.org)
1 point by matt_d 58 days ago | hide | past | pdf | discuss
9. SciCode-Verified: How Benchmark Defects Underestimated LLM Scientific-Coding (arxiv.org)
1 point by sbulaev 58 days ago | hide | past | pdf | discuss