about
3661. SPIRE: Semantic Prompt-Driven Image Restoration (2024) (arxiv.org)
3 points by fzliu on Feb 16, 2025 | hide | past | pdf | discuss
3662. Scaling Test-Time Compute Can Be More Effective Than Scaling Parameters (2024) (arxiv.org)
2 points by fzliu on Feb 16, 2025 | hide | past | pdf | discuss
3663. The hierarchy in HNSW is not necessary in high dimensions (arxiv.org)
5 points by blaise-muhirwa on Feb 15, 2025 | hide | past | pdf | 1 comment
3664. Direct Ascent Synthesis: Hidden Generative Capabilities in Discriminative Models (arxiv.org)
2 points by kynez on Feb 15, 2025 | hide | past | pdf | discuss
3665. What makes math problems hard for reinforcement learning: a case study (arxiv.org)
2 points by rntn on Feb 15, 2025 | hide | past | pdf | discuss
3666. The Danger of Overthinking: The Reasoning-Action Dilemma in Agentic Tasks (arxiv.org)
2 points by stared on Feb 15, 2025 | hide | past | pdf | discuss
3667. Can We Trust AI Benchmarks? A Review of Current Issues in AI Evaluation (arxiv.org)
23 points by rntn on Feb 15, 2025 | hide | past | pdf | 4 comments
3668. Geometry of Prompting: Distinct Mechanisms of Task Adaptation in Language Models (arxiv.org)
2 points by PlatinumSoup on Feb 15, 2025 | hide | past | pdf | discuss
3669. Competitive Programming with Large Reasoning Models (arxiv.org)
1 point by xdavidliu on Feb 15, 2025 | hide | past | pdf | discuss
3670. Matryoshka Quantization (arxiv.org)
3 points by ofou on Feb 14, 2025 | hide | past | pdf | discuss
3671. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
1 point by azhenley on Feb 14, 2025 | hide | past | pdf | discuss
3672. Commercial LLM Agents Are Already Vulnerable to Simple yet Dangerous Attacks (arxiv.org)
1 point by TaurenHunter on Feb 14, 2025 | hide | past | pdf | discuss
3673. Benchmarking vision-language models on OCR in dynamic video environments (arxiv.org)
142 points by ashu_trv on Feb 14, 2025 | hide | past | pdf | 58 comments
3674. High-Throughput SAT Sampling (arxiv.org)
2 points by ahsillyme on Feb 14, 2025 | hide | past | pdf | discuss
3675. LM2: Large Memory Models (arxiv.org)
110 points by fzliu on Feb 13, 2025 | hide | past | pdf | 30 comments
3676. Distillation Scaling Laws (arxiv.org)
3 points by fzliu on Feb 13, 2025 | hide | past | pdf | discuss
3677. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
3 points by asah on Feb 13, 2025 | hide | past | pdf | 1 comment
3678. Decoding AI Judgment: How LLMs Assess News Credibility and Bias (arxiv.org)
1 point by alphadelphi on Feb 13, 2025 | hide | past | pdf | discuss
3679. Vision-Language Models vs. Traditional OCR in Video – New Benchmark (arxiv.org)
6 points by ashu_trv on Feb 13, 2025 | hide | past | pdf | 1 comment
3680. NoLima: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
3 points by apsec112 on Feb 12, 2025 | hide | past | pdf | discuss
3681. Competitive programming with large language models (arxiv.org)
2 points by highfrequency on Feb 12, 2025 | hide | past | pdf | discuss
3682. Mixture-of-Agents Enhances Large Language Model Capabilities (arxiv.org)
2 points by wluk on Feb 12, 2025 | hide | past | pdf | discuss
3683. Automated Capability Discovery via Foundation Model Self-Exploration (arxiv.org)
63 points by f14t on Feb 12, 2025 | hide | past | pdf | 14 comments
3684. Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving [pdf] (arxiv.org)
6 points by bikenaga on Feb 12, 2025 | hide | past | pdf | 2 comments
3685. We Can't Understand AI Using Our Existing Vocabulary [pdf] (arxiv.org)
4 points by bikenaga on Feb 12, 2025 | hide | past | pdf | discuss
3686. Emergent Response Planning in LLM (arxiv.org)
1 point by Jimmc414 on Feb 12, 2025 | hide | past | pdf | discuss
3687. Competitive Programming with Large Reasoning Models (arxiv.org)
6 points by z7 on Feb 12, 2025 | hide | past | pdf | discuss
3688. RelBench: A Benchmark for Deep Learning on Relational Databases (PDF) (arxiv.org)
2 points by gk1 on Feb 12, 2025 | hide | past | pdf | discuss
3689. Matryoshka Quantization (arxiv.org)
4 points by qianli_cs on Feb 12, 2025 | hide | past | pdf | 1 comment
3690. A Survey on Large Language Models (2025) (arxiv.org)
1 point by OutOfHere on Feb 12, 2025 | hide | past | pdf | discuss