| 3661. |
SPIRE: Semantic Prompt-Driven Image Restoration (2024) (arxiv.org) |
|
3 points by fzliu on Feb 16, 2025 | hide | past | pdf | discuss
|
| 3662. |
Scaling Test-Time Compute Can Be More Effective Than Scaling Parameters (2024) (arxiv.org) |
|
2 points by fzliu on Feb 16, 2025 | hide | past | pdf | discuss
|
| 3663. |
The hierarchy in HNSW is not necessary in high dimensions (arxiv.org) |
|
5 points by blaise-muhirwa on Feb 15, 2025 | hide | past | pdf | 1 comment
|
| 3664. |
Direct Ascent Synthesis: Hidden Generative Capabilities in Discriminative Models (arxiv.org) |
|
2 points by kynez on Feb 15, 2025 | hide | past | pdf | discuss
|
| 3665. |
What makes math problems hard for reinforcement learning: a case study (arxiv.org) |
|
2 points by rntn on Feb 15, 2025 | hide | past | pdf | discuss
|
| 3666. |
The Danger of Overthinking: The Reasoning-Action Dilemma in Agentic Tasks (arxiv.org) |
|
2 points by stared on Feb 15, 2025 | hide | past | pdf | discuss
|
| 3667. |
Can We Trust AI Benchmarks? A Review of Current Issues in AI Evaluation (arxiv.org) |
|
23 points by rntn on Feb 15, 2025 | hide | past | pdf | 4 comments
|
| 3668. |
Geometry of Prompting: Distinct Mechanisms of Task Adaptation in Language Models (arxiv.org) |
|
2 points by PlatinumSoup on Feb 15, 2025 | hide | past | pdf | discuss
|
| 3669. |
Competitive Programming with Large Reasoning Models (arxiv.org) |
|
1 point by xdavidliu on Feb 15, 2025 | hide | past | pdf | discuss
|
| 3670. |
Matryoshka Quantization (arxiv.org) |
|
3 points by ofou on Feb 14, 2025 | hide | past | pdf | discuss
|
| 3671. |
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
1 point by azhenley on Feb 14, 2025 | hide | past | pdf | discuss
|
| 3672. |
Commercial LLM Agents Are Already Vulnerable to Simple yet Dangerous Attacks (arxiv.org) |
|
1 point by TaurenHunter on Feb 14, 2025 | hide | past | pdf | discuss
|
| 3673. |
Benchmarking vision-language models on OCR in dynamic video environments (arxiv.org) |
|
142 points by ashu_trv on Feb 14, 2025 | hide | past | pdf | 58 comments
|
| 3674. |
High-Throughput SAT Sampling (arxiv.org) |
|
2 points by ahsillyme on Feb 14, 2025 | hide | past | pdf | discuss
|
| 3675. |
LM2: Large Memory Models (arxiv.org) |
|
110 points by fzliu on Feb 13, 2025 | hide | past | pdf | 30 comments
|
| 3676. |
Distillation Scaling Laws (arxiv.org) |
|
3 points by fzliu on Feb 13, 2025 | hide | past | pdf | discuss
|
| 3677. |
NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
3 points by asah on Feb 13, 2025 | hide | past | pdf | 1 comment
|
| 3678. |
Decoding AI Judgment: How LLMs Assess News Credibility and Bias (arxiv.org) |
|
1 point by alphadelphi on Feb 13, 2025 | hide | past | pdf | discuss
|
| 3679. |
Vision-Language Models vs. Traditional OCR in Video – New Benchmark (arxiv.org) |
|
6 points by ashu_trv on Feb 13, 2025 | hide | past | pdf | 1 comment
|
| 3680. |
NoLima: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
3 points by apsec112 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3681. |
Competitive programming with large language models (arxiv.org) |
|
2 points by highfrequency on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3682. |
Mixture-of-Agents Enhances Large Language Model Capabilities (arxiv.org) |
|
2 points by wluk on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3683. |
Automated Capability Discovery via Foundation Model Self-Exploration (arxiv.org) |
|
63 points by f14t on Feb 12, 2025 | hide | past | pdf | 14 comments
|
| 3684. |
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving [pdf] (arxiv.org) |
|
6 points by bikenaga on Feb 12, 2025 | hide | past | pdf | 2 comments
|
| 3685. |
We Can't Understand AI Using Our Existing Vocabulary [pdf] (arxiv.org) |
|
4 points by bikenaga on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3686. |
Emergent Response Planning in LLM (arxiv.org) |
|
1 point by Jimmc414 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3687. |
Competitive Programming with Large Reasoning Models (arxiv.org) |
|
6 points by z7 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3688. |
RelBench: A Benchmark for Deep Learning on Relational Databases (PDF) (arxiv.org) |
|
2 points by gk1 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 3689. |
Matryoshka Quantization (arxiv.org) |
|
4 points by qianli_cs on Feb 12, 2025 | hide | past | pdf | 1 comment
|
| 3690. |
A Survey on Large Language Models (2025) (arxiv.org) |
|
1 point by OutOfHere on Feb 12, 2025 | hide | past | pdf | discuss
|
| More |