about
Stories from February 12, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. Automated Capability Discovery via Foundation Model Self-Exploration (arxiv.org)
63 points by f14t on Feb 12, 2025 | hide | past | pdf | 14 comments
2. Competitive Programming with Large Reasoning Models (arxiv.org)
16 points by t55 on Feb 12, 2025 | hide | past | pdf | 1 comment
3. Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving [pdf] (arxiv.org)
6 points by bikenaga on Feb 12, 2025 | hide | past | pdf | 2 comments
4. Competitive Programming with Large Reasoning Models (arxiv.org)
6 points by z7 on Feb 12, 2025 | hide | past | pdf | discuss
5. We Can't Understand AI Using Our Existing Vocabulary [pdf] (arxiv.org)
4 points by bikenaga on Feb 12, 2025 | hide | past | pdf | discuss
6. Matryoshka Quantization (arxiv.org)
4 points by qianli_cs on Feb 12, 2025 | hide | past | pdf | 1 comment
7. NoLima: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
3 points by apsec112 on Feb 12, 2025 | hide | past | pdf | discuss
8. Competitive programming with large language models (arxiv.org)
2 points by highfrequency on Feb 12, 2025 | hide | past | pdf | discuss
9. Mixture-of-Agents Enhances Large Language Model Capabilities (arxiv.org)
2 points by wluk on Feb 12, 2025 | hide | past | pdf | discuss
10. RelBench: A Benchmark for Deep Learning on Relational Databases (PDF) (arxiv.org)
2 points by gk1 on Feb 12, 2025 | hide | past | pdf | discuss
11. Reducing the Transformer Architecture to a Minimum [pdf] (arxiv.org)
2 points by DoctorOetker on Feb 12, 2025 | hide | past | pdf | discuss
12. Emergent Response Planning in LLM (arxiv.org)
1 point by Jimmc414 on Feb 12, 2025 | hide | past | pdf | discuss
13. A Survey on Large Language Models (2025) (arxiv.org)
1 point by OutOfHere on Feb 12, 2025 | hide | past | pdf | discuss