| 1. |
Automated Capability Discovery via Foundation Model Self-Exploration (arxiv.org) |
|
63 points by f14t on Feb 12, 2025 | hide | past | pdf | 14 comments
|
| 2. |
Competitive Programming with Large Reasoning Models (arxiv.org) |
|
16 points by t55 on Feb 12, 2025 | hide | past | pdf | 1 comment
|
| 3. |
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving [pdf] (arxiv.org) |
|
6 points by bikenaga on Feb 12, 2025 | hide | past | pdf | 2 comments
|
| 4. |
Competitive Programming with Large Reasoning Models (arxiv.org) |
|
6 points by z7 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 5. |
We Can't Understand AI Using Our Existing Vocabulary [pdf] (arxiv.org) |
|
4 points by bikenaga on Feb 12, 2025 | hide | past | pdf | discuss
|
| 6. |
Matryoshka Quantization (arxiv.org) |
|
4 points by qianli_cs on Feb 12, 2025 | hide | past | pdf | 1 comment
|
| 7. |
NoLima: Long-Context Evaluation Beyond Literal Matching (arxiv.org) |
|
3 points by apsec112 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 8. |
Competitive programming with large language models (arxiv.org) |
|
2 points by highfrequency on Feb 12, 2025 | hide | past | pdf | discuss
|
| 9. |
Mixture-of-Agents Enhances Large Language Model Capabilities (arxiv.org) |
|
2 points by wluk on Feb 12, 2025 | hide | past | pdf | discuss
|
| 10. |
RelBench: A Benchmark for Deep Learning on Relational Databases (PDF) (arxiv.org) |
|
2 points by gk1 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 11. |
Reducing the Transformer Architecture to a Minimum [pdf] (arxiv.org) |
|
2 points by DoctorOetker on Feb 12, 2025 | hide | past | pdf | discuss
|
| 12. |
Emergent Response Planning in LLM (arxiv.org) |
|
1 point by Jimmc414 on Feb 12, 2025 | hide | past | pdf | discuss
|
| 13. |
A Survey on Large Language Models (2025) (arxiv.org) |
|
1 point by OutOfHere on Feb 12, 2025 | hide | past | pdf | discuss
|