about
Stories from February 21, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. Some critical issues with the SWE-bench dataset (arxiv.org)
350 points by joshwa on Feb 21, 2025 | hide | past | pdf | 116 comments
2. Performance of Zero-Shot Time Series Foundation Models on Cloud Data (arxiv.org)
4 points by wanderingmind on Feb 21, 2025 | hide | past | pdf | discuss
3. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
2 points by azhenley on Feb 21, 2025 | hide | past | pdf | discuss
4. Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging (arxiv.org)
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
5. The Case for Cognitive-Dissonance-Aware Knowledge Updates in LLMs (arxiv.org)
2 points by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
6. Fire-Flyer AI-HPC: Cost-Effective Software-Hardware Co-Design for Deep Learning (arxiv.org)
2 points by doener on Feb 21, 2025 | hide | past | pdf | discuss
7. Too Noisy to Learn: Enhancing Data Quality for Code Review (arxiv.org)
1 point by PaulHoule on Feb 21, 2025 | hide | past | pdf | discuss
8. MLGym: A New Framework and Benchmark for Advancing AI Research Agents (arxiv.org)
1 point by jonbaer on Feb 21, 2025 | hide | past | pdf | discuss
9. Presumed Cultural Identity: How Names Shape LLM Responses (arxiv.org)
1 point by SerCe on Feb 21, 2025 | hide | past | pdf | discuss