|
|
Stories from August 1, 2024 (UTC)
|
| 1. |
Baidu's Improving Retrieval Augmented Language Model with Self-Reasoning (arxiv.org) |
|
66 points by a-s-k-af on Aug 1, 2024 | hide | past | pdf | 4 comments
|
| 2. |
Beating GPT-4o and Claude 3.5 on SWE-bench Lite through repeated sampling (arxiv.org) |
|
5 points by aglazer on Aug 1, 2024 | hide | past | pdf | discuss
|
| 3. |
Human-Like Episodic Memory for Infinite Context LLMs (arxiv.org) |
|
2 points by wslh on Aug 1, 2024 | hide | past | pdf | discuss
|
| 4. |
Things Come from Having Many Good Models (arxiv.org) |
|
2 points by PaulHoule on Aug 1, 2024 | hide | past | pdf | discuss
|
| 5. |
Meta-Rewarding Language Models:Self-Improving Alignment with LLM-as-a-Meta-Judge (arxiv.org) |
|
2 points by sssummer on Aug 1, 2024 | hide | past | pdf | discuss
|
| 6. |
Algorithmic Language Models with Neurally Compiled Libraries (arxiv.org) |
|
2 points by PaulHoule on Aug 1, 2024 | hide | past | pdf | discuss
|
| 7. |
Eyeballvul: A future-proof benchmark for vulnerability detection in the wild (arxiv.org) |
|
1 point by wslh on Aug 1, 2024 | hide | past | pdf | discuss
|
| 8. |
The Llama 3 Herd of Models (arxiv.org) |
|
1 point by GaggiX on Aug 1, 2024 | hide | past | pdf | discuss
|
| 9. |
From pixels to planning: scale-free active inference (arxiv.org) |
|
1 point by vlotar on Aug 1, 2024 | hide | past | pdf | discuss
|
| 10. |
Do AI Safety Benchmarks Measure Safety Progress? (arxiv.org) |
|
1 point by hendrycks on Aug 1, 2024 | hide | past | pdf | discuss
|
|