|
|
Stories from June 3, 2023 (UTC)
|
| 1. |
Bytes are all you need: Transformers operating directly on file bytes (arxiv.org) |
|
263 points by pmoriarty on Jun 3, 2023 | hide | past | pdf | 96 comments
|
| 2. |
CodeCompose: A large-scale industrial deployment of AI-assisted code authoring (arxiv.org) |
|
165 points by azhenley on Jun 3, 2023 | hide | past | pdf | 70 comments
|
| 3. |
Brainformers: Trading Simplicity for Efficiency (arxiv.org) |
|
81 points by PaulHoule on Jun 3, 2023 | hide | past | pdf | 3 comments
|
| 4. |
Thought Cloning: Learning to think while acting by imitating human thinking (arxiv.org) |
|
57 points by gardenfelder on Jun 3, 2023 | hide | past | pdf | 38 comments
|
| 5. |
Grokking of Hierarchical Structure in Vanilla Transformers (arxiv.org) |
|
6 points by wwarner on Jun 3, 2023 | hide | past | pdf | discuss
|
| 6. |
The Curse of Recursion: Training on Generated Data Makes Models Forget (arxiv.org) |
|
5 points by YeGoblynQueenne on Jun 3, 2023 | hide | past | pdf | discuss
|
| 7. |
The Impact of Positional Encoding on Length Generalization in Transformers (arxiv.org) |
|
5 points by teleforce on Jun 3, 2023 | hide | past | pdf | discuss
|
| 8. |
Bytes Are All You Need: Transformers Operating Directly On File Bytes (arxiv.org) |
|
3 points by optimalsolver on Jun 3, 2023 | hide | past | pdf | discuss
|
| 9. |
LLM Itself Can Read and Generate CXR Images (arxiv.org) |
|
2 points by famouswaffles on Jun 3, 2023 | hide | past | pdf | discuss
|
| 10. |
OpenAI: How to Train Reasoning in LLMs (arxiv.org) |
|
2 points by EnnioEvo on Jun 3, 2023 | hide | past | pdf | discuss
|
| 11. |
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in LLMs (arxiv.org) |
|
1 point by pilingual on Jun 3, 2023 | hide | past | pdf | 1 comment
|
| 12. |
Stress Testing Social Reasoning in Large Language Models (arxiv.org) |
|
1 point by Jimmc414 on Jun 3, 2023 | hide | past | pdf | discuss
|
|