| 5371. |
Forecasting the Future: Advancements in Large Meteorological Models (arxiv.org) |
|
2 points by PaulHoule on Apr 18, 2024 | hide | past | pdf | discuss
|
| 5372. |
What are human values, and how do we align AI to them? (arxiv.org) |
|
2 points by tosh on Apr 18, 2024 | hide | past | pdf | discuss
|
| 5373. |
Chinchilla Scaling: A replication attempt (arxiv.org) |
|
124 points by tosh on Apr 18, 2024 | hide | past | pdf | 68 comments
|
| 5374. |
Chinchilla Scaling: A Replication Attempt (arxiv.org) |
|
2 points by chewxy on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5375. |
Confidential Federated Computations (arxiv.org) |
|
1 point by tiziano88 on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5376. |
Collapse of self-trained language models (arxiv.org) |
|
92 points by PaulHoule on Apr 17, 2024 | hide | past | pdf | 30 comments
|
| 5377. |
Show Your Work with Confidence: Confidence Bands for Tuning Curves (arxiv.org) |
|
1 point by nicholaslourie on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5378. |
What are human values, and how do we align AI to them?[pdf] (arxiv.org) |
|
1 point by kelseyfrog on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5379. |
Chinchilla Scaling: A Replication Attempt (arxiv.org) |
|
6 points by apsec112 on Apr 17, 2024 | hide | past | pdf | 1 comment
|
| 5380. |
Long-form music generation with latent diffusion (arxiv.org) |
|
1 point by chaosprint on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5381. |
The illusion of state in state-space models (arxiv.org) |
|
4 points by canjobear on Apr 17, 2024 | hide | past | pdf | discuss
|
| 5382. |
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length (arxiv.org) |
|
168 points by amichail on Apr 16, 2024 | hide | past | pdf | 28 comments
|
| 5383. |
Writing Wikipedia-Like Articles from Scratch with LLMs (arxiv.org) |
|
1 point by talonx on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5384. |
LLM in a Flash: Efficient Large Language Model Inference with Limited Memory (arxiv.org) |
|
2 points by abhinavk on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5385. |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022) (arxiv.org) |
|
2 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5386. |
TransformerFAM: Feedback attention is working memory (arxiv.org) |
|
4 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5387. |
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length (arxiv.org) |
|
4 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5388. |
A Model-Based and Imitation Learning Deep Reinforcement Learning Hybrid (arxiv.org) |
|
2 points by lucaspauker on Apr 16, 2024 | hide | past | pdf | discuss
|
| 5389. |
ResearchAgent: Iterative Research Idea Generation Using LLMs (arxiv.org) |
|
124 points by milliondreams on Apr 16, 2024 | hide | past | pdf | 63 comments
|
| 5390. |
Why do small language models underperform? (arxiv.org) |
|
4 points by tosh on Apr 15, 2024 | hide | past | pdf | 1 comment
|
| 5391. |
CodecLM: Aligning Language Models with Tailored Synthetic Data (arxiv.org) |
|
2 points by asah on Apr 15, 2024 | hide | past | pdf | discuss
|
| 5392. |
Can Large Language Models Reason and Plan? (arxiv.org) |
|
4 points by mnk47 on Apr 15, 2024 | hide | past | pdf | 1 comment
|
| 5393. |
You Need to Pay Better Attention (arxiv.org) |
|
5 points by tosh on Apr 14, 2024 | hide | past | pdf | discuss
|
| 5394. |
CodecLM: Aligning Language Models with Tailored Synthetic Data (arxiv.org) |
|
2 points by milliondreams on Apr 14, 2024 | hide | past | pdf | discuss
|
| 5395. |
ChatGPT Can Predict the Future Telling Stories Set in the Future About the Past (arxiv.org) |
|
29 points by rntn on Apr 14, 2024 | hide | past | pdf | 8 comments
|
| 5396. |
Rho-1: Not All Tokens Are What You Need (arxiv.org) |
|
1 point by jacquesm on Apr 14, 2024 | hide | past | pdf | discuss
|
| 5397. |
Power Hungry Processing: Watts Driving the Cost of AI Deployment? (arxiv.org) |
|
2 points by todsacerdoti on Apr 14, 2024 | hide | past | pdf | discuss
|
| 5398. |
Chinchilla Debunked (arxiv.org) |
|
3 points by andai on Apr 13, 2024 | hide | past | pdf | 1 comment
|
| 5399. |
Algorithmic Collusion by Large Language Models (arxiv.org) |
|
3 points by mooreds on Apr 13, 2024 | hide | past | pdf | discuss
|
| 5400. |
Mechanics of Next Token Prediction with Self-Attention (arxiv.org) |
|
13 points by georgehill on Apr 13, 2024 | hide | past | pdf | discuss
|
| More |