about
5371. Forecasting the Future: Advancements in Large Meteorological Models (arxiv.org)
2 points by PaulHoule on Apr 18, 2024 | hide | past | pdf | discuss
5372. What are human values, and how do we align AI to them? (arxiv.org)
2 points by tosh on Apr 18, 2024 | hide | past | pdf | discuss
5373. Chinchilla Scaling: A replication attempt (arxiv.org)
124 points by tosh on Apr 18, 2024 | hide | past | pdf | 68 comments
5374. Chinchilla Scaling: A Replication Attempt (arxiv.org)
2 points by chewxy on Apr 17, 2024 | hide | past | pdf | discuss
5375. Confidential Federated Computations (arxiv.org)
1 point by tiziano88 on Apr 17, 2024 | hide | past | pdf | discuss
5376. Collapse of self-trained language models (arxiv.org)
92 points by PaulHoule on Apr 17, 2024 | hide | past | pdf | 30 comments
5377. Show Your Work with Confidence: Confidence Bands for Tuning Curves (arxiv.org)
1 point by nicholaslourie on Apr 17, 2024 | hide | past | pdf | discuss
5378. What are human values, and how do we align AI to them?[pdf] (arxiv.org)
1 point by kelseyfrog on Apr 17, 2024 | hide | past | pdf | discuss
5379. Chinchilla Scaling: A Replication Attempt (arxiv.org)
6 points by apsec112 on Apr 17, 2024 | hide | past | pdf | 1 comment
5380. Long-form music generation with latent diffusion (arxiv.org)
1 point by chaosprint on Apr 17, 2024 | hide | past | pdf | discuss
5381. The illusion of state in state-space models (arxiv.org)
4 points by canjobear on Apr 17, 2024 | hide | past | pdf | discuss
5382. Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length (arxiv.org)
168 points by amichail on Apr 16, 2024 | hide | past | pdf | 28 comments
5383. Writing Wikipedia-Like Articles from Scratch with LLMs (arxiv.org)
1 point by talonx on Apr 16, 2024 | hide | past | pdf | discuss
5384. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory (arxiv.org)
2 points by abhinavk on Apr 16, 2024 | hide | past | pdf | discuss
5385. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022) (arxiv.org)
2 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
5386. TransformerFAM: Feedback attention is working memory (arxiv.org)
4 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
5387. Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length (arxiv.org)
4 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss
5388. A Model-Based and Imitation Learning Deep Reinforcement Learning Hybrid (arxiv.org)
2 points by lucaspauker on Apr 16, 2024 | hide | past | pdf | discuss
5389. ResearchAgent: Iterative Research Idea Generation Using LLMs (arxiv.org)
124 points by milliondreams on Apr 16, 2024 | hide | past | pdf | 63 comments
5390. Why do small language models underperform? (arxiv.org)
4 points by tosh on Apr 15, 2024 | hide | past | pdf | 1 comment
5391. CodecLM: Aligning Language Models with Tailored Synthetic Data (arxiv.org)
2 points by asah on Apr 15, 2024 | hide | past | pdf | discuss
5392. Can Large Language Models Reason and Plan? (arxiv.org)
4 points by mnk47 on Apr 15, 2024 | hide | past | pdf | 1 comment
5393. You Need to Pay Better Attention (arxiv.org)
5 points by tosh on Apr 14, 2024 | hide | past | pdf | discuss
5394. CodecLM: Aligning Language Models with Tailored Synthetic Data (arxiv.org)
2 points by milliondreams on Apr 14, 2024 | hide | past | pdf | discuss
5395. ChatGPT Can Predict the Future Telling Stories Set in the Future About the Past (arxiv.org)
29 points by rntn on Apr 14, 2024 | hide | past | pdf | 8 comments
5396. Rho-1: Not All Tokens Are What You Need (arxiv.org)
1 point by jacquesm on Apr 14, 2024 | hide | past | pdf | discuss
5397. Power Hungry Processing: Watts Driving the Cost of AI Deployment? (arxiv.org)
2 points by todsacerdoti on Apr 14, 2024 | hide | past | pdf | discuss
5398. Chinchilla Debunked (arxiv.org)
3 points by andai on Apr 13, 2024 | hide | past | pdf | 1 comment
5399. Algorithmic Collusion by Large Language Models (arxiv.org)
3 points by mooreds on Apr 13, 2024 | hide | past | pdf | discuss
5400. Mechanics of Next Token Prediction with Self-Attention (arxiv.org)
13 points by georgehill on Apr 13, 2024 | hide | past | pdf | discuss