about
5461. No "Zero-Shot" Without Exponential Data (arxiv.org)
1 point by passwordoops on Apr 8, 2024 | hide | past | pdf | discuss
5462. Is Model Collapse Inevitable? Breaking the Curse with Real and Synthetic Data (arxiv.org)
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
5463. Strum-LLM: Attributed and Structured Contrastive Summarization (arxiv.org)
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
5464. A Neuro-Inspired Topological Sparse Training Algorithm for Large Language Models (arxiv.org)
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
5465. Studying Large Language Model Generalization with Influence Functions (arxiv.org)
6 points by tosh on Apr 8, 2024 | hide | past | pdf | 1 comment
5466. Cramming: Training a Language Model on a Single GPU in One Day (2022) (arxiv.org)
2 points by tosh on Apr 8, 2024 | hide | past | pdf | discuss
5467. Formal Aspects of Language Modeling (arxiv.org)
4 points by Anon84 on Apr 8, 2024 | hide | past | pdf | discuss
5468. No "Zero-Shot" Without Exponential Data (arxiv.org)
2 points by kens on Apr 8, 2024 | hide | past | pdf | discuss
5469. Stream of Search: Learning to Search in Language (arxiv.org)
1 point by Jimmc414 on Apr 8, 2024 | hide | past | pdf | discuss
5470. H2O-Danube-1.8B Technical Report (arxiv.org)
7 points by tosh on Apr 7, 2024 | hide | past | pdf | discuss
5471. Human vs. Machine: Language Models and Wargames (arxiv.org)
1 point by CharlesW on Apr 7, 2024 | hide | past | pdf | discuss
5472. Mixture-of-Depths: Dynamically allocating compute in transformers (arxiv.org)
281 points by milliondreams on Apr 7, 2024 | hide | past | pdf | 83 comments
5473. Arizona State University – Can Large Language Models Reason and Plan? (arxiv.org)
2 points by milliondreams on Apr 7, 2024 | hide | past | pdf | discuss
5474. Sophia: Scalable Stochastic 2nd-Order Optimizer for Language Model Pre-Training (arxiv.org)
54 points by tosh on Apr 7, 2024 | hide | past | pdf | 2 comments
5475. In-Context Learning with Retrieved Demonstrations for Language Models (arxiv.org)
1 point by tosh on Apr 7, 2024 | hide | past | pdf | discuss
5476. Vulnerability Detection with Code Language Models: How Far Are We? (arxiv.org)
1 point by waffleshype on Apr 7, 2024 | hide | past | pdf | discuss
5477. Fine-Tuning LLMs with ORPO: Odds Ratio Preference Optimization (arxiv.org)
3 points by TaurenHunter on Apr 7, 2024 | hide | past | pdf | discuss
5478. Faithfulness and content selection in book-length summarization (arxiv.org)
2 points by tkgally on Apr 7, 2024 | hide | past | pdf | discuss
5479. More Agents Is All You Need: LLMs performance scales with the number of agents (arxiv.org)
288 points by TaurenHunter on Apr 6, 2024 | hide | past | pdf | 206 comments
5480. Learning to Infer Generative Template Programs for Visual Concepts (arxiv.org)
2 points by PaulHoule on Apr 6, 2024 | hide | past | pdf | discuss
5481. Using Hallucinations to Bypass RLHF Filters (arxiv.org)
9 points by TaurenHunter on Apr 6, 2024 | hide | past | pdf | discuss
5482. Long-form factuality in large language models (arxiv.org)
26 points by PaulHoule on Apr 6, 2024 | hide | past | pdf | 16 comments
5483. Language models are Super Mario: Absorbing abilities from homologous models (arxiv.org)
105 points by tosh on Apr 6, 2024 | hide | past | pdf | 70 comments
5484. Can LLMs Every Reason? (arxiv.org)
1 point by milliondreams on Apr 6, 2024 | hide | past | pdf | 1 comment
5485. Dynamically allocating compute in transformer-based language models (arxiv.org)
2 points by Anon84 on Apr 6, 2024 | hide | past | pdf | discuss
5486. ReFT: Representation Finetuning for Language Models (arxiv.org)
3 points by kmdupree on Apr 5, 2024 | hide | past | pdf | discuss
5487. H2O-Danube-1.8B: Technical Report (arxiv.org)
3 points by tosh on Apr 5, 2024 | hide | past | pdf | discuss
5488. Understanding Diffusion Models: A Unified Perspective (arxiv.org)
2 points by panabee on Apr 5, 2024 | hide | past | pdf | discuss
5489. Making Sentence Embeddings Robust to User-Generated Content (arxiv.org)
1 point by PaulHoule on Apr 5, 2024 | hide | past | pdf | discuss
5490. Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks (arxiv.org)
3 points by max-andr on Apr 5, 2024 | hide | past | pdf | 1 comment