| 5461. |
No "Zero-Shot" Without Exponential Data (arxiv.org) |
|
1 point by passwordoops on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5462. |
Is Model Collapse Inevitable? Breaking the Curse with Real and Synthetic Data (arxiv.org) |
|
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5463. |
Strum-LLM: Attributed and Structured Contrastive Summarization (arxiv.org) |
|
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5464. |
A Neuro-Inspired Topological Sparse Training Algorithm for Large Language Models (arxiv.org) |
|
3 points by PaulHoule on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5465. |
Studying Large Language Model Generalization with Influence Functions (arxiv.org) |
|
6 points by tosh on Apr 8, 2024 | hide | past | pdf | 1 comment
|
| 5466. |
Cramming: Training a Language Model on a Single GPU in One Day (2022) (arxiv.org) |
|
2 points by tosh on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5467. |
Formal Aspects of Language Modeling (arxiv.org) |
|
4 points by Anon84 on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5468. |
No "Zero-Shot" Without Exponential Data (arxiv.org) |
|
2 points by kens on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5469. |
Stream of Search: Learning to Search in Language (arxiv.org) |
|
1 point by Jimmc414 on Apr 8, 2024 | hide | past | pdf | discuss
|
| 5470. |
H2O-Danube-1.8B Technical Report (arxiv.org) |
|
7 points by tosh on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5471. |
Human vs. Machine: Language Models and Wargames (arxiv.org) |
|
1 point by CharlesW on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5472. |
Mixture-of-Depths: Dynamically allocating compute in transformers (arxiv.org) |
|
281 points by milliondreams on Apr 7, 2024 | hide | past | pdf | 83 comments
|
| 5473. |
Arizona State University – Can Large Language Models Reason and Plan? (arxiv.org) |
|
2 points by milliondreams on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5474. |
Sophia: Scalable Stochastic 2nd-Order Optimizer for Language Model Pre-Training (arxiv.org) |
|
54 points by tosh on Apr 7, 2024 | hide | past | pdf | 2 comments
|
| 5475. |
In-Context Learning with Retrieved Demonstrations for Language Models (arxiv.org) |
|
1 point by tosh on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5476. |
Vulnerability Detection with Code Language Models: How Far Are We? (arxiv.org) |
|
1 point by waffleshype on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5477. |
Fine-Tuning LLMs with ORPO: Odds Ratio Preference Optimization (arxiv.org) |
|
3 points by TaurenHunter on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5478. |
Faithfulness and content selection in book-length summarization (arxiv.org) |
|
2 points by tkgally on Apr 7, 2024 | hide | past | pdf | discuss
|
| 5479. |
More Agents Is All You Need: LLMs performance scales with the number of agents (arxiv.org) |
|
288 points by TaurenHunter on Apr 6, 2024 | hide | past | pdf | 206 comments
|
| 5480. |
Learning to Infer Generative Template Programs for Visual Concepts (arxiv.org) |
|
2 points by PaulHoule on Apr 6, 2024 | hide | past | pdf | discuss
|
| 5481. |
Using Hallucinations to Bypass RLHF Filters (arxiv.org) |
|
9 points by TaurenHunter on Apr 6, 2024 | hide | past | pdf | discuss
|
| 5482. |
Long-form factuality in large language models (arxiv.org) |
|
26 points by PaulHoule on Apr 6, 2024 | hide | past | pdf | 16 comments
|
| 5483. |
Language models are Super Mario: Absorbing abilities from homologous models (arxiv.org) |
|
105 points by tosh on Apr 6, 2024 | hide | past | pdf | 70 comments
|
| 5484. |
Can LLMs Every Reason? (arxiv.org) |
|
1 point by milliondreams on Apr 6, 2024 | hide | past | pdf | 1 comment
|
| 5485. |
Dynamically allocating compute in transformer-based language models (arxiv.org) |
|
2 points by Anon84 on Apr 6, 2024 | hide | past | pdf | discuss
|
| 5486. |
ReFT: Representation Finetuning for Language Models (arxiv.org) |
|
3 points by kmdupree on Apr 5, 2024 | hide | past | pdf | discuss
|
| 5487. |
H2O-Danube-1.8B: Technical Report (arxiv.org) |
|
3 points by tosh on Apr 5, 2024 | hide | past | pdf | discuss
|
| 5488. |
Understanding Diffusion Models: A Unified Perspective (arxiv.org) |
|
2 points by panabee on Apr 5, 2024 | hide | past | pdf | discuss
|
| 5489. |
Making Sentence Embeddings Robust to User-Generated Content (arxiv.org) |
|
1 point by PaulHoule on Apr 5, 2024 | hide | past | pdf | discuss
|
| 5490. |
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks (arxiv.org) |
|
3 points by max-andr on Apr 5, 2024 | hide | past | pdf | 1 comment
|
| More |