| 111. |
Continual Learning Mechanisms Compose for Long-Horizon Memorization (arxiv.org) |
| They tested ways to stop a model forgetting old question-answer pairs as it learns 100 new ones in a row, then combined the best tricks that protect different things. The mix remembered 34.9% of answers at the end, versus 1.2% for plain sequential training. |
|
2 points by donk8r 18 days ago | hide | past | pdf | discuss
|
| 112. |
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs (arxiv.org) |
| A weaker model chops a harmful task into harmless-looking pieces, asks a stronger, safety-trained model about each piece, then stitches the answers together. This trick, called capability laundering, gets around refusals, lifting the weak model's bioweapon score from 62.3 to 83.1 out of 100. |
|
2 points by sbulaev 18 days ago | hide | past | pdf | discuss
|
| 113. |
AdaBoost Does Not Always Cycle (arxiv.org) |
| A 2012 question asked whether AdaBoost's exhaustive version, which keeps retraining on the hardest cases, always settles into a repeating cycle. A two-part example whose growth rates have an irrational ratio shows it does not: the winners never settle into a cycle, all checked exactly. |
|
2 points by gone35 18 days ago | hide | past | pdf | discuss
|
| 114. |
StepAudio 3 Gen Technical Report (arxiv.org) |
| One model turns text into speech, music, or sound effects by predicting compressed audio codes one after another, instead of diffusion models that paint continuous audio. It beat the best on text-to-speech and voice design while still handling speech, vocals, sound effects, and music. |
|
2 points by gmays 18 days ago | hide | past | pdf | discuss
|
| 115. |
Breaking the Token Ceiling (arxiv.org) |
| Big token-based models can teach small models that read text one byte at a time by turning the big model's word guesses into byte guesses. As training data grew, byte students started weaker but passed token students, matching them with one-sixth the data. |
|
2 points by sonabinu 18 days ago | hide | past | pdf | discuss
|
| 116. |
NCP-ArchPreview: 8.9B latent LM matches OLMo-3-7B on 51% of tokens (arxiv.org) |
| Instead of guessing only the next word, this model also predicts concepts—multi-token ideas from its internal states—to guide what it writes next. It matched a similar-size model's pretraining loss using half the tokens, then beat it by 2.45 points on average. |
|
2 points by iamsyr 19 days ago | hide | past | pdf | discuss
|
| 117. |
Solve the Loop: Attractor Models for Language and Reasoning (arxiv.org) |
| A backbone proposes an answer, then a second module refines it until it stops changing, keeping memory flat and letting the model choose its own number of rounds. It beat a 1.3B-parameter standard language model trained on twice the data despite having only 770M parameters. |
|
2 points by jerlendds 20 days ago | hide | past | pdf | 1 comment
|
| 118. |
MOBA: Mixture of Block Attention for Long-Context LLMs (arxiv.org) |
| Instead of letting every word check every other word, this splits long text into blocks and lets each word pick which blocks matter. It matches the accuracy of checking everything while cutting the cost, and now powers Kimi's long-context requests. |
|
2 points by ur-whale 20 days ago | hide | past | pdf | discuss
|
| 119. |
Datasets for Large Language Models: A Comprehensive Survey (arxiv.org) |
| This survey maps the text datasets that feed large language models, sorting them by role: raw pre-training text, instruction-tuning examples, preference pairs, evaluation sets, and language-processing collections. It catalogs 444 datasets across many languages and domains, with details on size, quality, and sources. |
|
2 points by Anon84 21 days ago | hide | past | pdf | discuss
|
| 120. |
Domain-Specific Hallucination Detection in Large Language Models (arxiv.org) |
| A detector combines a fine-tuned text classifier with repeated-run uncertainty checks and calibrated confidence to flag unfaithful claims in AI answers. It scores 0.915 on general tasks but just 0.52 on biomedical text, showing detectors need retraining for each domain. |
|
2 points by Betelbuddy 21 days ago | hide | past | pdf | 1 comment
|
| More |