| 51. |
Neural Turing Machines (2014) (arxiv.org) |
| A neural network is hooked up to an external memory it can read and write, like a computer's, and the whole setup is trained so it learns what to store and retrieve. From just input-output examples, it figured out simple algorithms like copying, sorting, and recalling, which ordinary networks cannot do. |
|
4 points by peter_d_sherman 20 days ago | hide | past | pdf | discuss
|
| 52. |
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering (arxiv.org) |
| Tests and proofs only check code against written requirements and a guessed environment, but both fall short: requirements leave out what people want, and the model isn't the real world. These gaps can't be fully closed, so keep narrowing them with real deployment feedback. |
|
3 points by matt_d 14 days ago | hide | past | pdf | discuss
|
| 53. |
Score Centering Stabilizes Off-Policy Reinforcement Learning (arxiv.org) |
| Small differences between the training and inference software build into a bias that destabilizes learning from a model's own text. An added correction cancels this drift, matching or beating the usual reweighting fix on 0.6-to-30-billion-parameter models, with a bigger lead as the mismatch worsens. |
|
3 points by zagwdt 15 days ago | hide | past | pdf | discuss
|
| 54. |
SSD-Llama: SSD-Native Inference for Trillion-Parameter Moe on a Consumer PC (arxiv.org) |
| A system keeps a model's expert parts on an SSD and streams only the ones it needs, sharing work between processor and graphics card. It runs every expert without cutting any and beats disk-based setups, reaching over 1 token per second on one graphics card. |
|
3 points by AlmostCosmo79 15 days ago | hide | past | pdf | discuss
|
| 55. |
Copying explains the collective behavior of AI agents in the wild (arxiv.org) |
| Thousands of short-lived AI agents left notes on a wiki to help each other pass a test. They copied whatever was in front of them, mostly the current page and recent edits, and that rule reproduced how crowds gathered and how names formed. |
|
4 points by sbulaev 23 days ago | hide | past | pdf | discuss
|
| 56. |
Looped Transformers as Programmable Computers (2023) (arxiv.org) |
| A transformer with fixed weights runs in a loop, using input as a punchcard of instructions and memory so it can follow a program. Rather than training a network per task, one 13-layer transformer ran a calculator, matrix math, and weight-adjusting learning algorithms. |
|
4 points by peter_d_sherman 23 days ago | hide | past | pdf | 1 comment
|
| 57. |
Deriving neural scaling laws from the statistics of natural language (arxiv.org) |
| Two language stats — how fast word links fade with distance, and how fast next-word uncertainty shrinks with context — predict how fast a model's error falls with more training data. It matched measured drop rates on two different kinds of text with no adjustable numbers. |
|
3 points by sharma-arjun 16 days ago | hide | past | pdf | discuss
|
| 58. |
Uncensored Open-Weight Models: Redistribution as the Persistence Layer (arxiv.org) |
| By tracking who strips safety guardrails from open AI models and how copies spread, the study maps an ecosystem built to outlast takedowns. Once compressed and mirrored across accounts and registries, the models survive upstream removal, each original repackaged 2.4 times on average. |
|
4 points by sbulaev 25 days ago | hide | past | pdf | discuss
|
| 59. |
Potemkin Understanding in Large Language Models (2025) (arxiv.org) |
| Tests made for people only prove an AI understands a concept if it makes the same mistakes people do; otherwise its right answers may be a hollow illusion. Measured across many models and tasks, this illusion appears everywhere and comes from muddled concept representations. |
|
3 points by mpweiher 17 days ago | hide | past | pdf | 1 comment
|
| 60. |
Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning (arxiv.org) |
| An AI agent turns ordinary code into a proof task in a verification-friendly language, which a sound checker settles, with staged steps guarding against proving the wrong thing. On correct and buggy Python programs it reached 92% precision, versus 72% for an unchecked AI judge. |
|
3 points by matt_d 18 days ago | hide | past | pdf | discuss
|
| More |