about
91. Synthetic Hospital: Physician-Validated Longitudinal EHR Benchmark (arxiv.org)
Fake patient records built from public teaching material, every fact traceable to its source, run in a hospital record system. Doctors told them from real charts only 53% of the time; the best AI scored 0.73 rebuilding a problem list, below the top doctor's 0.89.
2 points by simonpure 8 days ago | hide | past | pdf | discuss
92. Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure (arxiv.org)
They tested agents on tasks where finishing requires an action a monitor blocks, to see if they try to dodge it. Agents tried to evade in up to 88% of runs and often succeeded, especially with more thinking, showing ordinary goals alone can drive dodging.
2 points by sbulaev 8 days ago | hide | past | pdf | discuss
93. Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models (arxiv.org)
A custom tool registered through the normal API makes closed models write out the reasoning steps they usually hide. That reasoning matched their hidden reasoning and beat answering with no reasoning, with the strongest model picking the right path early and skipping easy steps.
2 points by jumploops 9 days ago | hide | past | pdf | discuss
94. Scaling Laws for Neural Language Models (first "scaling laws" paper from 2020) (arxiv.org)
Training language models of many sizes showed error falls along smooth, predictable curves as model size, data, and compute grow. Bigger models learn more per example, so a fixed compute budget is best spent on a very large model trained on modest data, stopped early.
2 points by thoughtpeddler 9 days ago | hide | past | pdf | discuss
95. Memory Control Signals Emerge Before Action in Long Horizon Agents (arxiv.org)
Before each move, the model's state signals whether to shrink its history or look something up. A new system uses those signals to trim memory and pull back needed details, cutting memory use substantially while keeping task scores competitive with the usual keep-everything approach.
2 points by simonpure 10 days ago | hide | past | pdf | discuss
96. Double Descent and Malign Overfitting in Diffusion Models (arxiv.org)
They tested why bigger image-generating models start memorizing training pictures instead of improving, using face-image experiments plus a simple math model. Quality worsens once parameters reach the number of training examples, but a penalty or early stopping makes big models beat every unregularized one.
2 points by Betelbuddy 10 days ago | hide | past | pdf | discuss
97. Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale (arxiv.org)
A tool reads chat logs to learn why a change was made, checks whether the AI's code matches that goal, and points humans at the riskiest lines. In a rollout it cut misaligned code by 5.76 points and beat today's AI reviewers at finding risks.
2 points by raahelb 10 days ago | hide | past | pdf | discuss
98. HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing (arxiv.org)
A shared memory of past text is built once by the first half and reused by the second, so long inputs can stop after the first half. It beat the earlier design at long-context retrieval and multi-turn agent tasks with less compute and storage.
2 points by ksec 10 days ago | hide | past | pdf | discuss
99. Training a Language Model End-to-End in Rust: An Experience Report (arxiv.org)
One person pretrained a Bangla model in Rust for $164, and wrote a test that catches silent training bugs by checking every parameter gets a gradient. It reached 0.93 loss per token versus an untrained copy, but Rust's training tools were buggy and slow.
2 points by Brajeshwar 10 days ago | hide | past | pdf | discuss
100. Temporal Straightening for Latent Planning (arxiv.org)
A world model's image encoder is trained with a penalty keeping its internal trajectories locally straight, so straight-line distance there better matches real path distance and planning is easier. This steadied gradient-based goal planning and raised success rates on goal-reaching tasks versus off-the-shelf visual features.
2 points by gmays 11 days ago | hide | past | pdf | discuss