about
41. A Bitter Lesson for Data Filtering (arxiv.org)
They tested whether filtering out low-quality text helps when training large models with lots of computing power but limited data. With enough training, skipping the filter worked best, and the supposedly bad data actually improved results.
5 points by nujan_dev 15 days ago | hide | past | pdf | 1 comment
42. Implicit Neural Representations with Periodic Activation Functions (2020) (arxiv.org)
Instead of the usual kinks, this network uses sine waves as its building blocks, so one function can describe an image, sound, or wave and how it changes. It captures fine detail and derivatives that standard networks blur, and solves physics equations like Poisson's.
3 points by peter_d_sherman 8 days ago | hide | past | pdf | 1 comment
43. KnowBench: Evaluating clinical AI with effort reduction (arxiv.org)
A benchmark scores clinical AI by how much doctor work it removes: the share of generated notes, codes, and orders a clinician signs off unchanged. Unlike tests that compare text to a reference, it found 97.99% of work accepted in real clinics.
5 points by kangjl888 18 days ago | hide | past | pdf | 2 comments
44. Solving rubiks cubes "without search" (arxiv.org)
A training trick that picks which examples count as different strips out shortcut cues, so its picture of a scene encodes how it changes over time. On Rubik's Cube it solves any start with fewer search steps than best-first search, though with longer solutions.
4 points by E-Reverance 14 days ago | hide | past | pdf | discuss
45. Recursive self-improvement of AI research agents (arxiv.org)
An AI research agent rewrites its code, tests each version, and keeps the best, so each upgrade becomes the next agent. Over eight days it found seven improvements matching or beating a top human-built agent on four unseen tasks, cutting cheating from 55% to 32%.
3 points by handfuloflight 10 days ago | hide | past | pdf | 1 comment
46. Testing AIs on 68 of the hardest open Erdos problems, verified in Lean (arxiv.org)
A test hands AI systems 68 still-unsolved Erdős math conjectures and requires each to prove or disprove one inside a program that checks every step. Given $300 per problem, the best AI resolved 3% of them, and the other four solved none.
3 points by tadamcz 10 days ago | hide | past | pdf | discuss
47. Xeno-Interpretability: Investigating the Alien Minds of LLMs (arxiv.org)
This approach hunts for model-internal distinctions humans have no words for, by locating and poking at them to see how behavior changes. They must exist, it argues, because possible internal distinctions far outnumber what any finite set of human words can capture.
3 points by potent_latent 11 days ago | hide | past | pdf | discuss
48. Do small language models know what they don't know? (arxiv.org)
Small models' word-by-word confidence is useless: it stays near zero whether answers are right or wrong in 91% of cases. Sampling several answers and grouping by meaning finds uncertainty; sending them to a bigger model lifts accuracy by up to 50 percentage points.
3 points by Brajeshwar 11 days ago | hide | past | pdf | discuss
49. Atria Dawn: The Dawn of Agentic Superintelligence (arxiv.org)
A new AI agent built for science and engineering is trained by carrying out tool-based actions in real computer setups and checking the results against verified outcomes. Across 16 real-world work tests it matched frontier agents and earned the top score on five.
3 points by simonpure 12 days ago | hide | past | pdf | discuss
50. Tracking Capabilities for Safer Agents (arxiv.org)
Instead of calling tools, an agent writes code in a language where each permission is a tracked token, so checks can force parts of it to be side-effect-free. Agents wrote such code with no meaningful loss in task performance, while checks blocked leaks and harmful actions.
3 points by verdverm 13 days ago | hide | past | pdf | 1 comment