about
131. AutoUVM: Automated Prefetching Framework for LLMs Under UVM Oversubscription (arxiv.org)
When a language model is too big for a GPU, data spills into regular memory and page swaps drag it down. AutoUVM watches which tensors the framework needs and pulls them in ahead of time, running 3.1x faster than the usual spill-and-fetch approach.
2 points by torsi0n 25 days ago | hide | past | pdf | discuss
132. Continual Harness: Online Adaptation for Self-Improving Foundation Agents (arxiv.org)
A wrapper that lets a game agent, in one run with no restarts, rewrite its own instructions, helpers, skills, and memory. On two Pokémon games it needed fewer button presses than a bare setup and closed most of the gap to a hand-built expert one.
2 points by wslh 25 days ago | hide | past | pdf | discuss
133. Hakken: Predicting future discoveries to fill gaps in today's knowledge (arxiv.org)
Hakken studies how links between scientific concepts appear over time in research papers, then guesses missing connections and explains each guess for scientists. In biomedicine it proposed 1.5 million candidate links about aging; lab tests confirmed previously unknown TP53–BAMBI and RAF1–TNF links.
2 points by anigbrowl 25 days ago | hide | past | pdf | discuss
134. Secret Collusion Among AI Agents (arxiv.org)
AI agents could secretly coordinate by hiding messages inside ordinary-looking text, so this paper defines the risk and builds tests to measure how well models can do it. Today's models are mostly weak at it, but GPT-4 showed a clear jump.
2 points by phaseonebig 26 days ago | hide | past | pdf | discuss
135. Visual General Intelligence: A White Paper (arxiv.org)
Vision researchers lay out how learning from images, video, and 3D shapes could lead to artificial general intelligence — the kind that handles any task. It maps the principles, inputs, tests, and learning setups to pursue, without settling on one definition of visual intelligence.
2 points by Anon84 26 days ago | hide | past | pdf | discuss
136. Transformers Are Efficient Compilers, Provably (arxiv.org)
They prove a transformer can compile a small C-like language—building its syntax tree, linking names, and checking types—when code nesting stays shallow, with parameters growing only as the logarithm of program length. Older word-by-word networks need linearly more, and tests confirm this gap.
2 points by Bluestein 27 days ago | hide | past | pdf | discuss
137. WeatherNext 3: Increasing resolution and performance of global weather models (arxiv.org)
This AI weather model reads raw satellite and station data instead of pre-processed analysis data, so it can forecast anywhere and update every hour instead of every six. It beats global models on temperature and dewpoint error, even at stations it never saw.
2 points by droidjj 27 days ago | hide | past | pdf | discuss
138. Legibility Is Not Interpretability (arxiv.org)
They measured each step's importance by how much it changes the chance of a right answer, then asked judges to spot the key steps from the text. Judges beat a simple guess but stayed far from perfect, so readable reasoning only partly shows what mattered.
2 points by iamsyr 27 days ago | hide | past | pdf | discuss
139. Oversight Has Capacity: Calibrating Agent Guards to a Subjective Fatiguing Human (arxiv.org)
A guard that pauses risky agent actions for human approval is tuned to the reviewer's limits, since people disagree and tire under load. On 125 actions, reviewers only moderately agreed (0.52 on an agreement scale), and safety peaked at partial escalation, which better resisted a flooding attack.
1 point by Jimmc414 5 hours ago | hide | past | pdf | discuss
140. Decoding Looped Transformers Better for Almost Free (arxiv.org)
A looped transformer reruns one block, so early passes make rougher guesses; this method steers the final guess using an earlier pass, with no training. It lifted a math score from 61.88% to 73.33%, and let the model halve loops while matching full runs.
1 point by simonpure 7 hours ago | hide | past | pdf | discuss