about
61. Can Theoretical Physics Research Benefit from Language Agents? (arxiv.org)
Today's AI chat models handle math and code but stumble on physical intuition and hard constraints, which better prompts can't fix. The fix proposed is agents trained on physics reasoning and checked by tools that enforce physical laws, needing new training data and reward signals.
3 points by num42 18 days ago | hide | past | pdf | discuss
62. Co-Evolving Harnesses and Models (arxiv.org)
They tune the prompt-and-tools wrapper around a small model, then instead of copying an expert's whole solution, they rewrite only the step where the model's attempt fails. Copying full runs hurt it on all seven tasks by 4 to 30 points; this fix kept the gains.
3 points by gmays 19 days ago | hide | past | pdf | discuss
63. Can AI agents conduct open-ended AI research? (arxiv.org)
Give an AI agent the main open-ended question from a strong unpublished paper, then let the paper's own authors judge its answer. In two six-day trials the agents handled all the coding alone but made no real progress, so both answers were rejected.
3 points by Betelbuddy 19 days ago | hide | past | pdf | discuss
64. The Measure of Intelligence (2019) (arxiv.org)
Judging a system by how well it plays one game can be faked with endless practice, so intelligence is how efficiently it picks up new skills. A new puzzle test uses only the basic knowledge humans are born with to compare people and AI fairly.
4 points by theanonymousone 29 days ago | hide | past | pdf | discuss
65. The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement (arxiv.org)
A roadmap for AI that improves itself and the way it improves, moving through stages from carrying out upgrades to choosing its own strategies and adapting to new settings. A new index shows today's language models still fall far short of this goal.
3 points by pella 20 days ago | hide | past | pdf | discuss
66. Large Language Models Reflect the Ideology of Their Creators (arxiv.org)
The study asked popular AI chatbots to describe political figures in all six UN languages, then read the moral judgments in their answers. Models from different countries and languages showed different values, matching their creators' worldviews rather than being neutral.
3 points by zvr 22 days ago | hide | past | pdf | discuss
67. An Overview of Large Language Models for Statisticians (arxiv.org)
This overview maps where statistics can help make large language models more trustworthy and transparent, from measuring uncertainty to protecting privacy and adapting models to new data. It asks how these text tools could aid statistical analysis, arguing the two fields should work together.
3 points by Anon84 23 days ago | hide | past | pdf | discuss
68. AI Coding Assistants Almost Never Check Supply-Chain Trust Signals (arxiv.org)
They tested whether AI coding assistants check trust signals—signed releases, build records, channel lists—before installing software, scoring behavior from container logs, not its words. Verification was nearly absent: 9 of 1,920 trials opened any signal, none ran a verification command, so signals changed nothing.
3 points by sbulaev 23 days ago | hide | past | pdf | discuss
69. Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning (arxiv.org)
An agent must figure out a hidden rule machine by asking whether strings are accepted and guessing the whole machine, like reverse-engineering a program. Language models solved small machines but failed as these grew, though they knew the standard algorithm and could code it.
3 points by wslh 24 days ago | hide | past | pdf | discuss
70. Kalman Delta Networks (arxiv.org)
A long-context model keeps a fixed-size memory of past words and tracks how sure it is about each stored association, so new information counts more when the memory is shaky. At two sizes, it beat top similar models on word-prediction error and task accuracy.
3 points by E-Reverance 24 days ago | hide | past | pdf | discuss