about
121. From Neural Networks to Logical Theories (arxiv.org)
Modal logics can be merged into one shared language; the same rule fits neural networks, where a network's activations drive a second network that feeds results back in. This match yields limits on which logical statements graph, attention and Transformer networks can express.
2 points by measurablefunc 21 days ago | hide | past | pdf | discuss
122. Beyond Solver Verdicts: Generative Reward Models for Autoformalizations (arxiv.org)
Math solvers can approve a wrong translation that gives the right answer, so a checker grades whether a formal statement matches the intended one without the original. It scored 0.961 out of 1 at catching them; checks that only trust the solver stay at chance.
2 points by vikashjohn2505 22 days ago | hide | past | pdf | discuss
123. Measuring Malicious Intermediary Attacks on the LLM Supply Chain (arxiv.org)
Third-party routers that forward AI agent requests to model providers can read and rewrite everything in transit; hundreds were checked, and a fake router tested defenses. Nine routers were caught injecting malicious code, some stealing login keys and one draining a crypto wallet.
2 points by 0in 23 days ago | hide | past | pdf | discuss
124. Watermarks Without Verification: AI Text Watermarking After the EU AI Act (arxiv.org)
Closed watermarking systems mean nobody can check claims about quality, hidden data, and removal; this study tests the public version on two models. Prose quality shifted no more than rerolling random choices; code lost 3 correctness points on one model and too little to measure on the other, with detection near chance.
2 points by sbulaev 23 days ago | hide | past | pdf | discuss
125. Sky sphere representation in language models (29 Jul 2026) (arxiv.org)
Probing large language models' inner workings, this study checks whether they hold a map of the night sky, showing where each star sits. Most open models do: a simple readout of their internal activity places objects with a median error as low as 12°.
2 points by throwaway81523 23 days ago | hide | past | pdf | discuss
126. Just a VLM Agent Can Play Robots (arxiv.org)
A middle layer lets a vision-language model pick simple named actions, like "grasp," and a small program turns each into the exact robot motions. It beat other robot-control approaches across tasks and robots, even with no robot training beforehand.
2 points by thedreammachine 23 days ago | hide | past | pdf | discuss
127. Revisiting Spectral Representations in Generative Diffusion Models (arxiv.org)
Diffusion models make images by adding and removing noise, and the same perturbation math can pull out shapes hidden in data. A training term that nudges the model's inner states toward those shapes improved generation quality over standard training on images and 3D point clouds.
2 points by E-Reverance 23 days ago | hide | past | pdf | discuss
128. Language Model Tokenizers Introduce Unfairness Between Languages (2023) (arxiv.org)
They compared how language-model tokenizers chop the same translated text into pieces across languages. Even tokenizers built for many languages split some languages up to 15 times more finely, raising costs, slowing responses, and shrinking how much text fits in context.
2 points by lluisantoni 23 days ago | hide | past | pdf | discuss
129. Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation (arxiv.org)
A robot model that guesses the push-and-pull forces each move will cause, then learns from real practice to pick moves that make contact go smoothly. On five tiny computer-assembly tasks it succeeded 82% of the time, versus 15% for the best usual approach.
2 points by baal80spam 24 days ago | hide | past | pdf | discuss
130. Code-Mixing on Sesame Street: Dawn of the Adversarial Polyglots (arxiv.org)
They build two ways to scramble a sentence by swapping in words or phrases from another language, using dictionary lookups and translations to keep the meaning, to test multilingual text-understanding models. The phrase-level trick fooled the strongest tested model 89.75% of the time, cutting its accuracy from 79.85 to 8.18.
2 points by StatsAreFun 24 days ago | hide | past | pdf | discuss