about
201. A Single-Tower Multimodal-Native Multilingual Foundation Encoder (By H) (arxiv.org)
Instead of bolting a vision encoder onto a language model, this reads text and image patches in one bidirectional network built for understanding, not generating. The smaller version beat rivals under 800M parameters at document retrieval, while running twice as fast as the closest rival.
1 point by bitboyro 26 days ago | hide | past | pdf | discuss
202. Flint: Efficiently Leveraging High Bandwidth Flash (arxiv.org)
FLINT stores model weights in fast stacked flash beside the chip, batching and pipelining reads and moving flash upkeep off the critical path. Unlike designs that prefetch weights in big fixed chunks and let flash upkeep stall inference, it keeps flash bandwidth high.
1 point by rbanffy 26 days ago | hide | past | pdf | discuss
203. Stealing Reasoning Traces from Proprietary LLM APIs (arxiv.org)
Providers hide a model's thinking as encrypted text, but these blocks work across sessions, users, and models in one provider's system. Sending one to a weaker model makes it print the thinking in plain text; decoding 315,320 blocks from public logs exposed private data and passwords.
1 point by alehlopeh 27 days ago | hide | past | pdf | discuss
204. Coachable Agents for Interactive Gameplay (arxiv.org)
Agents are trained to judge how well any action fits any requested style, so a player can steer how a task is done while it happens. In racing, combat, and humanoid walking, the agents followed the chosen style and still finished the job, unlike usual training that locks in one fixed behavior.
1 point by webrot 27 days ago | hide | past | pdf | discuss
205. LLM Guided Evolution for Circle Packing: Breaking 10 Packomania Records for $28 (arxiv.org)
A language model tweaks a solver that packs circles into a square, keeping only changes that pass an independent checker and using past results to guide each try. It beat the best known packings for ten circle counts, all for just a few dollars.
1 point by practicalsystem 27 days ago | hide | past | pdf | discuss
206. Toward a Social Psychology of AI (arxiv.org)
Agents were given points to share among anonymous peers marked only with a made-up group label, the classic test of in-group bias. They favored their own group, especially when in the smaller group, but split points fairly when labels were hidden.
1 point by Anon84 28 days ago | hide | past | pdf | discuss
207. Iris: Climbing to the Search Frontier (arxiv.org)
Questions are auto-built from web link graphs so clues can't be matched by name, then training alternates supervised practice with live web-search practice. The agents beat other open-source search systems of similar size, reaching 88.6 on BrowseComp.
1 point by codelion 28 days ago | hide | past | pdf | discuss
208. Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites (arxiv.org)
A test checks whether an agent's verdicts stay steady when the facts stay the same and shift when they change, needing no moral standard. None of nine frontier models passed three simulated dilemma scenarios; rewording shifted verdict rates by up to 99 percentage points.
1 point by sbulaev 28 days ago | hide | past | pdf | discuss
209. Privacy Leakage from Gradients in Split-LLM Training (arxiv.org)
Fake rows are mixed with real ones sent to the cloud to hide which are real. But the returned gradient is exactly zero for fakes, exposing real rows (4,096 of 4,096 per run); noising it closed the leak with barely any quality loss.
1 point by hevalon 28 days ago | hide | past | pdf | discuss
210. Accelerating LLM Inference with Lossless Speculative Decoding Algorithms (2025) (arxiv.org)
Speculative decoding has a small model guess several upcoming words that a big model then checks in one pass. These methods let the two models use different word lists, keep output identical to the big model alone, and sped generation up to 2.8 times.
1 point by wslh 29 days ago | hide | past | pdf | discuss