| 101. |
Complex KDA:Understanding and Enhancing the Expressivity of Kimi Delta Attention (arxiv.org) |
| Letting the gate and update weight go negative lets one KDA step perform a 2D rotation instead of needing two stacked steps. This matches the two-step version's ability to track states at single-step cost, and gave the best length extrapolation among the settings tried. |
|
2 points by jul8234 12 days ago | hide | past | pdf | discuss
|
| 102. |
Scaling Discovery Through Test-Time Communication (arxiv.org) |
| Agents working on hard puzzles share their progress in a common folder, so one breakthrough can lift the whole team. A team of k matched 4k working alone, and beat the best-known human answers on packing and compression tasks. |
|
2 points by simonpure 12 days ago | hide | past | pdf | discuss
|
| 103. |
An Introduction to Compression-Based Machine Learning (arxiv.org) |
| Any lossless compressor like gzip can become a classifier by comparing how well it squeezes data or picking the label that compresses best. A survey finds these designs match usual baselines, are far stronger on malware, and vary in accuracy by up to 0.62. |
|
2 points by jackhurwitz 12 days ago | hide | past | pdf | discuss
|
| 104. |
Steerable Cultural Preference Optimization of Reward Models (arxiv.org) |
| Reward models score AI answers by how much people like them; a weighting scheme keeps one country's tastes from drowning out the rest. It beat usual training by up to 7 points for the minority group and cut bias toward bigger groups. |
|
2 points by measurablefunc 12 days ago | hide | past | pdf | discuss
|
| 105. |
A Black-Box Audit of Provider-Side Token Inflation in LLM Services (arxiv.org) |
| A dishonest pay-per-token service can quietly stretch answers so users pay for extra words. A simple test nudges it to write longer: if it is already inflating, output barely grows, catching 85.1% of attacks with almost no false alarms. |
|
2 points by donk8r 15 days ago | hide | past | pdf | discuss
|
| 106. |
Federated Learning Is Not Private for Google GBoard Next Word Prediction [pdf] (arxiv.org) |
| Federated learning trains a phone keyboard's next-word model without sending typed text, but attackers can compare two model versions to read the words back out. Typed words and their order were recovered with high accuracy even with small batches or added random noise. |
|
2 points by thunderbong 15 days ago | hide | past | pdf | discuss
|
| 107. |
Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents (arxiv.org) |
| ScienceBuddy turns scientists' requests and feedback into tasks and grading rules for its AI. Two loops feed each other: the tool around a fixed model is improved first, then the model is trained inside it; case studies across four task kinds showed it working. |
|
2 points by Betelbuddy 16 days ago | hide | past | pdf | discuss
|
| 108. |
200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management (arxiv.org) |
| JustFit is a laptop runtime that squeezes the model's short-term memory to four bits and loads only the pieces each step needs, so execution and chat history share limited RAM. It kept 327,680 text positions alive versus the best tested setup's 30,720. |
|
2 points by Betelbuddy 16 days ago | hide | past | pdf | discuss
|
| 109. |
Competitive Market Behavior of LLMs (arxiv.org) |
| Classic economics experiments were rerun with AI chatbots playing buyers and sellers in a double auction, where traders call out prices until deals are made. These markets settled on fair prices more slowly or not at all, wasting more value than human-run markets. |
|
2 points by paulpauper 17 days ago | hide | past | pdf | discuss
|
| 110. |
Coding Agents Have Converged: Why the SWE-Bench Leaderboard Can No Longer Order (arxiv.org) |
| They checked 254 leaderboard submissions without rerunning anything, comparing which problems each solved to see if small gaps rank them. The leaders solve nearly the same problems, and adjacent top entries can't be separated; changing the surrounding code shifted scores by up to 29.8 points. |
|
2 points by sbulaev 17 days ago | hide | past | pdf | discuss
|
| More |