about
271. Reward Guided Speculative Decoding (arxiv.org)
1 point by E-Reverance 34 days ago | hide | past | pdf | discuss
272. Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills (arxiv.org)
5 points by swolpers 34 days ago | hide | past | pdf | discuss
273. Scaling Domain Data Repetition in LLM Pretraining (arxiv.org)
1 point by gmays 34 days ago | hide | past | pdf | discuss
274. DumpsterCluster: From Dumpster Diving to Serving Llama-70B on $60 GPUs (arxiv.org)
2 points by bookofjoe 34 days ago | hide | past | pdf | 1 comment
275. FreeToken: Efficient Edge-Native Moe Serving with Bandwidth-Adaptive Execution (arxiv.org)
2 points by gmays 34 days ago | hide | past | pdf | discuss
276. Deep Learning Is Not So Mysterious or Different (arxiv.org)
2 points by Anon84 34 days ago | hide | past | pdf | discuss
277. PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical (arxiv.org)
2 points by sbulaev 35 days ago | hide | past | pdf | discuss
278. Full-Bandwidth Transformer (arxiv.org)
1 point by gmays 35 days ago | hide | past | pdf | discuss
279. Eight Things to Know about Large Language Models(2023) (arxiv.org)
2 points by turingbook 35 days ago | hide | past | pdf | 1 comment
280. SwarmWorld: Stigmergic technological evolution in societies of LLM agents (arxiv.org)
3 points by wintercarver 35 days ago | hide | past | pdf | 1 comment
281. 30% on AR-AGI-1 at $0.0007 per task (arxiv.org)
3 points by randomizedalgs 35 days ago | hide | past | pdf | discuss
282. WikiSkill: Compiling Agent Experience for Skill Evolution (arxiv.org)
1 point by joshcorbin 35 days ago | hide | past | pdf | discuss
283. Beyond the Editing Canvas: Evidence Divergence in Ooxml-to-LLM Ingestion (arxiv.org)
3 points by sbulaev 36 days ago | hide | past | pdf | discuss
284. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents (arxiv.org)
3 points by sbulaev 36 days ago | hide | past | pdf | discuss
285. Why not to use the Gaussian kernel (arxiv.org)
2 points by E-Reverance 36 days ago | hide | past | pdf | discuss
286. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment (arxiv.org)
122 points by stephenchung 36 days ago | hide | past | pdf | 40 comments
287. Compiling Agent Experience into Persistent Knowledge for Skill Evolution (arxiv.org)
7 points by tcp_handshaker 36 days ago | hide | past | pdf | discuss
288. LLMs Can Design Near-Optimal OR Algorithms (arxiv.org)
4 points by tcp_handshaker 36 days ago | hide | past | pdf | discuss
289. AutoSaddler: Automatic Harness Optimization (arxiv.org)
23 points by drseu55 36 days ago | hide | past | pdf | 1 comment
290. Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect (arxiv.org)
2 points by sbulaev 36 days ago | hide | past | pdf | discuss
291. Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation (arxiv.org)
1 point by sbulaev 37 days ago | hide | past | pdf | discuss
292. GenRec: An LLM-Backed Recommendation Ranker at Netflix (arxiv.org)
2 points by armcat 37 days ago | hide | past | pdf | discuss
293. WAFT: Warping-Alone Field Transforms for Optical Flow (2025) (arxiv.org)
2 points by peter_d_sherman 37 days ago | hide | past | pdf | 1 comment
294. Evomal: Self-Poisoning in Self-Evolving Coding Agents (arxiv.org)
1 point by sbulaev 37 days ago | hide | past | pdf | discuss
295. Flare: Verifying MILP Reformulations with LLM-Based Theorem Proving (arxiv.org)
2 points by henryrobbins00 37 days ago | hide | past | pdf | discuss
296. Foundation Model 3x better at predicting cancer treatment (arxiv.org)
3 points by shcheklein 37 days ago | hide | past | pdf | discuss
297. Flint: Efficiently Leveraging High Bandwidth Flash for LLM Inference (arxiv.org)
2 points by metrofun 37 days ago | hide | past | pdf | discuss
298. Noise-Driven Escape from Metastable Phases Explains Grokking in DNNs (arxiv.org)
2 points by jerlendds 38 days ago | hide | past | pdf | 4 comments
299. Metaⁿ: Recursive Self-Improvement Through Emergent Depth (arxiv.org)
1 point by sbulaev 38 days ago | hide | past | pdf | discuss
300. Prefix Sliding for efficient test-time scaling (arxiv.org)
2 points by E-Reverance 38 days ago | hide | past | pdf | discuss