about
31. Pluralis Towards a Multicultural Multimodal, Multilingual Benchmark for AI Risk (arxiv.org)
1 point by thinkevolve 4 days ago | hide | past | pdf | discuss
32. AmpleGCG: Learning a Universal Generative Model for Jailbreaking (arxiv.org)
1 point by Anon84 4 days ago | hide | past | pdf | discuss
33. Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs (arxiv.org)
2 points by matt_d 4 days ago | hide | past | pdf | discuss
34. Learning How to Forget: Fine-Tuning for Long-Context Sparse Attention (arxiv.org)
1 point by theanonymousone 4 days ago | hide | past | pdf | discuss
35. Fathom: Per-query read depth for sparse decoding over offloaded KV caches (arxiv.org)
1 point by vivekkalyanaran 4 days ago | hide | past | pdf | discuss
36. Netflix replaced their recommendation algorithm with an LLM (arxiv.org)
2 points by nreece 4 days ago | hide | past | pdf | 1 comment
37. Self-Play Pretraining with Zero Data (arxiv.org)
3 points by 7777777phil 4 days ago | hide | past | pdf | discuss
38. Shockingly Simple Self-retrospection Improves Agentic Models Without RL (arxiv.org)
2 points by Betelbuddy 4 days ago | hide | past | pdf | discuss
39. Language Models Act on Hidden Valence (arxiv.org)
2 points by sbulaev 4 days ago | hide | past | pdf | discuss
40. User Model Extraction via Belief Self-Distillation (arxiv.org)
1 point by sbulaev 4 days ago | hide | past | pdf | discuss
41. CORVUS: Context Optimization and Reduction via Underlying Synchronization (arxiv.org)
2 points by matt_d 5 days ago | hide | past | pdf | discuss
42. An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries (arxiv.org)
2 points by matt_d 5 days ago | hide | past | pdf | discuss
43. GenRec: An LLM-Backed Recommendation Ranker at NetflixConference (arxiv.org)
3 points by throwthrowrow 5 days ago | hide | past | pdf | discuss
44. EnigmaForge – an LLM benchmark where the question is hidden in the story (arxiv.org)
1 point by robottwo 5 days ago | hide | past | pdf | discuss
45. Recursive Self-Improvement via On-Policy Distillation for Reasoning (arxiv.org)
2 points by simonpure 5 days ago | hide | past | pdf | discuss
46. A-Mem: Agentic Memory for LLM Agents (arxiv.org)
2 points by jerlendds 5 days ago | hide | past | pdf | discuss
47. CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents (arxiv.org)
1 point by 6bitquant 5 days ago | hide | past | pdf | discuss
48. Foundations of Large Language Models (arxiv.org)
3 points by rramadass 5 days ago | hide | past | pdf | 1 comment
49. Open-Source E2E FHE Implementation for Privacy-Preserving Llama 3 8B Inference (arxiv.org)
2 points by simonpure 6 days ago | hide | past | pdf | discuss
50. TuxBot: Semantic-Aware Online OS Tuning with Large Language Models (arxiv.org)
1 point by matt_d 6 days ago | hide | past | pdf | discuss
51. Quantized Reasoning Models Think They Need to Think Longer, but They Do Not (arxiv.org)
13 points by theanonymousone 6 days ago | hide | past | pdf | 1 comment
52. "As a Language Model": Chat Template Switches LLM Self-Referential Voice (arxiv.org)
103 points by yu3zhou4 6 days ago | hide | past | pdf | 110 comments
53. LLM Agents Can Easily Tamper with Their Own Traces (arxiv.org)
3 points by sbulaev 6 days ago | hide | past | pdf | 1 comment
54. Just Ask Jev: Reinforcement Learning for Calibrated Decisions (arxiv.org)
1 point by Anon84 7 days ago | hide | past | pdf | discuss
55. DeepSeek Elastic Compute (DSec) (arxiv.org)
323 points by shenli3514 7 days ago | hide | past | pdf | 116 comments
56. Grow the Harness, Not the Context (arxiv.org)
1 point by acossta 7 days ago | hide | past | pdf | discuss
57. Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion (arxiv.org)
1 point by sbulaev 7 days ago | hide | past | pdf | discuss
58. What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models (arxiv.org)
2 points by daniel_iversen 7 days ago | hide | past | pdf | discuss
59. Transformer Can Hold Two Thoughts at Once: Evidence of Linear (arxiv.org)
3 points by sbulaev 7 days ago | hide | past | pdf | discuss
60. Skill-Guided Mining and Compilation of LLM Agent Traces (arxiv.org)
1 point by nlpnerd 8 days ago | hide | past | pdf | discuss