about
31. Self-Play Pretraining with Zero Data (arxiv.org)
4 points by E-Reverance 8 days ago | hide | past | pdf | discuss
32. Loopjacking: Hijacking Human-in-the-Loop Approval (arxiv.org)
4 points by sbulaev 11 days ago | hide | past | pdf | discuss
33. Solving rubiks cubes "without search" (arxiv.org)
4 points by E-Reverance 14 days ago | hide | past | pdf | discuss
34. Neural Turing Machines (2014) (arxiv.org)
4 points by peter_d_sherman 20 days ago | hide | past | pdf | discuss
35. Copying explains the collective behavior of AI agents in the wild (arxiv.org)
4 points by sbulaev 23 days ago | hide | past | pdf | discuss
36. Uncensored Open-Weight Models: Redistribution as the Persistence Layer (arxiv.org)
4 points by sbulaev 25 days ago | hide | past | pdf | discuss
37. The Measure of Intelligence (2019) (arxiv.org)
4 points by theanonymousone 29 days ago | hide | past | pdf | discuss
38. Foundations of Large Language Models (arxiv.org)
3 points by rramadass 5 days ago | hide | past | pdf | 1 comment
39. LLM Agents Can Easily Tamper with Their Own Traces (arxiv.org)
3 points by sbulaev 6 days ago | hide | past | pdf | 1 comment
40. Implicit Neural Representations with Periodic Activation Functions (2020) (arxiv.org)
3 points by peter_d_sherman 8 days ago | hide | past | pdf | 1 comment
41. Recursive self-improvement of AI research agents (arxiv.org)
3 points by handfuloflight 10 days ago | hide | past | pdf | 1 comment
42. Tracking Capabilities for Safer Agents (arxiv.org)
3 points by verdverm 13 days ago | hide | past | pdf | 1 comment
43. Potemkin Understanding in Large Language Models (2025) (arxiv.org)
3 points by mpweiher 17 days ago | hide | past | pdf | 1 comment
44. CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI (arxiv.org)
3 points by taubek 25 days ago | hide | past | pdf | 1 comment
45. ByteX: A Unified AI Search Engine at ByteDance (arxiv.org)
3 points by softwaredoug 27 days ago | hide | past | pdf | 1 comment
46. PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization (arxiv.org)
3 points by matt_d 1 day ago | hide | past | pdf | discuss
47. AI Agents Are Vulnerable to Radicalization (arxiv.org)
3 points by Anon84 1 day ago | hide | past | pdf | discuss
48. Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text (arxiv.org)
3 points by sbulaev 1 day ago | hide | past | pdf | discuss
49. Pretraining Latent Information Feedback Transformers with Teacher Supervision (arxiv.org)
3 points by gmays 2 days ago | hide | past | pdf | discuss
50. Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning (arxiv.org)
3 points by matt_d 3 days ago | hide | past | pdf | discuss
51. Compiling Triton kernels without the Triton compiler (arxiv.org)
3 points by 50kIters 3 days ago | hide | past | pdf | discuss
52. Purlin: Separating Orchestration from the Datapath of Collectives (arxiv.org)
3 points by matt_d 3 days ago | hide | past | pdf | discuss
53. GenRec: An LLM-Backed Recommendation Ranker at NetflixConference (arxiv.org)
3 points by throwthrowrow 5 days ago | hide | past | pdf | discuss
54. Transformer Can Hold Two Thoughts at Once: Evidence of Linear (arxiv.org)
3 points by sbulaev 7 days ago | hide | past | pdf | discuss
55. Testing AIs on 68 of the hardest open Erdos problems, verified in Lean (arxiv.org)
3 points by tadamcz 10 days ago | hide | past | pdf | discuss
56. Xeno-Interpretability: Investigating the Alien Minds of LLMs (arxiv.org)
3 points by potent_latent 11 days ago | hide | past | pdf | discuss
57. Do small language models know what they don't know? (arxiv.org)
3 points by Brajeshwar 11 days ago | hide | past | pdf | discuss
58. Atria Dawn: The Dawn of Agentic Superintelligence (arxiv.org)
3 points by simonpure 12 days ago | hide | past | pdf | discuss
59. Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering (arxiv.org)
3 points by matt_d 14 days ago | hide | past | pdf | discuss
60. Score Centering Stabilizes Off-Policy Reinforcement Learning (arxiv.org)
3 points by zagwdt 15 days ago | hide | past | pdf | discuss