about
1321. Prism: Demystifying Retention and Interaction in Mid-Training (arxiv.org)
1 point by xhevahir 197 days ago | hide | past | pdf | discuss
1322. Dissociating Direct Access from Inference in AI Introspection (arxiv.org)
3 points by 3willows 197 days ago | hide | past | pdf | discuss
1323. Interwhen: A Generalizable Framework for Verifiable Reasoning (arxiv.org)
2 points by dimmuborgir 198 days ago | hide | past | pdf | discuss
1324. SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels (arxiv.org)
3 points by matt_d 198 days ago | hide | past | pdf | discuss
1325. Transformers Are Bayesian Networks (arxiv.org)
40 points by Anon84 198 days ago | hide | past | pdf | 33 comments
1326. Matrix Valued Residuals (arxiv.org)
5 points by E-Reverance 198 days ago | hide | past | pdf | discuss
1327. MemEvolve: Meta-Evolution of Agent Memory Systems (arxiv.org)
2 points by dennisy 198 days ago | hide | past | pdf | 1 comment
1328. M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Modeling (arxiv.org)
3 points by gmays 198 days ago | hide | past | pdf | discuss
1329. Democratizing GraphRAG: Linear, CPU-Only Graph Retrieval for Multi-Hop QA (arxiv.org)
2 points by PaulHoule 198 days ago | hide | past | pdf | discuss
1330. AlgoVeri: An Aligned Benchmark for Verified Code Gen. On Classical Algorithms (arxiv.org)
2 points by matt_d 198 days ago | hide | past | pdf | discuss
1331. The Missing Memory Hierarchy: Demand Paging for LLM Context Windows (arxiv.org)
3 points by brewcrew 198 days ago | hide | past | pdf | discuss
1332. How Do LLMs Compute Verbal Confidence (DeepMind) (arxiv.org)
3 points by armcat 199 days ago | hide | past | pdf | discuss
1333. Evaluating Genuine Reasoning in LLMs via Esoteric Programming Languages (arxiv.org)
1 point by kerneis 199 days ago | hide | past | pdf | 1 comment
1334. Reasoning Core: Procedural Data Generation Suite for Symbolic Pre-Training (arxiv.org)
2 points by jean-porte 199 days ago | hide | past | pdf | discuss
1335. Have LLMs Learned to Reason? A Characterization via 3-SAT Phase Transition (arxiv.org)
3 points by jacklondon 199 days ago | hide | past | pdf | 2 comments
1336. M^2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling (arxiv.org)
2 points by matt_d 199 days ago | hide | past | pdf | discuss
1337. NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL (arxiv.org)
3 points by matt_d 199 days ago | hide | past | pdf | discuss
1338. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (arxiv.org)
3 points by RickJWagner 199 days ago | hide | past | pdf | discuss
1339. Do Large Language Models Get Caught in Hofstadter-Mobius Loops? (arxiv.org)
2 points by pndy 200 days ago | hide | past | pdf | 1 comment
1340. Design Conductor: agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU (arxiv.org)
4 points by EvgeniyZh 200 days ago | hide | past | pdf | 1 comment
1341. SkillNet: Create, Evaluate, and Connect AI Skills (arxiv.org)
1 point by navikohli 200 days ago | hide | past | pdf | discuss
1342. Epiplexity: Rethinking Information for Computationally Bounded Intelligence (arxiv.org)
2 points by fritzo 200 days ago | hide | past | pdf | 2 comments
1343. Mamba-3: Improved Sequence Modeling Using State Space Principles (arxiv.org)
4 points by anentropic 201 days ago | hide | past | pdf | 1 comment
1344. UC Irvine researchers bring down AI powered drones with painted umbrellas (arxiv.org)
23 points by jcalvinowens 201 days ago | hide | past | pdf | 8 comments
1345. Why AI systems don't learn – On autonomous learning from cognitive science (arxiv.org)
205 points by aanet 201 days ago | hide | past | pdf | 116 comments
1346. Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters (arxiv.org)
1 point by PaulHoule 201 days ago | hide | past | pdf | 1 comment
1347. Benchmarking Distilled Language Models for Performance and Efficiency (arxiv.org)
2 points by PaulHoule 201 days ago | hide | past | pdf | discuss
1348. Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on SeL4 (arxiv.org)
1 point by PaulHoule 201 days ago | hide | past | pdf | discuss
1349. We Ran the Largest AI Pokemon Tournament Ever. Now It's an Open Benchmark (arxiv.org)
1 point by milkkarten 201 days ago | hide | past | pdf | discuss
1350. Automating Forecasting Question Generation and Resolution for AI Evaluation (arxiv.org)
3 points by nbosse 201 days ago | hide | past | pdf | discuss