about
4291. Revisiting Reliability in Large-Scale Machine Learning Research Clusters (arxiv.org)
1 point by mfiguiere on Oct 31, 2024 | hide | past | pdf | discuss
4292. The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org)
2 points by robg on Oct 30, 2024 | hide | past | pdf | discuss
4293. Chain-of-thought can hurt performance on tasks where thinking makes humans worse (arxiv.org)
371 points by benocodes on Oct 30, 2024 | hide | past | pdf | 250 comments
4294. One Model to Learn Them All (arxiv.org)
5 points by Anon84 on Oct 30, 2024 | hide | past | pdf | discuss
4295. Deep Optimizer States: Scalable Training of Transformer Interleaved Offloading (arxiv.org)
1 point by sandwichsphinx on Oct 30, 2024 | hide | past | pdf | discuss
4296. Conditional Hallucinations for Image Compression (arxiv.org)
1 point by Hard_Space on Oct 30, 2024 | hide | past | pdf | discuss
4297. LLMs know more than they show: On the intrinsic representation of hallucinations (arxiv.org)
137 points by benocodes on Oct 30, 2024 | hide | past | pdf | 140 comments
4298. Designing Robust Cyber-Defense Agents with Evolving Behavior Trees (arxiv.org)
3 points by PaulHoule on Oct 30, 2024 | hide | past | pdf | discuss
4299. Generator Matching: Generative modeling with arbitrary Markov processes (arxiv.org)
1 point by lnyan on Oct 30, 2024 | hide | past | pdf | discuss
4300. The AI Scientist: Towards Automated Open-Ended Scientific Discovery (arxiv.org)
2 points by Anon84 on Oct 30, 2024 | hide | past | pdf | discuss
4301. The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org)
2 points by roboboffin on Oct 30, 2024 | hide | past | pdf | discuss
4302. LLMmap: Fingerprinting for Large Language Models (arxiv.org)
2 points by emk_709 on Oct 30, 2024 | hide | past | pdf | 2 comments
4303. Hacking Back the AI-Hacker: Prompt Injection as a Defense for LLM-Attackers (arxiv.org)
2 points by emk_709 on Oct 30, 2024 | hide | past | pdf | discuss
4304. GPT-4o System Card [pdf] (arxiv.org)
1 point by SerCe on Oct 30, 2024 | hide | past | pdf | 1 comment
4305. SQFT: Low-Cost Model Adaptation in Low-Precision Sparse Foundation Models (arxiv.org)
3 points by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
4306. LLM Code Generation with Formal Specifications and Reactive Program Synthesis (arxiv.org)
1 point by sandwichsphinx on Oct 29, 2024 | hide | past | pdf | discuss
4307. Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-Wise LoRA (arxiv.org)
1 point by famouswaffles on Oct 29, 2024 | hide | past | pdf | discuss
4308. Acer: Automatic Language Model Context Extension via Retrieval (arxiv.org)
1 point by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
4309. LLMProxy: Reducing Cost to Access Large Language Models (arxiv.org)
2 points by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
4310. Encoding Agent Trajectories as Representations with Sequence Transformers (arxiv.org)
1 point by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
4311. Mathematical Theory of Deep Learning (arxiv.org)
3 points by nabla9 on Oct 28, 2024 | hide | past | pdf | discuss
4312. ReasonAgain: Using LLM Generated Symbolic Programs for Mathematical Reasoning (arxiv.org)
2 points by hendler on Oct 28, 2024 | hide | past | pdf | discuss
4313. Memory-augmented Transformers can implement Linear first-Order Optimization (arxiv.org)
1 point by PaulHoule on Oct 28, 2024 | hide | past | pdf | discuss
4314. O1 Replication Journey (arxiv.org)
1 point by fofoz on Oct 28, 2024 | hide | past | pdf | discuss
4315. Mixture of Parrots: Experts improve memorization more than reasoning (arxiv.org)
3 points by sandwichsphinx on Oct 28, 2024 | hide | past | pdf | discuss
4316. Faithful and Sound Reasoning on Knowledge Graphs (arxiv.org)
2 points by brainless on Oct 28, 2024 | hide | past | pdf | 1 comment
4317. Large Language Models Reflect the Ideology of Their Creators (arxiv.org)
4 points by delian66 on Oct 28, 2024 | hide | past | pdf | 2 comments
4318. Agents Thinking Fast and Slow: A Talker-Reasoner Architecture (arxiv.org)
9 points by benocodes on Oct 28, 2024 | hide | past | pdf | discuss
4319. Improving Pinterest Search Relevance Using LLMs (arxiv.org)
2 points by teej on Oct 27, 2024 | hide | past | pdf | discuss
4320. Solving Global Lyapunov functions: open problem in mathematics with transformers (arxiv.org)
3 points by famouswaffles on Oct 27, 2024 | hide | past | pdf | discuss