| 4291. |
Revisiting Reliability in Large-Scale Machine Learning Research Clusters (arxiv.org) |
|
1 point by mfiguiere on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4292. |
The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org) |
|
2 points by robg on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4293. |
Chain-of-thought can hurt performance on tasks where thinking makes humans worse (arxiv.org) |
|
371 points by benocodes on Oct 30, 2024 | hide | past | pdf | 250 comments
|
| 4294. |
One Model to Learn Them All (arxiv.org) |
|
5 points by Anon84 on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4295. |
Deep Optimizer States: Scalable Training of Transformer Interleaved Offloading (arxiv.org) |
|
1 point by sandwichsphinx on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4296. |
Conditional Hallucinations for Image Compression (arxiv.org) |
|
1 point by Hard_Space on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4297. |
LLMs know more than they show: On the intrinsic representation of hallucinations (arxiv.org) |
|
137 points by benocodes on Oct 30, 2024 | hide | past | pdf | 140 comments
|
| 4298. |
Designing Robust Cyber-Defense Agents with Evolving Behavior Trees (arxiv.org) |
|
3 points by PaulHoule on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4299. |
Generator Matching: Generative modeling with arbitrary Markov processes (arxiv.org) |
|
1 point by lnyan on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4300. |
The AI Scientist: Towards Automated Open-Ended Scientific Discovery (arxiv.org) |
|
2 points by Anon84 on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4301. |
The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org) |
|
2 points by roboboffin on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4302. |
LLMmap: Fingerprinting for Large Language Models (arxiv.org) |
|
2 points by emk_709 on Oct 30, 2024 | hide | past | pdf | 2 comments
|
| 4303. |
Hacking Back the AI-Hacker: Prompt Injection as a Defense for LLM-Attackers (arxiv.org) |
|
2 points by emk_709 on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4304. |
GPT-4o System Card [pdf] (arxiv.org) |
|
1 point by SerCe on Oct 30, 2024 | hide | past | pdf | 1 comment
|
| 4305. |
SQFT: Low-Cost Model Adaptation in Low-Precision Sparse Foundation Models (arxiv.org) |
|
3 points by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4306. |
LLM Code Generation with Formal Specifications and Reactive Program Synthesis (arxiv.org) |
|
1 point by sandwichsphinx on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4307. |
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-Wise LoRA (arxiv.org) |
|
1 point by famouswaffles on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4308. |
Acer: Automatic Language Model Context Extension via Retrieval (arxiv.org) |
|
1 point by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4309. |
LLMProxy: Reducing Cost to Access Large Language Models (arxiv.org) |
|
2 points by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4310. |
Encoding Agent Trajectories as Representations with Sequence Transformers (arxiv.org) |
|
1 point by PaulHoule on Oct 29, 2024 | hide | past | pdf | discuss
|
| 4311. |
Mathematical Theory of Deep Learning (arxiv.org) |
|
3 points by nabla9 on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4312. |
ReasonAgain: Using LLM Generated Symbolic Programs for Mathematical Reasoning (arxiv.org) |
|
2 points by hendler on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4313. |
Memory-augmented Transformers can implement Linear first-Order Optimization (arxiv.org) |
|
1 point by PaulHoule on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4314. |
O1 Replication Journey (arxiv.org) |
|
1 point by fofoz on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4315. |
Mixture of Parrots: Experts improve memorization more than reasoning (arxiv.org) |
|
3 points by sandwichsphinx on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4316. |
Faithful and Sound Reasoning on Knowledge Graphs (arxiv.org) |
|
2 points by brainless on Oct 28, 2024 | hide | past | pdf | 1 comment
|
| 4317. |
Large Language Models Reflect the Ideology of Their Creators (arxiv.org) |
|
4 points by delian66 on Oct 28, 2024 | hide | past | pdf | 2 comments
|
| 4318. |
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture (arxiv.org) |
|
9 points by benocodes on Oct 28, 2024 | hide | past | pdf | discuss
|
| 4319. |
Improving Pinterest Search Relevance Using LLMs (arxiv.org) |
|
2 points by teej on Oct 27, 2024 | hide | past | pdf | discuss
|
| 4320. |
Solving Global Lyapunov functions: open problem in mathematics with transformers (arxiv.org) |
|
3 points by famouswaffles on Oct 27, 2024 | hide | past | pdf | discuss
|
| More |