| 4261. |
Mastering the Craft of Data Synthesis for CodeLLMs (arxiv.org) |
|
1 point by sandwichsphinx on Nov 4, 2024 | hide | past | pdf | discuss
|
| 4262. |
Thinking LLMs: General Instruction Following with Thought Generation (arxiv.org) |
|
1 point by timbilt on Nov 4, 2024 | hide | past | pdf | discuss
|
| 4263. |
An embarrassingly simple approach to recover unlearned knowledge for LLMs (arxiv.org) |
|
259 points by PaulHoule on Nov 4, 2024 | hide | past | pdf | 121 comments
|
| 4264. |
Voice-Enabled AI Agents Can Perform Common Scams (arxiv.org) |
|
1 point by sandwichsphinx on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4265. |
Interpreting Affine Recurrence Learning in GPT-Style Transformers (arxiv.org) |
|
2 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4266. |
Improving Neuron-Level Interpretability with White-Box Language Models (arxiv.org) |
|
3 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4267. |
Improving Embedding Accuracy for Using ER Maps and Model-Aware Sampling (arxiv.org) |
|
2 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4268. |
EmbodiedRAG (arxiv.org) |
|
1 point by KeyurRamoliya on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4269. |
Length-Induced Embedding Collapse in Transformer-Based Models (arxiv.org) |
|
3 points by Wheatman on Nov 3, 2024 | hide | past | pdf | discuss
|
| 4270. |
Spann: Highly-Efficient Billion-Scale Approximate Nearest Neighbor Search (2021) (arxiv.org) |
|
124 points by ksec on Nov 2, 2024 | hide | past | pdf | 33 comments
|
| 4271. |
Context-Augmented Code Generation Using Programming Knowledge Graphs (arxiv.org) |
|
2 points by PaulHoule on Nov 2, 2024 | hide | past | pdf | discuss
|
| 4272. |
Llama-Berry: Pairwise Optimization for O1-Like Mathematical Reasoning (arxiv.org) |
|
2 points by pongogogo on Nov 2, 2024 | hide | past | pdf | discuss
|
| 4273. |
Language Models Learn to Mislead Humans via RLHF (arxiv.org) |
|
3 points by Anon84 on Nov 2, 2024 | hide | past | pdf | 1 comment
|
| 4274. |
Video-ChatGPT: Towards Video Understanding via Large Vision and Language Models (arxiv.org) |
|
2 points by godelmachine on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4275. |
Hypothetical Document Embeddings (HyDE) for Precise Zero-Shot Retrieval [pdf] (arxiv.org) |
|
2 points by TaurenHunter on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4276. |
Understanding Warmup-Stable-Decay Learning Rates (arxiv.org) |
|
1 point by fzliu on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4277. |
Fast and Accurate Deep Reconfigurable Spiking Inference Accelerator Architecture (arxiv.org) |
|
2 points by PaulHoule on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4278. |
FVEval: Language Model Capabilities in Formal Verification of Digital Hardware (arxiv.org) |
|
1 point by sandwichsphinx on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4279. |
BrainTransformers: SNN-LLM (arxiv.org) |
|
2 points by PaulHoule on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4280. |
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (arxiv.org) |
|
174 points by famouswaffles on Nov 1, 2024 | hide | past | pdf | 33 comments
|
| 4281. |
The AI Scientist: Towards Automated Open-Ended Scientific Discovery (arxiv.org) |
|
2 points by belter on Nov 1, 2024 | hide | past | pdf | discuss
|
| 4282. |
A Large Recurrent Action Model: xLSTM Enables Fast Inference for Robotics Tasks (arxiv.org) |
|
2 points by tosh on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4283. |
Transformers Are Efficient Compilers, Provably (arxiv.org) |
|
1 point by PaulHoule on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4284. |
Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges (arxiv.org) |
|
2 points by Anon84 on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4285. |
Tokenformer: Rethinking transformer scaling with tokenized model parameters (arxiv.org) |
|
3 points by andy12_ on Oct 31, 2024 | hide | past | pdf | 1 comment
|
| 4286. |
Universality of the π²/6 Pathway in Avoiding Model Collapse [pdf] (arxiv.org) |
|
1 point by bikenaga on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4287. |
A Prescriptive Theory for Brain-Like Inference (arxiv.org) |
|
2 points by liamdgray on Oct 31, 2024 | hide | past | pdf | 1 comment
|
| 4288. |
Revisiting Reliability in Large-Scale Machine Learning Research Clusters (arxiv.org) |
|
1 point by mfiguiere on Oct 31, 2024 | hide | past | pdf | discuss
|
| 4289. |
The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org) |
|
2 points by robg on Oct 30, 2024 | hide | past | pdf | discuss
|
| 4290. |
Chain-of-thought can hurt performance on tasks where thinking makes humans worse (arxiv.org) |
|
371 points by benocodes on Oct 30, 2024 | hide | past | pdf | 250 comments
|
| More |