about
4261. Mastering the Craft of Data Synthesis for CodeLLMs (arxiv.org)
1 point by sandwichsphinx on Nov 4, 2024 | hide | past | pdf | discuss
4262. Thinking LLMs: General Instruction Following with Thought Generation (arxiv.org)
1 point by timbilt on Nov 4, 2024 | hide | past | pdf | discuss
4263. An embarrassingly simple approach to recover unlearned knowledge for LLMs (arxiv.org)
259 points by PaulHoule on Nov 4, 2024 | hide | past | pdf | 121 comments
4264. Voice-Enabled AI Agents Can Perform Common Scams (arxiv.org)
1 point by sandwichsphinx on Nov 3, 2024 | hide | past | pdf | discuss
4265. Interpreting Affine Recurrence Learning in GPT-Style Transformers (arxiv.org)
2 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
4266. Improving Neuron-Level Interpretability with White-Box Language Models (arxiv.org)
3 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
4267. Improving Embedding Accuracy for Using ER Maps and Model-Aware Sampling (arxiv.org)
2 points by PaulHoule on Nov 3, 2024 | hide | past | pdf | discuss
4268. EmbodiedRAG (arxiv.org)
1 point by KeyurRamoliya on Nov 3, 2024 | hide | past | pdf | discuss
4269. Length-Induced Embedding Collapse in Transformer-Based Models (arxiv.org)
3 points by Wheatman on Nov 3, 2024 | hide | past | pdf | discuss
4270. Spann: Highly-Efficient Billion-Scale Approximate Nearest Neighbor Search (2021) (arxiv.org)
124 points by ksec on Nov 2, 2024 | hide | past | pdf | 33 comments
4271. Context-Augmented Code Generation Using Programming Knowledge Graphs (arxiv.org)
2 points by PaulHoule on Nov 2, 2024 | hide | past | pdf | discuss
4272. Llama-Berry: Pairwise Optimization for O1-Like Mathematical Reasoning (arxiv.org)
2 points by pongogogo on Nov 2, 2024 | hide | past | pdf | discuss
4273. Language Models Learn to Mislead Humans via RLHF (arxiv.org)
3 points by Anon84 on Nov 2, 2024 | hide | past | pdf | 1 comment
4274. Video-ChatGPT: Towards Video Understanding via Large Vision and Language Models (arxiv.org)
2 points by godelmachine on Nov 1, 2024 | hide | past | pdf | discuss
4275. Hypothetical Document Embeddings (HyDE) for Precise Zero-Shot Retrieval [pdf] (arxiv.org)
2 points by TaurenHunter on Nov 1, 2024 | hide | past | pdf | discuss
4276. Understanding Warmup-Stable-Decay Learning Rates (arxiv.org)
1 point by fzliu on Nov 1, 2024 | hide | past | pdf | discuss
4277. Fast and Accurate Deep Reconfigurable Spiking Inference Accelerator Architecture (arxiv.org)
2 points by PaulHoule on Nov 1, 2024 | hide | past | pdf | discuss
4278. FVEval: Language Model Capabilities in Formal Verification of Digital Hardware (arxiv.org)
1 point by sandwichsphinx on Nov 1, 2024 | hide | past | pdf | discuss
4279. BrainTransformers: SNN-LLM (arxiv.org)
2 points by PaulHoule on Nov 1, 2024 | hide | past | pdf | discuss
4280. TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters (arxiv.org)
174 points by famouswaffles on Nov 1, 2024 | hide | past | pdf | 33 comments
4281. The AI Scientist: Towards Automated Open-Ended Scientific Discovery (arxiv.org)
2 points by belter on Nov 1, 2024 | hide | past | pdf | discuss
4282. A Large Recurrent Action Model: xLSTM Enables Fast Inference for Robotics Tasks (arxiv.org)
2 points by tosh on Oct 31, 2024 | hide | past | pdf | discuss
4283. Transformers Are Efficient Compilers, Provably (arxiv.org)
1 point by PaulHoule on Oct 31, 2024 | hide | past | pdf | discuss
4284. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges (arxiv.org)
2 points by Anon84 on Oct 31, 2024 | hide | past | pdf | discuss
4285. Tokenformer: Rethinking transformer scaling with tokenized model parameters (arxiv.org)
3 points by andy12_ on Oct 31, 2024 | hide | past | pdf | 1 comment
4286. Universality of the π²/6 Pathway in Avoiding Model Collapse [pdf] (arxiv.org)
1 point by bikenaga on Oct 31, 2024 | hide | past | pdf | discuss
4287. A Prescriptive Theory for Brain-Like Inference (arxiv.org)
2 points by liamdgray on Oct 31, 2024 | hide | past | pdf | 1 comment
4288. Revisiting Reliability in Large-Scale Machine Learning Research Clusters (arxiv.org)
1 point by mfiguiere on Oct 31, 2024 | hide | past | pdf | discuss
4289. The Geometry of Concepts: Sparse Autoencoder Feature Structure (arxiv.org)
2 points by robg on Oct 30, 2024 | hide | past | pdf | discuss
4290. Chain-of-thought can hurt performance on tasks where thinking makes humans worse (arxiv.org)
371 points by benocodes on Oct 30, 2024 | hide | past | pdf | 250 comments