about
5191. Gradient Diversity: A Key Ingredient for Scalable Distributed Learning (arxiv.org)
3 points by veryluckyxyz on May 14, 2024 | hide | past | pdf | discuss
5192. Learning to Slice Wi-Fi Networks: A State-Augmented Primal-Dual Approach (arxiv.org)
1 point by loongloong on May 14, 2024 | hide | past | pdf | discuss
5193. AccEar: Accelerometer Acoustic Eavesdropping (2022) (arxiv.org)
2 points by api on May 14, 2024 | hide | past | pdf | discuss
5194. Scaling Autoregressive Multimodal Models with Vision, Text, Audio, and Action (arxiv.org)
4 points by georgehill on May 14, 2024 | hide | past | pdf | discuss
5195. Generalization in diffusion models arises from geometry-adaptive harmonic reps (arxiv.org)
1 point by RafelMri on May 14, 2024 | hide | past | pdf | discuss
5196. GPT-4 passes most of the 297 written Polish Board Certification Examinations (arxiv.org)
1 point by PaulHoule on May 14, 2024 | hide | past | pdf | discuss
5197. Towards Guaranteed Safe AI (arxiv.org)
2 points by belter on May 13, 2024 | hide | past | pdf | discuss
5198. Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models (arxiv.org)
1 point by veryluckyxyz on May 13, 2024 | hide | past | pdf | discuss
5199. LLMs on Tabular Data: Prediction, Generation, and Understanding – A Survey (arxiv.org)
2 points by Anon84 on May 13, 2024 | hide | past | pdf | discuss
5200. Meta FAIR: Memory Mosaics (arxiv.org)
1 point by georgehill on May 13, 2024 | hide | past | pdf | discuss
5201. Do Llamas Work in English? On the Latent Language of Multilingual Transformers (arxiv.org)
1 point by georgehill on May 13, 2024 | hide | past | pdf | discuss
5202. Linearizing Large Language Models (arxiv.org)
2 points by jasondavies on May 13, 2024 | hide | past | pdf | discuss
5203. Wit, Creativity, and Detectability of LLMs Adapted to Reddit's Showerthoughts (arxiv.org)
3 points by PaulHoule on May 13, 2024 | hide | past | pdf | discuss
5204. Plan of Thoughts: Heuristic-Guided Problem Solving with Large Language Models (arxiv.org)
3 points by jemoka on May 12, 2024 | hide | past | pdf | discuss
5205. In-Context Symbolic Regression: Using Language Models for Function Discovery (arxiv.org)
3 points by PaulHoule on May 12, 2024 | hide | past | pdf | discuss
5206. Automatically Detecting Under-Trained Tokens in Large Language Models (arxiv.org)
182 points by veryluckyxyz on May 12, 2024 | hide | past | pdf | 26 comments
5207. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs (SOTA sub-4-bit inference) (arxiv.org)
1 point by Hugsun on May 11, 2024 | hide | past | pdf | discuss
5208. Language Modeling Using Tensor Trains (arxiv.org)
2 points by PaulHoule on May 11, 2024 | hide | past | pdf | discuss
5209. Multidirectional joint distribution neurons reducing to KAN (arxiv.org)
36 points by jarekd on May 11, 2024 | hide | past | pdf | 13 comments
5210. You Only Cache Once: Decoder-Decoder Architectures for Language Models (arxiv.org)
3 points by reqo on May 10, 2024 | hide | past | pdf | 1 comment
5211. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? (arxiv.org)
36 points by Jimmc414 on May 10, 2024 | hide | past | pdf | 17 comments
5212. Aligning Large Language Models with Recommendation Knowledge (arxiv.org)
1 point by rntn on May 10, 2024 | hide | past | pdf | discuss
5213. Energy-Efficient Llama 2 Inference on FPGAs via High Level Synthesis (arxiv.org)
94 points by PaulHoule on May 10, 2024 | hide | past | pdf | 29 comments
5214. CNN-Based Equalization for Communications: Gigabit Throughput with FPGA (arxiv.org)
3 points by PaulHoule on May 9, 2024 | hide | past | pdf | discuss
5215. Chain of Thoughtlessness: An Analysis of Cot in Planning (arxiv.org)
2 points by mnk47 on May 9, 2024 | hide | past | pdf | discuss
5216. A Survey on the Real Power of ChatGPT (arxiv.org)
3 points by PaulHoule on May 9, 2024 | hide | past | pdf | discuss
5217. LLMs Can Patch Up Missing Relevance Judgments in Evaluation (arxiv.org)
1 point by lnyan on May 9, 2024 | hide | past | pdf | discuss
5218. Instructing Robots by Sketching: Learning from Demonstration (arxiv.org)
2 points by programd on May 9, 2024 | hide | past | pdf | discuss
5219. Graph Neural Network Approach to Semantic Type Detection in Tables (arxiv.org)
3 points by PaulHoule on May 9, 2024 | hide | past | pdf | discuss
5220. No "Zero-Shot" Without Exponential Data (arxiv.org)
187 points by zerojames on May 9, 2024 | hide | past | pdf | 118 comments