about
4351. Bypassing the Popularity Bias: Repurposing Models for Long-Tail Recommendation (arxiv.org)
2 points by PaulHoule on Oct 22, 2024 | hide | past | pdf | discuss
4352. What Makes Large Language Models Reason in (Multi-Turn) Code Generation? (arxiv.org)
1 point by dyekuu on Oct 22, 2024 | hide | past | pdf | discuss
4353. Guide to Fine-Tuning LLMs (arxiv.org)
157 points by ignoramous on Oct 22, 2024 | hide | past | pdf | 16 comments
4354. Artificial Kuramoto Oscillatory Neurons (arxiv.org)
4 points by lukeplato on Oct 22, 2024 | hide | past | pdf | discuss
4355. Allegro: Open the Black Box of Commercial-Level Video Generation Model (arxiv.org)
2 points by jinqueeny on Oct 22, 2024 | hide | past | pdf | discuss
4356. Transformers Utilization in Chart Understanding: A Review of Advances and Future (arxiv.org)
39 points by sandwichsphinx on Oct 22, 2024 | hide | past | pdf | 2 comments
4357. (Somewhat) Recent paper on model collapse (arxiv.org)
1 point by Wheatman on Oct 21, 2024 | hide | past | pdf | 2 comments
4358. Persistent Pre-Training Poisoning of LLMs (arxiv.org)
1 point by geox on Oct 21, 2024 | hide | past | pdf | discuss
4359. Machine Learning to Computational Plasma Physics Reduced-Order Plasma Modeling (arxiv.org)
20 points by sandwichsphinx on Oct 21, 2024 | hide | past | pdf | 1 comment
4360. LLMs are overparameterized for text embedding tasks and can be easily pruned (arxiv.org)
3 points by hdvr on Oct 21, 2024 | hide | past | pdf | discuss
4361. Good Parenting is all you need – Multi-agentic LLM Hallucination Mitigation (arxiv.org)
1 point by belter on Oct 21, 2024 | hide | past | pdf | discuss
4362. Arcee's MergeKit: A Toolkit for Merging Large Language Models (arxiv.org)
3 points by AnhTho_FR on Oct 20, 2024 | hide | past | pdf | discuss
4363. Thinking LLMs: General Instruction Following with Thought Generation (arxiv.org)
3 points by ed on Oct 20, 2024 | hide | past | pdf | 1 comment
4364. Seeing Faces in Things: A Model and Dataset for Pareidolia (arxiv.org)
1 point by PaulHoule on Oct 20, 2024 | hide | past | pdf | discuss
4365. Large Language Models in Finance: A Survey (arxiv.org)
3 points by Anon84 on Oct 19, 2024 | hide | past | pdf | discuss
4366. Looking Inward: Language Models Can Learn About Themselves by Introspection (arxiv.org)
2 points by ovaqre on Oct 19, 2024 | hide | past | pdf | discuss
4367. Autoregressive Large Language Models Are Computationally Universal (arxiv.org)
1 point by Anon84 on Oct 19, 2024 | hide | past | pdf | discuss
4368. A (sorta) recent paper about model collapse has got me thinking (arxiv.org)
3 points by Wheatman on Oct 19, 2024 | hide | past | pdf | 3 comments
4369. Does Refusal Training in LLMs Generalize to the Past Tense? (arxiv.org)
2 points by fzliu on Oct 18, 2024 | hide | past | pdf | discuss
4370. Agents Thinking Fast and Slow: A Talker-Reasoner Architecture (arxiv.org)
13 points by robg on Oct 18, 2024 | hide | past | pdf | discuss
4371. Decoupling Visual Encoding for Unified Multimodal Understanding and Generation (arxiv.org)
2 points by jinqueeny on Oct 18, 2024 | hide | past | pdf | discuss
4372. NGPT: Normalized Transformer with Representation Learning on the Hypersphere (arxiv.org)
4 points by programd on Oct 18, 2024 | hide | past | pdf | discuss
4373. Understanding Slides and User Interfaces via Synthetic Data Generation (arxiv.org)
1 point by PaulHoule on Oct 18, 2024 | hide | past | pdf | discuss
4374. Reducing the Transformer Architecture to a Minimum (arxiv.org)
2 points by belter on Oct 18, 2024 | hide | past | pdf | discuss
4375. Numerical Precision Affects Mathematical Reasoning Capabilities of LLMs (arxiv.org)
66 points by belter on Oct 18, 2024 | hide | past | pdf | 49 comments
4376. LLMD: A Large Language Model for Interpreting Longitudinal Medical Records (arxiv.org)
48 points by troyastorino on Oct 18, 2024 | hide | past | pdf | 19 comments
4377. State-space models can learn in-context by gradient descent (arxiv.org)
2 points by trextrex on Oct 18, 2024 | hide | past | pdf | discuss
4378. Why do random forests work? They are self-regularizing adaptive smoothers (arxiv.org)
295 points by sebg on Oct 17, 2024 | hide | past | pdf | 41 comments
4379. From Commands to Prompts: LLM-Based Semantic File System for AIOS (arxiv.org)
1 point by sandwichsphinx on Oct 17, 2024 | hide | past | pdf | discuss
4380. Improving Spoken Language Modeling with Phoneme Classification (arxiv.org)
1 point by PaulHoule on Oct 17, 2024 | hide | past | pdf | discuss