| 4351. |
Bypassing the Popularity Bias: Repurposing Models for Long-Tail Recommendation (arxiv.org) |
|
2 points by PaulHoule on Oct 22, 2024 | hide | past | pdf | discuss
|
| 4352. |
What Makes Large Language Models Reason in (Multi-Turn) Code Generation? (arxiv.org) |
|
1 point by dyekuu on Oct 22, 2024 | hide | past | pdf | discuss
|
| 4353. |
Guide to Fine-Tuning LLMs (arxiv.org) |
|
157 points by ignoramous on Oct 22, 2024 | hide | past | pdf | 16 comments
|
| 4354. |
Artificial Kuramoto Oscillatory Neurons (arxiv.org) |
|
4 points by lukeplato on Oct 22, 2024 | hide | past | pdf | discuss
|
| 4355. |
Allegro: Open the Black Box of Commercial-Level Video Generation Model (arxiv.org) |
|
2 points by jinqueeny on Oct 22, 2024 | hide | past | pdf | discuss
|
| 4356. |
Transformers Utilization in Chart Understanding: A Review of Advances and Future (arxiv.org) |
|
39 points by sandwichsphinx on Oct 22, 2024 | hide | past | pdf | 2 comments
|
| 4357. |
(Somewhat) Recent paper on model collapse (arxiv.org) |
|
1 point by Wheatman on Oct 21, 2024 | hide | past | pdf | 2 comments
|
| 4358. |
Persistent Pre-Training Poisoning of LLMs (arxiv.org) |
|
1 point by geox on Oct 21, 2024 | hide | past | pdf | discuss
|
| 4359. |
Machine Learning to Computational Plasma Physics Reduced-Order Plasma Modeling (arxiv.org) |
|
20 points by sandwichsphinx on Oct 21, 2024 | hide | past | pdf | 1 comment
|
| 4360. |
LLMs are overparameterized for text embedding tasks and can be easily pruned (arxiv.org) |
|
3 points by hdvr on Oct 21, 2024 | hide | past | pdf | discuss
|
| 4361. |
Good Parenting is all you need – Multi-agentic LLM Hallucination Mitigation (arxiv.org) |
|
1 point by belter on Oct 21, 2024 | hide | past | pdf | discuss
|
| 4362. |
Arcee's MergeKit: A Toolkit for Merging Large Language Models (arxiv.org) |
|
3 points by AnhTho_FR on Oct 20, 2024 | hide | past | pdf | discuss
|
| 4363. |
Thinking LLMs: General Instruction Following with Thought Generation (arxiv.org) |
|
3 points by ed on Oct 20, 2024 | hide | past | pdf | 1 comment
|
| 4364. |
Seeing Faces in Things: A Model and Dataset for Pareidolia (arxiv.org) |
|
1 point by PaulHoule on Oct 20, 2024 | hide | past | pdf | discuss
|
| 4365. |
Large Language Models in Finance: A Survey (arxiv.org) |
|
3 points by Anon84 on Oct 19, 2024 | hide | past | pdf | discuss
|
| 4366. |
Looking Inward: Language Models Can Learn About Themselves by Introspection (arxiv.org) |
|
2 points by ovaqre on Oct 19, 2024 | hide | past | pdf | discuss
|
| 4367. |
Autoregressive Large Language Models Are Computationally Universal (arxiv.org) |
|
1 point by Anon84 on Oct 19, 2024 | hide | past | pdf | discuss
|
| 4368. |
A (sorta) recent paper about model collapse has got me thinking (arxiv.org) |
|
3 points by Wheatman on Oct 19, 2024 | hide | past | pdf | 3 comments
|
| 4369. |
Does Refusal Training in LLMs Generalize to the Past Tense? (arxiv.org) |
|
2 points by fzliu on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4370. |
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture (arxiv.org) |
|
13 points by robg on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4371. |
Decoupling Visual Encoding for Unified Multimodal Understanding and Generation (arxiv.org) |
|
2 points by jinqueeny on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4372. |
NGPT: Normalized Transformer with Representation Learning on the Hypersphere (arxiv.org) |
|
4 points by programd on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4373. |
Understanding Slides and User Interfaces via Synthetic Data Generation (arxiv.org) |
|
1 point by PaulHoule on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4374. |
Reducing the Transformer Architecture to a Minimum (arxiv.org) |
|
2 points by belter on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4375. |
Numerical Precision Affects Mathematical Reasoning Capabilities of LLMs (arxiv.org) |
|
66 points by belter on Oct 18, 2024 | hide | past | pdf | 49 comments
|
| 4376. |
LLMD: A Large Language Model for Interpreting Longitudinal Medical Records (arxiv.org) |
|
48 points by troyastorino on Oct 18, 2024 | hide | past | pdf | 19 comments
|
| 4377. |
State-space models can learn in-context by gradient descent (arxiv.org) |
|
2 points by trextrex on Oct 18, 2024 | hide | past | pdf | discuss
|
| 4378. |
Why do random forests work? They are self-regularizing adaptive smoothers (arxiv.org) |
|
295 points by sebg on Oct 17, 2024 | hide | past | pdf | 41 comments
|
| 4379. |
From Commands to Prompts: LLM-Based Semantic File System for AIOS (arxiv.org) |
|
1 point by sandwichsphinx on Oct 17, 2024 | hide | past | pdf | discuss
|
| 4380. |
Improving Spoken Language Modeling with Phoneme Classification (arxiv.org) |
|
1 point by PaulHoule on Oct 17, 2024 | hide | past | pdf | discuss
|
| More |