about
Guide to Fine-Tuning LLMs (arxiv.org)
157 points by ignoramous on Oct 22, 2024 | hide | past | pdf | 16 comments on HN

In plain words: A review of fine-tuning large language models, laying out a seven-stage workflow from data prep to deployment and comparing approaches from supervised training to lightweight updates and human-preference alignment. It maps which techniques suit which tasks and flags hurdles like data imbalance, cost, and privacy.

Abstract · The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

This report examines the fine-tuning of Large Language Models (LLMs), integrating theoretical insights with practical applications. It outlines the historical evolution of LLMs from traditional Natural Language Processing (NLP) models to their pivotal role in AI. A comparison of fine-tuning methodologies, including supervised, unsupervised, and instruction-based approaches, highlights their applicability to different tasks. The report introduces a structured seven-stage pipeline for fine-tuning LLMs, spanning data preparation, model initialization, hyperparameter tuning, and model deployment. Emphasis is placed on managing imbalanced datasets and optimization techniques. Parameter-efficient methods like Low-Rank Adaptation (LoRA) and Half Fine-Tuning are explored for balancing computational efficiency with performance. Advanced techniques such as memory fine-tuning, Mixture of Experts (MoE), and Mixture of Agents (MoA) are discussed for leveraging specialized networks and multi-agent collaboration. The report also examines novel approaches like Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO), which align LLMs with human preferences, alongside pruning and routing optimizations to improve efficiency. Further sections cover validation frameworks, post-deployment monitoring, and inference optimization, with attention to deploying LLMs on distributed and cloud-based platforms. Emerging areas such as multimodal LLMs, fine-tuning for audio and speech, and challenges related to scalability, privacy, and accountability are also addressed. This report offers actionable insights for researchers and practitioners navigating LLM fine-tuning in an evolving landscape.

Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, Arsalan Shahid
arXiv:2408.13296 · cs.LG, cs.CL · submitted Aug 23, 2024 · updated Oct 30, 2024
abstract · pdf · html

add comment on HN
Also discussed: Jun 2026 (2 points, 1 comment)

I saw this when it was making the rounds on X a few days ago. Fair warning: it seems like at least some sections are AI-generated, and there isn't much insight to be gained from reading the actual sections compared to eg. reading the relevant category pages on Huggingface.
I took a skim through it in the morning - I like the LoRA Learns Less and Forgets Less paper more https://openreview.net/forum?id=aloEru2qCG - it has much more signal in a few pages - also the original QLoRA paper from Dettmers https://arxiv.org/abs/2305.14314 has so many more important morsels.

But all in all, the review is a reasonable "manual" I guess. I would have liked maybe more instructive comprehensive practical examples, and maybe more mention of other OSS packages for finetuning :))

Not only the it seems to be AI generated, it seems these guys don't even know about best practices or even what works. e.g. It contains archaic comparison of optimizers and its pros and cons, but for LLMs no optimizer other than Adam and new ones like Lion works.
Is there a paper on this? Why do no other optimizers give good results? Adam requires insane amounts of memory so alternatives would be welcome.
Glancing at the authors' names, it's possible that none of them are native English speakers. Any chance that the sections you're referring to were just AI-polished rather than AI-generated?
No, this paper was edited yesterday. The original (you can verify on arxiv) contained this incredible section: "6.10 Optimised Routing and Pruning Operations (ORPO)"

The actual ORPO paper is "Odds Ratio Preference Optimisation" and it has nothing to do with pruning. This goes way beyond native language preference.

It takes no time at all to find other major mistakes. For instance, the Mixtral diagram § 6.6.1 shows a single router that selects separate 32-layer transformers. Instead, Mixtral has one router per layer (inside of each block), and it doesn’t select a transformer block: it selects a feedforward.
Wow, so significant parts of the paper could still be LLM confabulation.
I would not say that, as long as it is a good summary there is a value in having everything in the same document.

Obviously they should have stated that this is partially generated, but at least they are dog fooding it :)

Well, it sucks that we’re at the “best practices” phase already
It sucks that we’re still at “best practices” phase. We’ve been in this phase for the last three decades [1], and I really hope we enter “good theory” phase soon.

[1] https://cseweb.ucsd.edu/classes/wi08/cse253/Handouts/lecun-9...

there's likely still an infinite amount of things to figure out, transformers haven't been out for 10 years yet
Why is that?
Ctrl-F: Unsloth --> no results == bad paper
But you can find "delve"
FYI, some sections are LLM-generated bullshit: https://x.com/sam_paech/status/1848332471953448972

The example shown in the tweet has been edited out of the paper since, but there must be others. High noise, low signal content.