about
Hyperparameter Optimization for LLMs via Scaling Laws (arxiv.org)
84 points by Lindizz on Jun 6, 2023 | hide | past | pdf | 12 comments on HN

In plain words: A group of neural networks predicts each trial's learning curve as a power law—accuracy rising predictably with training—so the tuner can pause weak trials and keep training promising ones. Across 59 tasks in three benchmarks it beat seven rivals throughout.

Abstract · Scaling Laws for Hyperparameter Optimization

Hyperparameter optimization is an important subfield of machine learning that focuses on tuning the hyperparameters of a chosen algorithm to achieve peak performance. Recently, there has been a stream of methods that tackle the issue of hyperparameter optimization, however, most of the methods do not exploit the dominant power law nature of learning curves for Bayesian optimization. In this work, we propose Deep Power Laws (DPL), an ensemble of neural network models conditioned to yield predictions that follow a power-law scaling pattern. Our method dynamically decides which configurations to pause and train incrementally by making use of gray-box evaluations. We compare our method against 7 state-of-the-art competitors on 3 benchmarks related to tabular, image, and NLP datasets covering 59 diverse tasks. Our method achieves the best results across all benchmarks by obtaining the best any-time results compared to all competitors.

Arlind Kadra, Maciej Janowski, Martin Wistuba, Josif Grabocka
arXiv:2302.00441 · cs.LG · submitted Feb 1, 2023 · updated Oct 25, 2023
abstract · pdf · html · Accepted at NeurIPS 2023

add comment on HN

What’s better, train a model with 10X parameters once on some default hyperparameter setting or to search for a good hyperparameter configuration by training on X parameters 10 times? While I’m at it, how many LLMs of the size of GPT3 were trained until they landed on the capability of GPT3? How much of this is dependent on the data, or do good settings transcend the type of text that a model is trying to train on?
The proposed idea here is different.

You can train several smaller models with different hyperparameters with dynamic budgets, i.e. bad configurations are trained for only few epochs, and good ones for more epochs. Once you find a good hyperparameter configuration for the small-scale model, then you train the large model with that configuration.

What is being shown is that the overhead of doing hyperparameter optimization at a small scale, is comparable to a single optimization at the largest scale.

Overall, the idea looks very cool.

30 and 40B parameter models regularly crush GPT-3 175B (Davinci) on every benchmark. GPT-3.5 is probably a 13B parameter model and it beats the original GPT-3 175B on most benchmarks (but not the more recent finetunes like 175B Davinci-003), so hyperparameters are clearly very important.
How are hyperparameters tuned for GPT3.5, is there any leak on the method they use?
The questions you raise are very interesting. My question would be, where does the default hyperparameter configuration come from? Additionally, does there exist one hyperparameter configuration that performs well on all tasks?
Very interesting. Wondering what is the state of the art in Hyperparameter Optimization at the moment. Does this method apply to all Deep Learning systems?
For a general overview, this could be a good starting point [1]. As for deep learning, you may wanna start from here [2], but I personally had good results with Hyperband [3] for DL.

[1] https://wires.onlinelibrary.wiley.com/doi/full/10.1002/widm....

[2] https://github.com/google-research/tuning_playbook

[3] https://jmlr.csail.mit.edu/papers/v18/16-558.html

Hyperband [1] has been my go-to hyperparam optimization method over the past few years. Handily beats Bayesian search wherever I applied it, also implemented in most frameworks.

1. https://arxiv.org/abs/1603.06560

The work compares against Hyperband and the new method is significantly better (Figure 2, Hypothesis 2).
I just got back into hyperopt a couple weeks ago. It's easy enough and worked for me, but I was thinking there had to be some new things I'm not aware of.
I think so, they apply it to Computer Vision datasets as well
Really nice, it is very useful now with so many new models and datasets for NLP