about
The tide is shifting: 1.3B outperforms 7B Llama 2 (arxiv.org)
63 points by __vec__ on Sep 12, 2023 | hide | past | pdf | 15 comments on HN

In plain words: A 1.3-billion-part model trained on textbook-style text written by a bigger AI instead of web pages, to learn common-sense reasoning. It matches models five times larger on ordinary language tasks and beats most rivals outside the top tier on grade-school math and basic coding.

Abstract · Textbooks Are All You Need II: phi-1.5 technical report

We continue the investigation into the power of smaller Transformer-based language models as initiated by \textbf{TinyStories} -- a 10 million parameter model that can produce coherent English -- and the follow-up work on \textbf{phi-1}, a 1.3 billion parameter model with Python coding performance close to the state-of-the-art. The latter work proposed to use existing Large Language Models (LLMs) to generate ``textbook quality" data as a way to enhance the learning process compared to traditional web data. We follow the ``Textbooks Are All You Need" approach, focusing this time on common sense reasoning in natural language, and create a new 1.3 billion parameter model named \textbf{phi-1.5}, with performance on natural language tasks comparable to models 5x larger, and surpassing most non-frontier LLMs on more complex reasoning tasks such as grade-school mathematics and basic coding. More generally, \textbf{phi-1.5} exhibits many of the traits of much larger LLMs, both good -- such as the ability to ``think step by step" or perform some rudimentary in-context learning -- and bad, including hallucinations and the potential for toxic and biased generations -- encouragingly though, we are seeing improvement on that front thanks to the absence of web data. We open-source \textbf{phi-1.5} to promote further research on these urgent topics.

Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, Yin Tat Lee
arXiv:2309.05463 · cs.CL, cs.AI · submitted Sep 11, 2023
abstract · pdf · html

add comment on HN

> we are seeing improvement on that front thanks to the absence of web data

Bingo. This (not the parameter count) is the amazing thing to me.

Garbage in garbage out, and there is a ton of garbage in the Falcon/Llama (and OpenAI?) datasets. It feels like such a waste of compute and parameter space.

>Garbage in garbage out, and there is a ton of garbage in the Falcon/Llama (and OpenAI?) datasets. It feels like such a waste of compute and parameter space.

This is true and most researchers understand this on some level (https://arxiv.org/abs/2305.07759) but understanding it doesn't really make the problem any easier. How do you curate a general purpose model's dataset without throwing the baby out with the bathwater ?

EDIT: This paper seems to take a good stab at the question. https://arxiv.org/abs/2309.04564

> How do you curate a general purpose model's dataset without throwing the baby out with the bathwater

Time, money, and human work. The pruning process needs focused input from experts other than ML researchers and data scientists.

Maybe the Cyc team could provide some generalization tips.
>How do you curate a general purpose model's dataset without throwing the baby out with the bathwater ?

Maybe we could train an AI to do it ;)

it might not exactly be a waste, when you are inputting your training data you need some context on what the source is and then add a weight to how valid it is. That is how we learn, if i talk to a phd on some topic, I would pay close attention to what they say and give it more significance, maybe even write it down, compared to a youtube video on that topic, compared to a guy on a street talking about that same topic, ill probably nod my head and walk away and try to forget as soon as i can. but there probably still is something valid to learn from that last conversation, how a certain type of person structures sentences, responds, the words used etc...

do these llms get trained with something like a credibility weight on the training data? that was it seems they did in this paper, just manually curated that

Do people want that though? I don't want phd level responses for my queries. I want it to be better than what I could come up in a minute or by searching half an hour. Rather than some highly advanced highly detailed response I could probably not understand if the topic is not something I'm sufficient in to begin with.

Think common use cases. A lot of users are students, do I want it to write an essay like a linguist? Or solve my homework using the better but more advanced techniques and style?

It's a lot easier to add a pass to simplify an explanation by rephrasing or eliding information than it is to smarten up an overly simplified answer.

You do want the underlying model to be capable of the advanced answers, since if it is, it can be used to supply simple answers. You can't make that work the other way around in the same way.

I think they want it. What do you think the most perfect or ideal question answer-er or teacher is? It would probably be an expert in the field, but also with the ability to deliver that content in a level appropriate way for the recipient/student. Unreasonable for the most part, we get away with good enough in the real world. This is a skill we all try to learn though. Like when you need to give a technical presentation to non technical audience

If you have input data that includes a high rated reddit eli5 question, the content of that answer might be hard to verify, the style and way its delivered would be ideal to keep around in the training data. on the other side, technical in-depth answers have content that is worth keeping around, the style of its delivery would be very specific.

Keeping the entire internet around in your training data would still give you access to all these types of delivery still. hope that makes sense.

Maybe the answer is multi-step: first use curated primary sources, e.g. scientific papers. Then reinforce using well written summaries, perhaps by actual models or well graded student papers. Finally, somehow apply negative weights using wrong answers only. Bonus points if you can automate the whole process
Textbooks Are All You Need II: phi-1.5 technical report

"Perhaps achieving ChatGPT’s level of capability at the one billion parameters scale is actually achievable?"

On a long enough timeline, Id say yes.

Right now they're comparing to the 7B Llama 2, which is a shadow of 65B llama2, which is a couple steps off ChatGPT, which is a shadow of GPT4.

I'm still comfy saying yes because there's no reason to doubt it will follow the same logic as "eventually $phone will have the same FLOPS as $desktop"

Llama generally acheives much higher accuracy with very small amount of fine-tuning(on similar quality of dataset like this paper) on lot of tasks. So the model understanding is present in llama to get higher accuracy. e.g Hellaswag, ARC and MMLU for 7b model is 0.8, 0.57 and 0.52 respectively[0], while phi-1 is 0.48, 0.45 and 0.38.

I don't think finetuning phi-1 on good quality synthetic data will increase its accuracy as it is only trained on that.

[0]: https://huggingface.co/pankajmathur/orca_mini_v3_7b

The model can be downloaded here: https://huggingface.co/microsoft/phi-1_5
The title is clickbait and not the title of the paper.