about
WaveCoder: Enhanced instruction tuning with refined data generation (arxiv.org)
27 points by tosh on Jan 17, 2024 | hide | past | pdf | 10 comments on HN

In plain words: A tool turns open-source code into varied, high-quality training examples across four kinds of coding tasks, so a model learns more than just writing new code. Models trained on these 19,915 examples beat other open models on a wider range of coding work.

Abstract · WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Recent work demonstrates that, after instruction tuning, Code Large Language Models (Code LLMs) can obtain impressive capabilities to address a wide range of code-related tasks. However, current instruction tuning methods for Code LLMs mainly focus on the traditional code generation task, resulting in poor performance in complex multi-task scenarios. In this paper, we concentrate on multiple code-related tasks and present WaveCoder, a series of Code LLMs trained with Widespread And Versatile Enhanced instruction data. To enable the models to tackle complex code-related tasks, we propose a method to stably generate diverse, high-quality instruction data from open source code dataset in multi-task scenarios and obtain CodeSeaXDataset, a dataset comprising 19,915 instruction instances across 4 code-related tasks, which is aimed at improving the generalization ability of Code LLM. Our experiments demonstrate that WaveCoder models significantly outperform other open-source models in terms of the generalization ability across different code-related tasks. Moreover, WaveCoder-Ultra-6.7B presents the state-of-the-art generalization abilities on a wide range of code-related tasks.

Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, Qiufeng Yin
arXiv:2312.14187 · cs.CL, cs.AI, cs.SE · submitted Dec 20, 2023 · updated Jun 7, 2024
abstract · pdf · html

add comment on HN

is synthetic data a really big deal right now and LLM? if so, are there any take-home ideas that might apply to other areas, say analysis of MRI?
Yes and no. In terms of LLMs, it's basically figuring out how to exfiltrate information from GPT4 to remove costs of data gathering. The limitations of that are that the model will never be better than gpt4, and when gpt4 produces incorrect information, the model trained on synthetic data will also do so.

In other fields like computer vision, synthetic data is useful for generating ground truth data, like for depth masks.

This is mistaken. You use GPT-4 to generate new data using other data sources, for example text books.

Asking GPT-4 to create 50 new conversations from a chapter of a textbook creates higher quality data than most of “the pile” and can extend beyond what GPT-4 has in its existing dataset. This is exactly what MS has done with Phi.

I’ve disputed that fact before. It depends what part of your data is generated from GPT4 if I have high quality code and I’m using gpt4 to synthetically generate variations on how some could ask for the code to be written it’s entirely possible because of the high quality code for the model to be better. It’s not all or nothing if you’re mixing synthetic data with quality sources.

In this case even though parts of the dataset are synthetic the bound is on the code not necessarily the 50 ways I got gpt4 to say “write me a script to do x” or modeled other interactions with that code data source.

If we’re talking about model distillation[0] I don’t think the student can ever be better than the teacher as optimising for speed and smaller model sizes inherently means that there will be precision loss. Even if the student is as big as the teacher, there is still data loss.

[0] https://arxiv.org/pdf/2210.17332.pdf

Synthetic data is a big deal, essentially as a form of “knowledge distillation” from large models or for transforming high-quality text into training data (e.g. Q&A pairs). Almost everyone is using GPT-4 for this. Dunno about other domains, as it’s based on the mutability of text, relative to whatever ground truths are embedded therein. This seems less feasible for other kinds of inputs, but who knows.
I think the critical thing is you need some ground truth way of evaluating the synthetic data. You can generate 100 programs with your LLM and filter to the 1-2 that solve the problem, but there's not an equivalent option for things like MRI.
A self-debiasing estimator might become unreliable, and brains think that matters?
Did anyone find the source code yet?
They said on Twitter that they’re still conferring with Microsoft internally on the extent and nature of the open-source release:

https://nitter.net/TeamCodeLLM_AI/status/1747652471714144702