about
Synthetic Data Almost from Scratch (arxiv.org)
2 points by milliondreams on Mar 3, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Instead of copying from existing examples, this method builds a tree of human knowledge—fields, subjects, then class-by-class syllabi—and uses it to write training questions covering every discipline. The resulting model beat models trained on task-specific data across math, coding, exams, logic, and general instruction following.

Abstract · Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

We introduce Generalized Instruction Tuning (called GLAN), a general and scalable method for instruction tuning of Large Language Models (LLMs). Unlike prior work that relies on seed examples or existing datasets to construct instruction tuning data, GLAN exclusively utilizes a pre-curated taxonomy of human knowledge and capabilities as input and generates large-scale synthetic instruction data across all disciplines. Specifically, inspired by the systematic structure in human education system, we build the taxonomy by decomposing human knowledge and capabilities to various fields, sub-fields and ultimately, distinct disciplines semi-automatically, facilitated by LLMs. Subsequently, we generate a comprehensive list of subjects for every discipline and proceed to design a syllabus tailored to each subject, again utilizing LLMs. With the fine-grained key concepts detailed in every class session of the syllabus, we are able to generate diverse instructions with a broad coverage across the entire spectrum of human knowledge and skills. Extensive experiments on large language models (e.g., Mistral) demonstrate that GLAN excels in multiple dimensions from mathematical reasoning, coding, academic exams, logical reasoning to general instruction following without using task-specific training data of these tasks. In addition, GLAN allows for easy customization and new fields or skills can be added by simply incorporating a new node into our taxonomy.

Haoran Li, Qingxiu Dong, Zhengyang Tang, Chaojun Wang, Xingxing Zhang, Haoyang Huang, Shaohan Huang, Xiaolong Huang, Zeqiang Huang, Dongdong Zhang, Yuxian Gu, Xin Cheng, et al.
arXiv:2402.13064 · cs.CL · submitted Feb 20, 2024
abstract · pdf · html · Work in progress

add comment on HN

An interesting discussion around creating synthetic data with very little starting information. It introduces a smart way to build diverse datasets using something called taxonomies. This approach is intriguing and points towards new directions in AI development.

But, it also highlights some big challenges we need to think about. The richness of the English language is part of what makes it so successful, allowing for a wide range of expression. However, there's a growing trend towards making synthetic data more uniform, not taking into account this diversity.

This raises a crucial question: how will this uniformity affect the quality and variety of online content? Nowadays, there's already a lot of content online created by big AI models, making the internet feel more and more the same.

In this rush, major players in AI research—like OpenAI , Google , and Microsoft —are focusing more on turning AI models into new types of search engines. This shift could mean we're missing out on addressing the real challenges in creating really intelligent systems. It makes you wonder if we're even measuring AI success correctly.

With so much AI-created content out there, it's essential to think about new ways to push AI research forward. So, who's really breaking new ground in building smarter AI models? Who's tackling the important challenges that will shape the future of AI?