about
ChipNeMo: Domain-Adapted LLMs for Chip Design (arxiv.org)
50 points by RafelMri on Nov 7, 2023 | hide | past | pdf | 7 comments on HN

In plain words: A general language model was fed chip-design text and taught the field's vocabulary and instructions, turning it into a helper for chip engineers. Its largest version beat GPT-4 at answering engineering questions and writing chip-tool scripts, without losing its general skills.

Abstract

ChipNeMo aims to explore the applications of large language models (LLMs) for industrial chip design. Instead of directly deploying off-the-shelf commercial or open-source LLMs, we instead adopt the following domain adaptation techniques: domain-adaptive tokenization, domain-adaptive continued pretraining, model alignment with domain-specific instructions, and domain-adapted retrieval models. We evaluate these methods on three selected LLM applications for chip design: an engineering assistant chatbot, EDA script generation, and bug summarization and analysis. Our evaluations demonstrate that domain-adaptive pretraining of language models, can lead to superior performance in domain related downstream tasks compared to their base LLaMA2 counterparts, without degradations in generic capabilities. In particular, our largest model, ChipNeMo-70B, outperforms the highly capable GPT-4 on two of our use cases, namely engineering assistant chatbot and EDA scripts generation, while exhibiting competitive performance on bug summarization and analysis. These results underscore the potential of domain-specific customization for enhancing the effectiveness of large language models in specialized applications.

Mingjie Liu, Teodor-Dumitru Ene, Robert Kirby, Chris Cheng, Nathaniel Pinckney, Rongjian Liang, Jonah Alben, Himyanshu Anand, Sanmitra Banerjee, Ismet Bayraktaroglu, Bonita Bhaskaran, Bryan Catanzaro, et al.
arXiv:2311.00176 · cs.CL · submitted Oct 31, 2023 · updated Apr 4, 2024
abstract · pdf · html · Updated results for ChipNeMo-70B model

add comment on HN

The work is solid, the approach of domain-specific tokenizer is definitely effective. But the topic seems to be somehow misleading, as I expect the researchers to somehow find a way to connect LLMs for direct design of chips, while this paper focuses on three topics: 1. Engineering assistant chatbot, 2. EDA tool script generation, and 3. Bug summarization and analysis. Where (1) is mostly irrelavent with the core of "design", and only (2) is directly related with a useful workflow. What's more, it seems that they are only generating at most 10 lines of Python/Tcl code (more line "commands" in this case) for the EDA tools.
Their paper from ~6 weeks earlier is on actual LLM-designed chips, though "just" from an LLM benchmark POV: https://research.nvidia.com/index.php/publication/2023-09_ve...
It looks like availability of good quality training sets will be a stumbling block for LLM use in Verilog chip design since pretraining with other programming language corpus is not transferable. (see below for quote from Nvidia paper) A lot of high quality Verilog is locked in licensed, close source IP blocks covered by NDAs.

HDLBits problems are a toy level complexity circuits suitable for Verilog 101 course material.

I would set a benchmark for serious HDL design LLM at reaching ability to implement AXI bus components with specified by user functionality, e.g. AXI4 Slave (address, data widths, burst capability) with memory implemented as banked synchronous SRAM.

https://arxiv.org/pdf/2309.07544.pdf

* Despite the fact that multi models undergo pretraining on an extensive corpus of multi-lingual code data, they exhibit only marginal enhancements of approximately 3% when applied to Verilog coding task. This observation potentially suggests that there is limited positive knowledge transfer between software programming languages like C++ and hardware descriptive languages such as Verilog. This highlights the significance of pretraining on substantial Verilog corpora, as it can significantly enhance model performance in Verilog-related tasks *

I love the phrase "limited positive knowledge transfer between software programming languages like C++ and hardware descriptive languages such as Verilog."

Maybe someday the LLM's will just read the spec and come up with new designs for which they have no examples!

Ooh, this is what I've been looking for!

But I agree with pbazarnik that their benchmark are still tackling with toy-level problems. Bits operations, Karnaugh map and other stuff are just one-level of transformation logic, really want to see more in-depth examples such as high-speed circuit optimization (say a clk over 600MHz), complex filters and PID controls, etc. Wouldn't hurt even if they only achieve 20% correctness on these problems, as long as the approach can work once every few times.

Here’s an idea: to specialize an LLM for a specific field, invent a programming language for that field and fine-tune/restrict the LLM to that programming language.
That's called DSL(domain-specific language). Domains that have their DSLs will largely benefit as DSLs can work as an LLM interface. Even LangChain is inventing their own DSL (LCEL).