about
Program-as-Weights: A Programming Paradigm for Fuzzy Functions (arxiv.org)
54 points by simonpure 92 days ago | hide | past | pdf | 9 comments on HN

In plain words: A big language model reads a plain-English description of a fuzzy task, like flagging log lines, and compiles it into tiny extra weights that a small local model runs offline. The result matched a larger cloud model using about one fiftieth of the memory.

Abstract

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose fuzzy-function programming: compiling such a function from a natural-language specification into a compact, locally-executable neural artifact. We instantiate this paradigm with Program-as-Weights (PAW), in which a 4B compiler trained on FuzzyBench, a 10M-example dataset we release, emits parameter-efficient adapters for a frozen, lightweight interpreter. A 0.6B Qwen3 interpreter executing PAW programs matches the performance of direct prompting of Qwen3-32B, while using roughly one fiftieth of the inference memory and running at 30 tokens/s on a MacBook M3. PAW reframes the foundation model from a per-input problem solver into a tool builder: invoked once per function definition, it produces a small reusable artifact whose subsequent calls per function application are cheap and offline.

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie, Stuart Shieber, Yuntian Deng
arXiv:2607.02512 · cs.LG, cs.AI, cs.CL · submitted Jul 2, 2026
abstract · pdf · html

add comment on HN

This looks cool, but I wonder how well their trained compiler generalizes to new task families. They trained on 29 specific types of tasks, with 800 sub tasks and many rephrasings of each one (the specs). They hold out some specs for validation, but don’t seem to have held out a full task family and maybe not even full sub tasks?

If the compiler can’t generalize well to unseen tasks then it’s effectively acting as a fancy router to one of 29/800 predefined LoRAs.

I too found this extremely interesting, jotted some notes here: https://dennisy.me/notes/programs-as-weights

Will keep noodling on this, I feel there must be ways to scale this up!

I like the goal of this. As expected, I don't really understand the math/concept of this. It sounds like it caches some neural network activity and exports it to be run later. So I suppose this can't be used for things like image or video generation.
Super interesting. I am always wondering where the bridge between symbolic programs and inline connectionist approaches could be. LLMs can call traditional programs "in process" as tools, but the other way usually just looks an external, expensive call. Maybe this fits in for a certain class of program. I will definitely add some support to my little collection of browser based AI tools https://andergrove.com/tools/ai/
this reminds me of this other flavor of "program as weights" I saw few months ago, putting deterministic programs inside the weights of an LLM

https://www.percepta.ai/blog/constructing-llm-computer

Despite the appeal of such an approach, I find this extremely unsettling.

Imagine if we had declared that Math for FIR filter design in Signal Processing was too difficult, so we’d just test random FIR coefficients until something good came out.

That sounds pretty horrible but at least the frequency response of the resulting filter would be known. We’d at least understand the behavior of the final product.

With LLMs, we don’t even know what we’re getting out of it.

(And no, I don’t see anything wrong with adaptive filters and such. Their behavior can still be quantified)

> PAW reframes the foundation model from a per-input problem solver into a tool builder: invoked once per function definition, it produces a small reusable artifact whose subsequent calls per function application are cheap and offline.

Umm you can just get the LLM to spit out real functions instead of fuzzy functions and just run those real functions through real interpreters, which is also "cheap" and "offline".

> Each is the kind of fuzzy task that resists symbolic implementation but does not need an API call to a 30B-parameter model on every input.
For super simple programs, yes, this doesn't make sense. For very complicated programs, the generated code would begin to look like spaghetti so fine that it begins to look like weights.