about
Astrocyte-Enabled Spiking Neural Networks for Large Language Modeling (arxiv.org)
27 points by PaulHoule on Dec 21, 2023 | hide | past | pdf | 12 comments on HN

In plain words: Brain-like networks usually model only neurons; this design adds star-shaped helper cells that tune neuron activity at each connection, as astrocytes do in the brain. The new network remembered longer stretches and generated language better while using less memory and running faster than neuron-only versions.

Abstract · Astrocyte-Enabled Advancements in Spiking Neural Networks for Large Language Modeling

Within the complex neuroarchitecture of the brain, astrocytes play crucial roles in development, structure, and metabolism. These cells regulate neural activity through tripartite synapses, directly impacting cognitive processes such as learning and memory. Despite the growing recognition of astrocytes' significance, traditional Spiking Neural Network (SNN) models remain predominantly neuron-centric, overlooking the profound influence of astrocytes on neural dynamics. Inspired by these biological insights, we have developed an Astrocyte-Modulated Spiking Unit (AM-SU), an innovative framework that integrates neuron-astrocyte interactions into the computational paradigm, demonstrating wide applicability across various hardware platforms. Our Astrocyte-Modulated Spiking Neural Network (AstroSNN) exhibits exceptional performance in tasks involving memory retention and natural language generation, particularly in handling long-term dependencies and complex linguistic structures. The design of AstroSNN not only enhances its biological authenticity but also introduces novel computational dynamics, enabling more effective processing of complex temporal dependencies. Furthermore, AstroSNN shows low latency, high throughput, and reduced memory usage in practical applications, making it highly suitable for resource-constrained environments. By successfully integrating astrocytic dynamics into intelligent neural networks, our work narrows the gap between biological plausibility and neural modeling, laying the groundwork for future biologically-inspired neural computing research that includes both neurons and astrocytes.

Guobin Shen, Dongcheng Zhao, Yiting Dong, Yang Li, Jindong Li, Kang Sun, Yi Zeng
arXiv:2312.07625 · cs.NE, cs.AI · submitted Dec 12, 2023 · updated Dec 26, 2023
abstract · pdf · html

add comment on HN

I don't actually do LLM research or engineering myself - I hate Python :) So I am ignorant about a lot of the nitty-gritty in actually training and deploying an LLM. There's a lot in this paper I haven't read closely. But this struck me as suspicious:

"Building on Kocon et al.’s insights [37], we explored the zero-shot learning capabilities of the scaled-up AM-SNet model with 1.5 billion parameters, trained on The Pile dataset [35]. Zero-shot learning tasks, which test the model’s generalization to unseen tasks without specific training, are crucial for language models’ practical application."

...but a good chunk of the text generation examples begin with "As an AI language model..."

I may be mistaken: I thought "as an AI language model" was OpenAI's design, and they deliberately instilled this specific phrase into GPT via RLHF. It's not something LLMs say "naturally," and The Pile by itself shouldn't make an LLM preprend certain answers with "as an AI language model." The paper does not say they did any RLHF themselves, or even any prompt engineering about making the LLM respond like an AI chatbot. LLMs don't "know" they're AI unless they've been prompted to that effect.

[Edit: I missed this in Section 3, "Subsequent fine-tuning was conducted on the MOSS finetune dataset, which is tailored for dialogue systems." See causalmodels's helpful comment below. What I wrote below is wrong yet not totally off-base. The research wasn't "clumsily faked," but the post-training conditioning was certainly corrupted with GPT-3.5 output, which IMO makes the natural language generation analysis quite flawed - and in general should make us wonder how "zero-shot" these examples really are.]

I am suspicious this part of the research was clumsily faked. Happy to be proven wrong: maybe there's something in The Pile that explains it, or part of the modern LLM toolchain does some pre-training to make things more ChatGPT-like. Maybe the researchers actually trained it on GPT-4 and didn't mention that in the paper.

> I thought "as an AI language model" was OpenAI's design, and they deliberately instilled this specific phrase into GPT via RLHF. It's not something LLMs say "naturally," and The Pile by itself shouldn't make an LLM preprend certain answers with "as an AI language model." The paper does not say they did any RLHF themselves, or even any prompt engineering about making the LLM respond like an AI chatbot. LLMs don't "know" they're AI unless they've been prompted to that effect.

This is correct. OpenAI's GPT models say this because they've been trained to do so via RLHF. Newer language models typically say this because their training set is partially generated by OpenAI's models. It's quite fascinating to ask LLaMA finetunes about themselves and see which ones respond with "As an AI language model..." or even outright say they were created by OpenAI.

Yes, see causalmodel's comment. I missed that the researchers trained the AI on The Pile but did some post-conditioning on MOSS, which contains GPT-3.5 output.

To be clear: it's fine for commercial / personal / etc LLMs to train on GPT. But it's not so fine when research LLMs do it: those benchmarks don't mean very much if you're training the AI on GPT's answers to the benchmarks. Test contamination is bad enough in LLM research; using datasets like MOSS in academic papers is just reckless.

This got me curious, so I looked into it. The only post ChatGPT dataset was MOSS-3-SFT, which was used for dialogue fine tuning. MOSS-3 was generated using ChatGPT-3.5-turbo [1]. They explicitly call out their fine tuning data in footnote 13 so I wouldn't say it is faked per say, but it isn't a great look.

[1] https://huggingface.co/fnlp/moss-moon-003-base#data

Got it, thanks for looking into it more closely. I didn't actually finish reading the paper, and missed that this was in section 3: "Subsequent fine-tuning was conducted on the MOSS finetune dataset [36], which is tailored for dialogue systems."

Agreed this almost certainly isn't faked, but the validity of using this dataset at all seems questionable.

> I don't actually do LLM research or engineering myself - I hate Python :)

What programming language(s) do you like?

Idris, Coq, F#/OCaml, Scheme. C# is also nice, and has the significant advantage that people will pay me to write it....

I was being a bit facetious, Python is perfectly fine for writing programs and software. The problem is the AI/ML dialect of Python. I don't like working in domains that involve big 3rd-party ecosystems and libraries[1]. I prefer doing something bespoke for a smaller organization, even if that means solving a smaller problem.

[1] Interacting with tediously-but-necessarily detailed libraries and programming frameworks is an area where ChatGPT is genuinely useful - codegen which boils down to "translate this mathematical English into PyTorch." So at some point I might go back and actually program some of this stuff.

Do you have recommendations for getting into ML-family languages and functional approaches, or even just F#? Although Python's ecosystem and plentiful training data for LLMs are useful, trying Elm made me want to find something better at types.
As a former neuroscientist, I welcome attempts to add more biologically-inspired complexity to the relatively simple computer-based NNs.

That being said, I suspect the authors still need to take a Neuroscience 101 class. There's any number of neuronal features missing from NNs that would be more interesting to incorporate before looking towards glial cells like astrocytes. E.g., different neurotransmitters, G proteins, gap junctions/electrical synapses, pyramidal neurons, really all large-scale and long-range organizational structures in brain.

To my understanding, pyramidal cells are already partially emulated by certain structures of weights. It's not the same as dedicated multi-layer mechanisms outside the local flow, but that option seems difficult to implement efficiently with current compute technology.

Maybe more VRAM to hold the added inter-layer features would allow it? NVidia seems reluctant or unable to do that on the consumer side. I haven't had time to thoroughly read up on why. All of these seem plausible:

* Offsetting rising per-unit costs from other components

* Intentional market segmentation

* Export restrictions

I suspect cost/efficiency is a driving factor. Existing layered NN designs seem quite local, with fewer long-range connections than brains have. I'd bet that updating non-local weights increases memory requirements, slows down computation by orders of magnitude, or both.
If there's an architecture that could be 10% better, I think it's highly unlikely that we would conduct ablation experiments, since it's too expensive, It's really a pity compared to CV.