In plain words: Brain-like networks usually model only neurons; this design adds star-shaped helper cells that tune neuron activity at each connection, as astrocytes do in the brain. The new network remembered longer stretches and generated language better while using less memory and running faster than neuron-only versions.
Abstract · Astrocyte-Enabled Advancements in Spiking Neural Networks for Large Language Modeling
Within the complex neuroarchitecture of the brain, astrocytes play crucial roles in development, structure, and metabolism. These cells regulate neural activity through tripartite synapses, directly impacting cognitive processes such as learning and memory. Despite the growing recognition of astrocytes' significance, traditional Spiking Neural Network (SNN) models remain predominantly neuron-centric, overlooking the profound influence of astrocytes on neural dynamics. Inspired by these biological insights, we have developed an Astrocyte-Modulated Spiking Unit (AM-SU), an innovative framework that integrates neuron-astrocyte interactions into the computational paradigm, demonstrating wide applicability across various hardware platforms. Our Astrocyte-Modulated Spiking Neural Network (AstroSNN) exhibits exceptional performance in tasks involving memory retention and natural language generation, particularly in handling long-term dependencies and complex linguistic structures. The design of AstroSNN not only enhances its biological authenticity but also introduces novel computational dynamics, enabling more effective processing of complex temporal dependencies. Furthermore, AstroSNN shows low latency, high throughput, and reduced memory usage in practical applications, making it highly suitable for resource-constrained environments. By successfully integrating astrocytic dynamics into intelligent neural networks, our work narrows the gap between biological plausibility and neural modeling, laying the groundwork for future biologically-inspired neural computing research that includes both neurons and astrocytes.
Guobin Shen, Dongcheng Zhao, Yiting Dong, Yang Li, Jindong Li, Kang Sun, Yi Zeng
arXiv:2312.07625 · cs.NE, cs.AI · submitted Dec 12, 2023 · updated Dec 26, 2023
abstract · pdf · html
"Building on Kocon et al.’s insights [37], we explored the zero-shot learning capabilities of the scaled-up AM-SNet model with 1.5 billion parameters, trained on The Pile dataset [35]. Zero-shot learning tasks, which test the model’s generalization to unseen tasks without specific training, are crucial for language models’ practical application."
...but a good chunk of the text generation examples begin with "As an AI language model..."
I may be mistaken: I thought "as an AI language model" was OpenAI's design, and they deliberately instilled this specific phrase into GPT via RLHF. It's not something LLMs say "naturally," and The Pile by itself shouldn't make an LLM preprend certain answers with "as an AI language model." The paper does not say they did any RLHF themselves, or even any prompt engineering about making the LLM respond like an AI chatbot. LLMs don't "know" they're AI unless they've been prompted to that effect.
[Edit: I missed this in Section 3, "Subsequent fine-tuning was conducted on the MOSS finetune dataset, which is tailored for dialogue systems." See causalmodels's helpful comment below. What I wrote below is wrong yet not totally off-base. The research wasn't "clumsily faked," but the post-training conditioning was certainly corrupted with GPT-3.5 output, which IMO makes the natural language generation analysis quite flawed - and in general should make us wonder how "zero-shot" these examples really are.]
I am suspicious this part of the research was clumsily faked. Happy to be proven wrong: maybe there's something in The Pile that explains it, or part of the modern LLM toolchain does some pre-training to make things more ChatGPT-like. Maybe the researchers actually trained it on GPT-4 and didn't mention that in the paper.