about
Replacing Attention to Phase-Locking to Overcome Catastrophic Forgetting [pdf] (arxiv.org)
1 point by Yujivus 304 days ago | hide | past | pdf | 1 comment on HN

In plain words: Signals are waves with fixed strength, so meaning lives in their timing and waves cancel noise rather than the biggest signals winning. Mixed with normal attention, it beat baselines with fewer parameters, and scrambling the timing wrecked performance while leaving it intact barely hurt.

Abstract · Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks

In standard Transformer architectures, semantic importance is often conflated with activation magnitude, obscuring the geometric structure of latent representations. To disentangle these factors, we introduce PRISM, a complex-valued architecture designed to isolate the computational role of phase. By enforcing a strict unit-norm constraint ($|z| = 1$) and replacing attention with gated harmonic convolutions, the model is encouraged to utilize subtractive interference in the frequency domain to suppress noise, rather than relying on magnitude-based gating. We utilize this constrained regime to study a hybrid architecture -- fusing phase-based routing with standard attention -- which achieves improved parameter efficiency and representation quality compared to baselines in our evaluated settings. Mechanistically, interventional ablations indicate that the model carries substantial task-relevant information in phase: preserving phase largely maintains performance, whereas disrupting phase causes severe degradation. Together, these results suggest that phase-based spectral interference is a usable computational mechanism for neural sequence modeling at the evaluated scale.

Alper Yıldırım, İbrahim Yücedağ
arXiv:2512.01208 · cs.LG, cs.AI, cs.CL · submitted Dec 1, 2025 · updated Jul 16, 2026
abstract · pdf · html · ICML 2026

add comment on HN

Author here. I have been working on catastrophic forgetting on transformers. And I found a potential solution with strong results. I have a weird attention free encoder that treats embeddings as waves instead of vectors. Motivation and summary of solution is below.

Pretrained embeddings starts to learn faster, this is a known thing and used for low resource NLP. So, if we could scructure the embedding "map" faster, decoder should learn a lot faster.

So, we isolated the alignment cost of this map with a method we call ISMR. We found 20 layered model with 14.5% embeddings does not learn faster than 1 layered model with 80% embeddings.

Then, we invented a weird encoder we call "PRISM". It treats embeddings as waves instead of vectors. It teleports embeddings to relevant frequencies rapidly. It looks like it can learn new concepts 5-shot with nearly no forgetting (-0.7 to -0.84 BLEU) while standard transformer encoder decoder suffers catastrophic forgetting (more than 10 BLEU loss).