about
Architectural Effects on Maximum Dependency Lengths of Recurrent Neural Networks (arxiv.org)
23 points by PaulHoule on Aug 30, 2024 | hide | past | pdf | 3 comments on HN

In plain words: A new calculation determines the longest gap a recurrent network can still connect between two parts of a sequence, no matter how long the input gets. It then measures how layer count and neuron count change that limit in plain RNNs, gated units, and LSTMs.

Abstract · A Technical Note on the Architectural Effects on Maximum Dependency Lengths of Recurrent Neural Networks

This work proposes a methodology for determining the maximum dependency length of a recurrent neural network (RNN), and then studies the effects of architectural changes, including the number and neuron count of layers, on the maximum dependency lengths of traditional RNN, gated recurrent unit (GRU), and long-short term memory (LSTM) models.

Jonathan S. Kent, Michael M. Murray
arXiv:2408.11946 · cs.NE · submitted Jul 19, 2024
abstract · pdf · html · 13 pages, 12 figures

add comment on HN

Anyone know any research combining RNN like architecture to transformers? It’d be neat if a layer of a transformer could decide to loop its output vector to a previous layer n number of times to allow it to “think harder”. The implementation would be tricky to get right though.
I've been working on using spiking neural networks (highly recurrent) to directly transform text via buffer mechanisms actuated by spiking activity. Ive had some marginal degree of success but there seems to be a bit of a wall. In these experiments, "thinking harder" takes the form of a growing pending spike queue. For practical reasons, this activity needs to be clipped or inhibited at some point. In real world biology, I suspect the power law growth in spiking activity is what contributes to much of cognition/intelligence. There's really no practical way to simulate it at full scale outside of hypothetical neuromorphic settings.

For now, I think we need to rely on probability, statistics and clever functions to drive most approaches forward. "Thinking harder" seems to require unconstrained resources if you need a decision in near real time.

https://arxiv.org/abs/2006.16236 Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention