LSTM
An LSTM adds a separate memory line to a recurrent network, with taps that decide what to write, what to keep and what to read out, so information can survive hundreds of steps.
- 24 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An LSTM is a recurrent network with a separate memory line running through it. Three learned taps control that line: what gets written in, what stays, and what gets read out.
LSTM stands for long short-term memory — a short-term memory that has been made to last a long time.
The list in your pocket
Walk into a market with a paper shopping list in your pocket. Three separate things happen with it, and never at the same moment.
You add an item when someone messages you to pick up milk. You cross off an item once it is in the bag. And you read from it only when the shopkeeper is looking at you, not the whole time you are walking around.
Writing, crossing off and reading are three different decisions, and each one has its own moment. The list itself sits in your pocket untouched in between.
That is an LSTM. The list is the memory. The three decisions are the three taps, and they are what makes it different from the network in the previous lesson.
Why it exists
Recurrent neural networks hold a running memory and update it at every step. The lesson before this one measured what goes wrong. Over forty steps, the influence of the first word on the last decision shrank to about a millionth of a millionth.
The reason is that the memory is rewritten from scratch at every single step. Nothing is ever set aside. The shopkeeper holding a running total in his head has one number, and every item changes it.
Notice what the paper list does differently. When you are told to buy milk, you write "milk" down. Then nothing happens to that word while you walk the next four aisles. It is not averaged with anything. It is not squashed. It waits.
An LSTM copies that. It keeps a memory line that is left alone by default, and it changes only where a tap opens.
The four parts
The cell state is the list. A row of numbers that runs straight through the network from step to step, largely untouched.
The forget tap decides what to rub out. Wide open means keep everything, and shut means wipe it.
The input tap decides what to write in from the current step. Shut means this step contributes nothing to the memory.
The output tap decides how much of the memory to say out loud right now. The memory can hold something for two hundred steps and reveal it only at the end.
Every tap is a number between zero and one, for each slot of memory. The network learns when to open each one. Nobody programs the rules.
How it works
memory in ──────►[ forget tap ]────►( + )──────────────────► memory out
(the list) ▲ ▲ │
│ │ │
│ [ input tap ] │
│ ▲ ▼
│ │ [ output tap ]
this step's word ──────┴──────────────┴──────────────────────────┤
▼
what gets passed onThe important thing in that picture is the top line. The memory travels left to right and meets one multiplication and one addition. It is not passed through a squashing function every step, and it is not multiplied by a weight matrix every step.
That straight line across the top is the whole invention. A plain recurrent network has no such line. That is why its memory decays.
Where you have already seen this
- Voice typing on your phone, transcribing you while you are still speaking.
- Swipe and next-word prediction on a phone keyboard, running with no internet.
- Google Translate between 2016 and 2018, before transformers replaced the engine.
- Handwriting recognition, where a stroke's meaning depends on strokes made earlier.
- Machine health monitoring, watching a stream of vibration readings for a fault.
The honest part: transformers took most of this work
This matters more than any praise, so it goes in plain words.
For nearly all language work, LSTMs have been replaced by transformers. Every large language model you have used is a transformer. Translation moved across around 2017 and 2018. Nobody starts a new text project with an LSTM today.
The reason is not accuracy. It is training speed.
An LSTM reads step one, then step two, then step three. Step 500 cannot start before step 499 finishes. A transformer looks at all the positions at once, so a graphics card can work on the whole sequence in parallel. That difference decides how large a model you can afford to train, and scale turned out to matter enormously.
There is a second reason. An LSTM squeezes everything it has read into one fixed-size memory. A transformer keeps every earlier position available and looks directly at whichever one it needs. For "what was the name mentioned in paragraph two", direct access wins outright.
So where is an LSTM still the right answer?
- Live streams that never end. Speech recognition running as you speak, sensor readings, a live trading feed. An LSTM does the same amount of work at step 5 and at step 50,000. A transformer's cost per step grows as the context grows.
- Small devices. An LSTM's memory is one fixed-size row of numbers. A transformer must keep a growing store of every past position. On a phone, a hearing aid or a microcontroller, that difference decides whether the model fits.
- Forecasting a number from a modest history — electricity demand, stock levels, footfall. On these, well-tuned LSTMs remain competitive, and often the dataset is far too small to train a transformer well.
- Anywhere latency per step is the constraint rather than total throughput.
Speech recognition is the clearest live example. The standard on-device design still uses LSTM layers, because the model must produce words while the person is still talking.
Learn it, and know when not to reach for it. The gating idea also shows up inside newer architectures, so this is not a museum piece.
Remember this
- An LSTM adds a memory line that passes through with almost nothing done to it.
- Three learned taps decide what to write, what to keep and what to read out.
- Transformers replaced it for language, and it still wins on live streams and small devices.
What to learn next
- Transformers — the architecture that replaced this one for language, and why.
- Recurrent neural networks — the simpler cell, and the failure this lesson repairs.
- PyTorch basics — the training loop these examples assume.
Developer — Code and libraries.
Setup
pip install torch numpyEverything runs on CPU in seconds. No GPU, no downloads.
The taps, with no learning involved
Set the taps by hand and watch the memory. There are no trained weights here at all — this is the mechanism on its own.
import numpy as np
c = 0.0 # the cell state: the note in your pocket
h = 0.0 # the hidden state: what gets passed onward
# (label, this step's input, forget tap, input tap, output tap)
steps = [
("0.90 arrives", 0.90, 1.0, 1.0, 1.0),
("nothing", 0.00, 1.0, 0.0, 1.0),
("nothing", 0.00, 1.0, 0.0, 1.0),
("nothing", 0.00, 1.0, 0.0, 1.0),
("nothing", 0.00, 1.0, 0.0, 1.0),
("stop speaking", 0.00, 1.0, 0.0, 0.0), # keep the note, say nothing
("nothing", 0.00, 1.0, 0.0, 1.0),
("clear it", 0.00, 0.0, 0.0, 1.0), # forget tap shut: the note is erased
("nothing", 0.00, 1.0, 0.0, 1.0),
]
print("step event f i o cell c output h")
for n, (label, x, f, i, o) in enumerate(steps, 1):
g = np.tanh(x) # the value this step proposes writing
c = f * c + i * g # keep some of the old note, add some of the new
h = o * np.tanh(c)
print(f"{n:4d} {label:14s} {f:.1f} {i:.1f} {o:.1f} {c:+.4f} {h:+.4f}")step event f i o cell c output h 1 0.90 arrives 1.0 1.0 1.0 +0.7163 +0.6146 2 nothing 1.0 0.0 1.0 +0.7163 +0.6146 3 nothing 1.0 0.0 1.0 +0.7163 +0.6146 4 nothing 1.0 0.0 1.0 +0.7163 +0.6146 5 nothing 1.0 0.0 1.0 +0.7163 +0.6146 6 stop speaking 1.0 0.0 0.0 +0.7163 +0.0000 7 nothing 1.0 0.0 1.0 +0.7163 +0.6146 8 clear it 0.0 0.0 1.0 +0.0000 +0.0000 9 nothing 1.0 0.0 1.0 +0.0000 +0.0000
Read the cell c column downwards. It reads +0.7163 for seven consecutive steps, to the last decimal place. Not 0.716, then 0.712, then 0.708. The exact same number.
Compare that with the plain recurrent cell in the RNN lesson, where h changed at every single step because tanh was applied to it every time.
Now look at step 6, which is the part people miss. The output h fell to zero while c stayed at +0.7163. The memory was held and deliberately not spoken. At step 7 it comes straight back.
A plain recurrent network cannot do that. Its memory and its output are the same thing, so staying quiet means forgetting.
Step 8 shuts the forget tap and the memory is gone. Erasing is a decision the network makes, not something that happens to it by decay.
The measurement that justifies the whole design
The RNN lesson measured how much the first step of a sequence influences the last output, and watched it collapse. Here is the same measurement with an LSTM beside it.
import torch
import torch.nn as nn
rnn = nn.RNN(1, 1, batch_first=True)
with torch.no_grad(): # the same hand-set weights as the RNN lesson
rnn.weight_ih_l0.fill_(1.0)
rnn.weight_hh_l0.fill_(0.5)
rnn.bias_ih_l0.zero_()
rnn.bias_hh_l0.zero_()
lstm = nn.LSTM(1, 1, batch_first=True)
with torch.no_grad():
# PyTorch stacks the four gates in one matrix, in the order input, forget, candidate, output.
lstm.weight_ih_l0.copy_(torch.tensor([[0.], [0.], [1.], [0.]]))
lstm.weight_hh_l0.zero_()
lstm.bias_ih_l0.zero_()
lstm.bias_hh_l0.copy_(torch.tensor([0., 5., 0., 5.])) # forget bias 5 -> tap open at 0.993
def influence_of_first_step(model, length):
x = torch.zeros(1, length, 1, requires_grad=True)
out = model(x)[0]
out[0, -1, 0].backward() # how much does step 1 move the final output?
return x.grad[0, 0, 0].item()
print("length plain RNN LSTM")
for T in (5, 10, 20, 40, 100, 200):
print(f"{T:6d} {influence_of_first_step(rnn, T):.3e} {influence_of_first_step(lstm, T):.3e}")length plain RNN LSTM
5 6.250e-02 4.835e-01
10 1.953e-03 4.675e-01
20 1.907e-06 4.372e-01
40 1.819e-12 3.822e-01
100 1.578e-30 2.555e-01
200 0.000e+00 1.305e-01This table is the entire argument for LSTMs, in twelve numbers.
At 200 steps the plain RNN prints 0.000e+00. That is not rounding for display. The value underflowed to exactly zero in 32-bit floating point. There is no gradient at all, so training can never correct anything about step 1.
The LSTM at 200 steps still returns 0.1305, roughly a quarter of what it had at length 5.
Be honest about what this does and does not show. The LSTM's influence still decays, from 0.4835 to 0.1305. An LSTM does not give infinite memory. It changes the decay from a cliff into a slope, and a slope is something gradient descent can work with.
The reason sits in the setting forget bias 5, which makes the forget tap 0.993. The memory is multiplied by 0.993 each step instead of being crushed through a squashing function. Over 199 steps, 0.993 raised to that power is about 0.25, which is what the table shows.
The one-line change that decides whether it trains
Now train something. A sequence of 30 steps: the answer is announced at step 1, and the remaining 29 steps are filler. The model must carry the answer to the end.
import torch
import torch.nn as nn
T = 30 # the answer is announced at step 1 and must survive 29 more steps
H = 20 # size of the memory
def make_batch(n, gen):
x = torch.zeros(n, T, 3)
y = torch.randint(0, 2, (n, 1), generator=gen).float()
x[:, :, 2] = 1.0 # channel 2 marks "this step is filler"
x[:, 0, 2] = 0.0
x[:, 0, 0] = 1.0 - y[:, 0] # channel 0 = "the answer is no"
x[:, 0, 1] = y[:, 0] # channel 1 = "the answer is yes"
return x, y
class Model(nn.Module):
def __init__(self, cell, forget_bias=None):
super().__init__()
self.rnn = cell(3, H, batch_first=True)
self.out = nn.Linear(H, 1)
if forget_bias is not None:
with torch.no_grad():
# PyTorch stacks the gates as [input, forget, candidate, output].
self.rnn.bias_ih_l0[H:2 * H].fill_(forget_bias)
def forward(self, x):
seq, _ = self.rnn(x) # nn.LSTM returns (output, (h, c)); we want the output
return self.out(seq[:, -1]) # decide from the final step alone
gen_test = torch.Generator().manual_seed(99)
x_test, y_test = make_batch(500, gen_test)
lossfn = nn.BCEWithLogitsLoss()
for name, cell, fb in [("plain RNN ", nn.RNN, None),
("LSTM, default init ", nn.LSTM, None),
("LSTM, forget bias set to 1", nn.LSTM, 1.0)]:
torch.manual_seed(0)
gen = torch.Generator().manual_seed(1)
model = Model(cell, fb)
opt = torch.optim.Adam(model.parameters(), lr=0.01)
for step in range(600):
x, y = make_batch(64, gen)
opt.zero_grad()
lossfn(model(x), y).backward()
nn.utils.clip_grad_norm_(model.parameters(), 1.0) # RNN gradients explode readily
opt.step()
with torch.no_grad():
acc = ((model(x_test) > 0).float() == y_test).float().mean().item()
print(f"{name} test accuracy {acc:.3f}")plain RNN test accuracy 1.000 LSTM, default init test accuracy 0.512 LSTM, forget bias set to 1 test accuracy 1.000
Exact numbers may shift on a different PyTorch version, since weight initialisation changes between releases. The pattern — one line at chance, two at a perfect score — is stable.
There are two lessons here and one of them is uncomfortable.
The comfortable one. An LSTM with default initialisation scored 0.512, which is a coin flip. The same LSTM with one line changed scored a perfect 1.000.
PyTorch initialises every bias near zero, so the forget tap starts at about 0.5. The memory is halved at every step. Over 30 steps that is 0.5 to the power of 30, roughly one in a billion. The gradient signal linking step 1 to the answer never arrives, so the network never discovers that step 1 mattered.
Setting the forget bias to 1.0 starts the tap at about 0.73 and, more importantly, starts it in the region where its own gradient is healthy. From there the network can learn to push it towards 1.0. This is standard practice, recommended since Gers et al. (2000) and confirmed at scale by Jozefowicz et al. (2015), and PyTorch still does not do it for you.
The uncomfortable one. The plain RNN also scored 1.000.
At 30 steps, with filler that carries no interference, a plain recurrent network manages the task. This page is not evidence that an LSTM beats an RNN on every sequence problem, and presenting it that way would be dishonest.
The honest evidence for LSTMs is the gradient table above, plus thirty years of published benchmarks at lengths where a controlled experiment costs more than a few seconds on a laptop. Any short training race you can run in a lesson is dominated by initialisation luck, not by architecture.
What this experiment does prove is that forget-gate initialisation, not the choice of cell, is what decided the outcome here. That is a practical fact worth more than a rigged comparison.
Common mistakes
Unpacking the return value wrongly. nn.RNN returns (output, h) and nn.LSTM returns (output, (h, c)). Swapping a cell without changing the unpacking gives ValueError: too many values to unpack at best, and a silent shape bug at worst. Writing seq, _ = self.rnn(x) works for both.
Leaving the forget bias at zero. Demonstrated above. One line, and it decides whether the model learns.
Forgetting batch_first=True. Without it PyTorch expects (length, batch, features). This trips up everyone once.
Not packing variable-length batches. Padding contributes loss and gradient unless you mask it or use nn.utils.rnn.pack_padded_sequence. Otherwise the model dutifully learns to predict padding.
Using bidirectional=True on a live stream. A bidirectional LSTM reads the sequence backwards as well, so it needs the whole sequence before it can output anything. It is excellent for offline labelling and impossible for real-time transcription.
Stacking many layers by reflex. Two or three layers is the usual range. Depth in an LSTM adds sequential work per step and rarely helps the way it does in a feedforward network.
Not detaching the state between batches. Carrying (h, c) across batches without h.detach() keeps the previous batch's graph alive. You get RuntimeError: Trying to backward through the graph a second time, and rising memory use.
Choosing an LSTM for a long document. Above a few hundred tokens, use a transformer. Choose an LSTM for streams and for tight hardware, not for text.
Try it yourself
In does_it_survive.py, change the forget bias from 5.0 to 0.0, which puts the tap at exactly 0.5.
Predict the length-40 number before running it. Then run it. The LSTM should now fall off as fast as the plain RNN, or faster.
That single change turns the whole architecture off. The gates are not the magic on their own — a forget tap held near one is. This is why the initialisation trick matters so much, and why "I used an LSTM" is not by itself a claim about long-range memory.
Then swap nn.LSTM for nn.GRU in forget_bias.py. A GRU merges the forget and input taps into one, and its gate order is [reset, update, new], so the bias slice changes. Working out the correct slice yourself is the fastest way to learn to read these layers.
What to learn next
- Transformers — the architecture that replaced this one for language, and why.
- Recurrent neural networks — the simpler cell, and the failure this lesson repairs.
- PyTorch basics — the training loop these examples assume.
Researcher — Mathematics and papers.
The full recurrence
Hochreiter & Schmidhuber (1997) introduced the memory cell; Gers, Schmidhuber & Cummins (2000) added the forget gate, giving the form universally used today:
i_t = σ( W_i x_t + U_i h_{t−1} + b_i ) input gate
f_t = σ( W_f x_t + U_f h_{t−1} + b_f ) forget gate
o_t = σ( W_o x_t + U_o h_{t−1} + b_o ) output gate
g_t = tanh( W_g x_t + U_g h_{t−1} + b_g ) candidate update
c_t = f_t ⊙ c_{t−1} + i_t ⊙ g_t cell state
h_t = o_t ⊙ tanh( c_t ) hidden statex_t ∈ R^d— input at stept;dis the input dimensionh_t, c_t ∈ R^m— hidden and cell state;mis the hidden sizeW_• ∈ R^{m×d},U_• ∈ R^{m×m},b_• ∈ R^m— per-gate parametersσ— logistic sigmoid, mapping to(0,1), so each gate is a soft per-coordinate switch⊙— elementwise (Hadamard) product
The gates are vectors, not scalars. Each of the m memory slots is independently written, held and read.
Implementation note. PyTorch concatenates the four gates into single matrices weight_ih_l0 ∈ R^{4m×d} and weight_hh_l0 ∈ R^{4m×m}, ordered [i, f, g, o], and carries two bias vectors b_ih and b_hh — a redundancy inherited from cuDNN. Parameter count per layer is therefore
4 m (d + m + 2) PyTorch, with the duplicated bias
4 m (d + m + 1) the mathematical minimumFor d = m = 512: 2.10 M parameters per layer. Roughly four times an Elman RNN of the same width, which is the price of the gates.
Why the gradient survives: the constant error carousel
Differentiate the cell recurrence with respect to the previous cell state:
∂c_t / ∂c_{t−1} = diag(f_t) + ( terms via h_{t−1} inside i_t, f_t, g_t )The dominant term is diag(f_t). Composing across a span from s to t:
∂c_t / ∂c_s ≈ Π_{k=s+1..t} diag(f_k)Compare the Elman case, where each factor is diag(φ'(a_k)) · W_hh. Three differences matter.
- No weight matrix in the path. The product involves no repeated multiplication by
W_hh, so its spectral radius does not govern decay. Bengio, Simard & Frasconi's dichotomy — stable storage requires contraction, gradient flow forbids it — does not bind on this path. - No saturating derivative in the path.
tanh'has supremum 1 andσ'supremum 0.25; neither appears here. - The decay rate is learned and data-dependent. With
f_k → 1the product approaches the identity, giving the constant error carousel of the original paper.
The measurement in the Developer section is this expression made numeric. At f = σ(5) ≈ 0.9933, the product over 199 steps is 0.9933^199 ≈ 0.263, matching the observed drop from 0.4835 to 0.1305.
The corollary is the exercise at the end of that section: with b_f = 0, f ≈ 0.5 and 0.5^199 underflows. An LSTM whose forget gate sits near 0.5 has no more long-range gradient than a vanilla RNN. The architecture supplies a path; the gate values decide whether it is open.
Forget-gate bias initialisation
b_f = 1 (sometimes 2) is standard and the reason is exactly the above. Jozefowicz, Zaremba & Sutskever (2015) ran an architecture search over thousands of RNN variants and reported that adding a forget bias of 1 was worth more than most architectural changes they tested — including, on several tasks, the difference between LSTM and GRU.
nn.LSTM initialises all parameters from U(−1/√m, 1/√m), so the default forget bias is near zero. Correcting it takes one line, applied to the [m : 2m] slice of bias_ih_l0:
with torch.no_grad():
lstm.bias_ih_l0[m:2*m].fill_(1.0)There is no output block because there is nothing to print: fill_ mutates the
parameter in place and returns the tensor it wrote into.
TensorFlow's LSTM layer has unit_forget_bias=True by default. PyTorch does not, which is a live source of unexplained training failures.
Variants
Peephole connections (Gers & Schmidhuber, 2000) let the gates see c_{t−1} directly: i_t = σ(W_i x_t + U_i h_{t−1} + p_i ⊙ c_{t−1} + b_i). Needed for tasks requiring precise interval timing, where the output gate would otherwise be blind to the cell it controls. Rarely used now; not exposed by nn.LSTM.
Coupled input and forget gates set i_t = 1 − f_t, saving m(d + m + 1) parameters at negligible cost in quality.
GRU (Cho et al., 2014) merges the cell and hidden states and uses two gates:
z_t = σ( W_z x_t + U_z h_{t−1} + b_z ) update gate
r_t = σ( W_r x_t + U_r h_{t−1} + b_r ) reset gate
n_t = tanh( W_n x_t + r_t ⊙ (U_n h_{t−1}) + b_n )
h_t = (1 − z_t) ⊙ n_t + z_t ⊙ h_{t−1}3m(d + m + 1) parameters, 25 percent fewer than an LSTM. Chung et al. (2014) found the two comparable on the tasks they tested. Greff et al. (2017), LSTM: A Search Space Odyssey, ablated eight variants over 5,400 runs and found none reliably better than the standard LSTM, with the forget gate and the output activation the two components whose removal hurt most.
Layer-normalised LSTM (Ba, Kiros & Hinton, 2016) normalises the pre-activations per step, stabilising training at the cost of losing the fused cuDNN kernel — typically a 2–4× slowdown.
ConvLSTM (Shi et al., 2015) replaces the matrix products with convolutions, giving a recurrent model over spatial grids for precipitation nowcasting and video.
Cost, against a transformer
sequential steps train time / layer inference state
LSTM O(T) O(T · m²) O(m) constant
Transformer O(1) O(T² d + T d²) O(T · d) KV cacheThe first column is the decisive one. Vaswani et al. (2017) cite the number of sequential operations explicitly as a design driver. An LSTM's step t cannot begin before t−1 completes, so wall-clock training time scales with T irrespective of how many accelerators you have. This is what ended LSTM dominance in language, and it is a statement about hardware utilisation rather than about representational power.
The third column is why recurrence survives. An LSTM's decoding state is O(m) regardless of how much it has read; a transformer's KV cache grows linearly in context. For a model streaming for hours, or running inside a hearing aid, that is the binding constraint.
Maximum path length between two positions is O(1) under attention and O(T) under recurrence, which removes the long-range gradient argument entirely rather than mitigating it. That is the second, quieter reason transformers won.
Where LSTMs remain in production
Streaming ASR. The RNN-Transducer (Graves, 2012) pairs a recurrent encoder over audio with a prediction network over emitted labels, and emits tokens without seeing future audio. He et al. (2019) describe a fully on-device LSTM RNN-T shipped for mobile speech recognition. Conformer encoders have replaced LSTMs in many places, but chunked-streaming variants still carry recurrent state, and RNN-T remains the dominant streaming loss.
Probabilistic forecasting. DeepAR (Salinas et al., 2020) is an LSTM emitting the parameters of a likelihood per step, trained across many related series. It remains a live baseline in forecasting libraries, and the datasets in this domain are frequently far too small to train a transformer well.
Reinforcement learning with partial observability. R2D2 (Kapturowski et al., 2019) uses an LSTM over the agent's observation history with stored recurrent states in the replay buffer. Constant-size state per timestep is the reason.
Embedded and low-power inference. Keyword spotting, gesture recognition and sensor anomaly detection on microcontrollers, where the KV cache is not affordable.
The revival
xLSTM (Beck et al., 2024) revisits the cell directly. Two changes: exponential rather than sigmoid gating with a normaliser state, which permits revising a stored value the original could not overwrite; and mLSTM, which replaces the vector cell with a matrix memory and a covariance update rule, removing the sequential dependency in the recurrence so training parallelises. Reported competitive with transformers and Mamba at billion-parameter scale.
This sits inside the broader linear-recurrence line — S4, Mamba, RWKV — surveyed in the researcher block of the RNN lesson. The shared observation is that the fixed-size state which forced LSTMs out of language is also the property that makes constant-memory inference possible, and that the sequential training constraint, not the recurrence itself, was the fatal part.
Papers
- Hochreiter, S. & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation 9(8).
- Gers, F., Schmidhuber, J. & Cummins, F. (2000). Learning to Forget: Continual Prediction with LSTM. Neural Computation 12(10).
- Gers, F. & Schmidhuber, J. (2000). Recurrent Nets that Time and Count. IJCNN.
- Graves, A. (2012). Sequence Transduction with Recurrent Neural Networks. — arxiv.org/abs/1211.3711
- Cho, K. et al. (2014). Learning Phrase Representations using RNN Encoder-Decoder. — arxiv.org/abs/1406.1078
- Chung, J. et al. (2014). Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. — arxiv.org/abs/1412.3555
- Jozefowicz, R., Zaremba, W. & Sutskever, I. (2015). An Empirical Exploration of Recurrent Network Architectures. ICML.
- Shi, X. et al. (2015). Convolutional LSTM Network. NeurIPS — arxiv.org/abs/1506.04214
- Greff, K. et al. (2017). LSTM: A Search Space Odyssey. IEEE TNNLS — arxiv.org/abs/1503.04069
- Vaswani, A. et al. (2017). Attention Is All You Need. NeurIPS — arxiv.org/abs/1706.03762
- Kapturowski, S. et al. (2019). Recurrent Experience Replay in Distributed Reinforcement Learning. ICLR.
- Salinas, D. et al. (2020). DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks. IJF 36(3).
- Beck, M. et al. (2024). xLSTM: Extended Long Short-Term Memory. — arxiv.org/abs/2405.04517
What to learn next
- Transformers — the architecture that replaced this one for language, and why.
- Recurrent neural networks — the simpler cell, and the failure this lesson repairs.
- PyTorch basics — the training loop these examples assume.