LSTM (long short-term memory)
In one sentence An LSTM is an RNN with learned gates that decide what to remember, what to forget, and what to output — letting memory survive long sequences.
Updated
An LSTM is a recurrent network with built-in gates that learn what to keep, what to discard, and what to reveal at each step.
A plain RNN is a student taking notes without an eraser or a highlighter — everything gets scribbled onto the same crowded page, and early notes drown. An LSTM has both tools plus discipline. At each step it asks: what in my notes is now irrelevant (forget gate)? What from this new sentence deserves writing down (input gate)? And what do I need visible for the current question (output gate)?
The core trick is a separate cell state: a memory track that runs through time and is modified mostly by addition — erase a little here, write a little there — rather than being wholly rewritten each step. Because information rides along this track without repeated multiplication, gradients survive the journey backwards. That is the direct fix for the vanishing-gradient problem, and it let networks connect a word to context hundreds of steps earlier.
Invented in 1997 by Hochreiter and Schmidhuber, the LSTM powered a striking share of 2010s AI: Google Translate's first neural version, Siri-era speech recognition, and handwriting recognition. The transformer has since taken most of that territory, but LSTMs remain a sensible choice for small models, streaming audio, and modest time-series work.
Where to go next
- Full lesson: RNNs
- Related terms: rnn, vanishing-gradient, transformer, gradient-clipping