Dimensionality Reduction

Independent component analysis

ICA un-mixes blended signals back into their original independent sources — the algorithm behind separating voices, brain waves and instruments from recordings that captured them together.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

ICA takes signals that were mixed together and recovers the original separate sources, without ever being told what they were.

Stand in a wedding hall with two conversations happening — one about cricket to your left, one about property prices to your right. Each of your ears receives both conversations, blended: the left ear gets more cricket, the right more property. Yet your brain hands you either conversation on demand, cleanly separated.

This is the cocktail party problem: several sources, several sensors, every sensor hearing a different blend. ICA is the algorithm that solves it — from the blends alone.

Why it exists

The world rarely gives you pure signals. Microphones record all speakers at once. Scalp electrodes record millions of neurons plus every eye blink. A satellite pixel mixes reflections from soil, crop and water. In each case you have several recordings, each a different mixture of the same hidden sources, and you want the sources back.

PCA cannot do this. PCA finds directions of large variation, and mixed conversations vary loudly in directions that correspond to no single voice. The key insight of ICA is different: real sources are independent — the cricket chat carries no information about the property chat. So ICA searches for the un-mixing that makes the recovered signals as unrelated to each other as possible. Where PCA asks "where is the variation?", ICA asks "how do I make the outputs stop sharing information?".

How it works

hidden sources         what sensors record             ICA recovers
tabla:  _#_#_#       mic 1: mostly tabla, some flute      tabla:  _#_#_#
flute:  ~~~~~~       mic 2: mostly flute, some tabla      flute:  ~~~~~~
                          (both are mush)

ICA needs at least as many sensors as sources. It also exploits a strange fact. Mixtures of different signals look more ordinary, more bell-curve-like, than the originals do. So ICA turns the un-mixing knobs to make each output look as distinctive and un-bell-like as possible. When every output has stopped looking like a blend, the sources have been found.

A real example you have seen

Noise-cancelling earbuds and video-call background suppression continuously separate "voice" from "everything else" using multiple microphones. And in hospitals, ECG machines separate a baby's faint heartbeat from the mother's overpowering one — two sources, several chest electrodes, one un-mixing.

Remember this

  • ICA recovers hidden sources from recorded blends.
  • It works by making outputs independent — sharing no information — not by chasing the biggest spread.
  • It needs several sensors, each hearing a different mixture.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs verified with scikit-learn 1.7.2.

Un-mixing a tabla and a flute

Two synthetic instruments, two "microphones", and a race between ICA and PCA.

ica.py
import numpy as np
from sklearn.decomposition import FastICA, PCA

t = np.linspace(0, 8, 2000)
tabla = np.sign(np.sin(3 * t))            # sharp, drum-like square wave
flute = np.sin(1.3 * t)                   # smooth sine
sources = np.column_stack([tabla, flute])

# two microphones, each hearing a different blend
mixing = np.array([[0.6, 0.4], [0.35, 0.65]])
mics = sources @ mixing.T

ica = FastICA(n_components=2, random_state=0)
unmixed = ica.fit_transform(mics)

def best_match(recovered, original):
    """ICA can flip sign and swap order, so report the best absolute correlation."""
    c = np.abs(np.corrcoef(recovered.T, original.T)[:2, 2:])
    return c.max(axis=1).round(3)

print("ICA recovery correlation:", best_match(unmixed, sources))
print("PCA recovery correlation:", best_match(PCA(2).fit_transform(mics), sources))
Output
ICA recovery correlation: [0.999 0.997]
PCA recovery correlation: [0.821 0.799]

The walkthrough

0.999 versus 0.821. ICA's outputs correlate almost perfectly with the true tabla and flute — the blends are undone. PCA's best rotation still leaves each output around 20% contaminated by the other instrument. PCA found the loudest directions; the loudest directions of a mixture are still mixtures.

Why best_match shrugs at sign and order. ICA has two built-in ambiguities: it cannot know which source is "first", and it cannot know a source's original sign or volume — flipping a voice's waveform upside down changes nothing about its independence. Recovered sources come back in arbitrary order, arbitrarily scaled. Any honest evaluation, and any downstream code, must tolerate this.

Why a square wave and a sine. ICA's engine runs on non-ordinariness — technically, distance from a bell curve. Square waves (two spikes of values) and sines (values piling at the extremes) are both delightfully un-bell-like. Had both sources been bell-curve noise, ICA would be mathematically blind: mixtures of bell curves are indistinguishable from rotations of them. One bell-curve source is allowed; two is failure.

FastICA is the standard algorithm — fast, and dependable on well-conditioned problems. It internally whitens the data first (a PCA-like step), then rotates to maximise non-ordinariness.

Common mistakes

Interpreting component order as importance. Unlike PCA, ICA components come in no order at all. Run twice with different seeds and the tabla may swap slots with the flute. Never write components[0] expecting "the main one".

Using ICA to compress. ICA is a separator, not a compressor. It does not rank directions by information kept, and dropping ICA components is not a principled reduction. Compress with PCA; separate with ICA.

Feeding it fewer sensors than sources. Two microphones cannot cleanly separate three speakers with vanilla ICA — the mixing is no longer invertible. Overcomplete variants exist, but the sklearn tool assumes sensors >= sources.

Expecting magic on a single channel. One recording of band music cannot be ICA-separated into instruments — there is nothing to un-mix; every sample is one number. Single-channel separation is a different, harder problem solved today with neural methods — see speech separation approaches.

Try it yourself

Add a third source, rng.normal(size=2000) (bell-curve noise), a third microphone row to the mixing matrix, and set n_components=3. One Gaussian source is legal — check whether the tabla and flute still come back clean.

What to learn next

Researcher — Mathematics and papers.

Model and identifiability

ICA posits x = A s: observed x in R^m are an unknown invertible mixing A applied to sources s with statistically independent components. The task is estimating W = A^{-1} from samples of x alone. Identifiability (Comon, 1994, Independent component analysis, a new concept?, Signal Processing): sources are recoverable up to permutation and scaling if and only if at most one source is Gaussian. The Gaussian case fails because a rotation of an isotropic Gaussian is identical in distribution — independence plus Gaussianity is rotation-blind. Second-order statistics (covariance) cannot distinguish candidate un-mixings; ICA necessarily consumes higher-order structure.

Objectives and FastICA

Equivalent formulations: maximise non-Gaussianity of each output (kurtosis, or negentropy J(y) = H(y_gauss) - H(y)); minimise mutual information among outputs; maximum likelihood under non-Gaussian source priors; Infomax (Bell and Sejnowski, 1995). FastICA (Hyvarinen, 1999, IEEE Trans. Neural Networks) whitens x, then finds w maximising E[G(w^T z)] for a robust contrast G (log cosh being standard) via an approximate Newton fixed-point iteration:

w_new = E[z g(w^T z)] - E[g'(w^T z)] w, then orthonormalise

Convergence is cubic near solutions, and deflation or symmetric decorrelation extracts multiple components. Cost per iteration is O(n m^2) after whitening — cheap; the whitening itself is a PCA at O(n m^2 + m^3).

Applications with teeth

EEG/MEG artifact removal is the flagship: eye blinks, cardiac and muscle artifacts separate from neural sources given many electrodes (Makeig et al., 1996), and ICA remains standard preprocessing in that field. fMRI resting-state networks, hyperspectral unmixing, and financial factor extraction follow the same template. In each, the physical mixing really is close to linear and instantaneous — the model's core assumption. Convolutive mixing (echoes, room acoustics) violates it; frequency-domain and time-context extensions handle that, and modern single-channel separation abandoned the linear model for supervised deep networks entirely (Conv-TasNet, Luo and Mesgarani, 2019).

Relation to modern methods

ICA is the ancestor of disentangled representation learning: the wish that latent factors be independent and semantically atomic. Nonlinear ICA is unidentifiable without side information (Hyvarinen and Pajunen, 1999) — a negative result that haunts VAE-style disentanglement (Locatello et al., 2019 formalised the impossibility without inductive bias) — while auxiliary-variable and contrastive formulations (Hyvarinen et al., 2019) restore identifiability under time structure or labels, an active research seam.

What to learn next