Manufacturing and Predictive Maintenance

Yield analysis and wafer maps

Where a chip fails on a silicon wafer matters as much as how many failed, because the spatial shape of failures points directly at which step of the process went wrong.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Where a chip fails on a silicon wafer matters as much as how many failed — the shape of the failures points at what went wrong.

Think about a large tray of cookies baked in an oven. If a few cookies scattered randomly across the tray come out slightly overdone, that is probably ordinary variation. Maybe there was a stray hot spot here and there. But if every cookie along the edge burns while the middle ones are perfect, that tells you something specific. Your oven heats unevenly near its walls. The pattern on the tray reveals the cause, not only the count of burnt cookies.

Semiconductor manufacturing reads its own version of that tray constantly.

Why it exists

A silicon wafer is a thin circular disc, cut after manufacturing into hundreds of individual chips called dies. Each die is tested, and yield is the percentage that pass. Factories watch this number obsessively — a single wafer can be worth a great deal, and a low-yield wafer represents real, expensive waste.

Knowing the yield percentage alone is not very useful on its own. Knowing where the failing dies sit on the wafer is what actually helps engineers fix the problem. Failures scattered randomly across the whole wafer usually mean ordinary background defects. Failures concentrated in a specific shape tell a different story. A ring around the edge, a cluster in one corner, a scratch-like line — each shape points to something specific. It almost always traces back to one fixable step in the process.

How it works

Wafer A: failures scattered randomly              Wafer B: failures ringed around the edge

     . . x . . .                                        x x x x x x
     . x . . . .                                        x . . . . x
     . . . x . .                                        x . . . . x
     . . . . x .                                        x . . . . x
     . x . . . .                                        x x x x x x

  (x = failed die)                                (x = failed die)

Random pattern:      probably ordinary background defects, nothing specific to chase
Edge-ring pattern:    strongly suggests uneven temperature or gas flow near the wafer's rim

Engineers have built up decades of knowledge linking specific spatial patterns to specific causes. An edge ring often points at uneven chamber temperature. A scratch line often points at a mishandling event during transport. A full-wafer cluster often points at contamination.

Where you have already seen it

  • A pizza with burnt edges but a perfectly cooked centre, from an oven that runs hotter near its walls than in the middle.
  • A photograph with dark corners (vignetting), caused by light falling off unevenly toward the edges of a camera lens.
  • Grass growing unevenly in a lawn, thinner near a tree's shade or a garden edge, revealing something specific about that location rather than random bad luck.

Remember this

  • Overall yield tells you how much failed. The spatial pattern of failures tells you why.
  • A random scatter of failures usually means ordinary background defects, not a specific fixable cause.
  • A recognisable shape — a ring, a cluster, a line — usually points directly at a specific step or piece of equipment in the process.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

We build a simple round wafer from a grid, then compare a wafer with random defects against one with a real edge-ring defect signature — checking whether a simple edge-versus-core comparison can tell them apart.

wafer_yield_analysis.py
import numpy as np

rng = np.random.default_rng(31)

grid_size = 21
center = grid_size // 2
radius = center

# Build a round wafer out of a square grid of dies -- only positions
# within the circle are real, testable chips ("dies").
yy, xx = np.mgrid[0:grid_size, 0:grid_size]
dist_from_center = np.sqrt((xx - center) ** 2 + (yy - center) ** 2)
is_die = dist_from_center <= radius
is_edge = is_die & (dist_from_center > 0.8 * radius)
is_core = is_die & ~is_edge

n_dies = is_die.sum()
print(f"total dies on this wafer: {n_dies} ({is_edge.sum()} near the edge, {is_core.sum()} in the core)")
print()


def make_wafer(edge_fail_rate, core_fail_rate):
    fail = np.zeros((grid_size, grid_size), dtype=bool)
    fail[is_edge] = rng.random(is_edge.sum()) < edge_fail_rate
    fail[is_core] = rng.random(is_core.sum()) < core_fail_rate
    return fail


def summarize(name, fail_map):
    yield_pct = 100 * (1 - fail_map[is_die].mean())
    edge_fail_rate = fail_map[is_edge].mean()
    core_fail_rate = fail_map[is_core].mean()
    ratio = edge_fail_rate / max(core_fail_rate, 1e-6)
    print(f"{name}:")
    print(f"  overall yield: {yield_pct:.1f}%")
    print(f"  edge fail rate: {edge_fail_rate:.1%}   core fail rate: {core_fail_rate:.1%}")
    print(f"  edge/core ratio: {ratio:.1f}x", "-> EDGE-RING pattern suspected" if ratio > 3 else "-> looks like ordinary random defects")
    print()


# Wafer A: random defects scattered evenly -- normal process variation.
wafer_random = make_wafer(edge_fail_rate=0.05, core_fail_rate=0.05)
summarize("Wafer A (random defects)", wafer_random)

# Wafer B: a real edge-ring signature, common when a process chamber's
# temperature or gas flow is uneven near the wafer's rim.
wafer_edge_ring = make_wafer(edge_fail_rate=0.45, core_fail_rate=0.04)
summarize("Wafer B (edge-ring defect)", wafer_edge_ring)
Output
total dies on this wafer: 317 (120 near the edge, 197 in the core)

Wafer A (random defects):
  overall yield: 94.0%
  edge fail rate: 6.7%   core fail rate: 5.6%
  edge/core ratio: 1.2x -> looks like ordinary random defects

Wafer B (edge-ring defect):
  overall yield: 76.7%
  edge fail rate: 58.3%   core fail rate: 2.0%
  edge/core ratio: 28.7x -> EDGE-RING pattern suspected

What actually happened

dist_from_center computes how far every position on the grid is from the wafer's middle, using ordinary Euclidean distance — the same geometry as checking how far a point sits from the centre of a circle. is_die keeps only the positions actually within the round wafer, discarding the corners of the square grid that fall outside it.

Wafer A's edge and core fail rates come out close to each other, 6.7% and 5.6% — a ratio near 1, exactly what random, spatially unstructured defects should look like. Wafer B's edge fail rate, 58.3%, dwarfs its core fail rate of 2.0% — a ratio of nearly 29x, an unmistakable spatial signature rather than random chance.

This is a deliberately simple version of a real diagnostic. It captures the core idea — compare failure rates across meaningful spatial regions, not only the overall count — that underlies far more sophisticated wafer-map pattern recognition used in real semiconductor fabs.

Common mistakes

Only tracking the overall yield number. Two wafers with an identical 85% yield can have completely different root causes if one fails randomly and the other fails in a ring — averaging away the spatial pattern throws away the most useful diagnostic information available.

Picking an arbitrary edge threshold without checking it against the process physics. The 0.8 * radius cutoff here is a reasonable illustrative choice; a real fab tunes this boundary based on known equipment geometry, not guesswork.

Assuming every non-random pattern has an equally obvious explanation. Edge rings are a well-understood classic. Many real spatial patterns are subtler, and diagnosing them reliably takes combining the spatial statistics with domain expertise about the specific process step under suspicion.

Comparing wafers processed under different conditions without controlling for that. A meaningful edge-versus-core comparison assumes both regions were exposed to the same process step; conflating data across different equipment or process recipes muddies the signal.

Try it yourself

Change the pattern to a center cluster instead: set a new region is_center_cluster = dist_from_center < 0.3 * radius, give it a high fail rate, and leave the rest low. Adapt summarize to report a centre-versus-outer ratio instead of edge-versus-core, and check that the same basic idea catches this different spatial signature too.

What to learn next

Researcher — Mathematics and papers.

Wafer maps as a spatial pattern recognition problem

Formally, a wafer map is a binary (or multi-class, for different failure bins) spatial field defined on the disc of testable die positions. The edge-versus-core comparison above is a simple, hand-specified spatial statistic. The general problem — classifying which of several known failure pattern classes a wafer map belongs to (edge-ring, centre, scratch, donut, random, none) — is a standard image classification task once the wafer map is treated as a 2D image, and is dominated in the literature by convolutional neural networks trained on labelled historical wafer maps.

The WM-811K benchmark

Wu, Jang and Chen (2015), Wafer Map Failure Pattern Recognition and Similarity Ranking for Large-Scale Data Analysis, IEEE Transactions on Semiconductor Manufacturing, released the WM-811K dataset — over 800,000 real wafer maps from actual fabrication lines, a substantial fraction hand-labelled into nine defect pattern classes (including "none," the majority class). It remains the standard public benchmark for this task, and its class distribution is itself instructive: even in a curated dataset, most wafers show no identifiable pattern, meaning this is another domain-typical class-imbalance problem, closely related to the rare-failure setting in Building a model from four failures.

CNN-based approaches and their design considerations

CNNs applied to wafer maps (Nakazawa and Kulkarni, 2018, Wafer Map Defect Pattern Classification and Image Retrieval Using Convolutional Neural Network, IEEE Transactions on Semiconductor Manufacturing) generally outperform hand-engineered spatial statistics once enough labelled examples exist, because they can learn complex or mixed pattern types (a wafer showing both an edge-ring and a localised cluster simultaneously) that a single hand-specified statistic like the edge/core ratio cannot represent. A key domain-specific design choice: wafer maps are naturally represented on a disc, not a square grid, and rotation invariance matters (a scratch pattern rotated 90 degrees is the same underlying defect) — architectures and augmentation strategies for this task typically build in rotational and reflective symmetry explicitly, rather than relying on the network to learn it from data alone.

Root cause linkage: from pattern to process step

Pattern classification alone identifies that a spatial signature exists; linking it to a specific upstream process step (which chamber, which recipe parameter, which equipment maintenance event) is typically solved as a downstream correlation or causal analysis problem, joining the wafer-map classification output against process-tool logs and recipe metadata — an application of the same predict-then-decide thinking discussed generally in Predict then optimise, here applied to root-cause engineering rather than inventory allocation. This linkage step, not the pattern classification itself, is usually where the majority of real diagnostic value is created in a production yield-management system.

Key references

  • Wu, M.-J., Jang, J.-S. R. & Chen, J.-L. (2015). Wafer Map Failure Pattern Recognition and Similarity Ranking for Large-Scale Data Analysis. IEEE Transactions on Semiconductor Manufacturing 28(1).
  • Nakazawa, T. & Kulkarni, D. V. (2018). Wafer Map Defect Pattern Classification and Image Retrieval Using Convolutional Neural Network. IEEE Transactions on Semiconductor Manufacturing 31(2).

What to learn next