Error database

DataLoader worker is killed by signal / insufficient shared memory

The helper processes loading your data were killed, either because the machine ran out of RAM or because shared memory is too small. Rerun with num_workers=0 to see the real error first.

The message you saw
DataLoader worker is killed by signal / insufficient shared memory

By Updated

The error

Output
RuntimeError: DataLoader worker (pid 3821) is killed by signal: Killed.
Output
ERROR: Unexpected bus error encountered in worker. This might be caused by insufficient shared memory (shm).
RuntimeError: DataLoader worker (pid 3821) is killed by signal: Bus error.
Output
RuntimeError: DataLoader worker (pid 3821) exited unexpectedly with exit code 1. Details are lost due to multiprocessing. Rerunning with num_workers=0 may give better error messages.

What it means

DataLoader with num_workers greater than zero starts separate processes to load and prepare batches while the GPU trains on the previous one. One of those processes died, and the main process noticed it had gone.

The three messages have three different causes, and the wording tells you which you have:

  • signal: Killed — the Linux out-of-memory killer terminated the process. The machine ran out of RAM.
  • Bus error or a message naming shm — the shared memory area that workers use to hand tensors back is full. This is a Docker default, not a hardware limit.
  • exit code 1 — the worker hit an ordinary Python exception inside your dataset code, and multiprocessing lost the traceback.

Why it happens

Each worker is a full copy of the Python process. Linux fork shares memory pages until they are written to, which sounds like it should keep workers cheap — but CPython updates a reference count on every object it touches, and that write copies the page. So a dataset holding a large Python list or dictionary gradually becomes a private copy in every worker. Eight workers can quietly turn a 3 GB dataset into 24 GB of resident memory.

Loaded batches travel back to the main process through shared memory, which on Linux is the /dev/shm filesystem. Docker gives a container 64 MB of it by default, which is far too small for image or audio batches. This is the single most common cause on any containerised setup, and it has nothing to do with how much RAM the machine has.

Beyond those two: a genuine bug in __getitem__ — a corrupt file, a missing path, a PIL failure on one image out of fifty thousand — kills the worker with exit code 1. And num_workers set higher than the machine's core count multiplies memory use while adding no speed.

How to fix it

1. Set num_workers=0 and run again. Do this first, before changing anything else. With no worker processes, the loading code runs in the main process and Python prints the real traceback.

python
train_loader = DataLoader(dataset, batch_size=32, num_workers=0)

If a real exception appears — a missing file, a bad image, an index error — that is your bug, and the rest of this page does not apply. Fix it and put the workers back.

2. If you are in Docker, raise the shared memory size. This is the fix whenever the message mentions shm or a bus error.

bash
docker run --shm-size=8g --gpus all -it your-image
yaml
# docker-compose.yml
services:
  train:
    image: your-image
    shm_size: '8gb'

--ipc=host also works and gives the container the host's shared memory. Use it when you control the host and trust what runs in the container; --shm-size is the more contained option. In Kubernetes, mount an emptyDir with medium: Memory at /dev/shm. You can confirm the current size with df -h /dev/shm.

3. Confirm whether the machine actually ran out of RAM. On Linux, the kernel records every kill:

bash
dmesg -T | grep -i "killed process"
Output
[Wed Aug 20 02:14:07 2026] Out of memory: Killed process 3821 (python) total-vm:38412992kB, anon-rss:15884224kB

If your process is listed there, it is a RAM problem, so continue with steps 4 and 5. Watching free -h or htop in a second terminal during the first minute of training shows the same thing live.

4. Reduce workers, prefetching and batch size. Memory used by loading is roughly num_workers × prefetch_factor × batch size.

python
train_loader = DataLoader(
    dataset,
    batch_size=32,
    num_workers=4,            # a sensible start; more is not automatically better
    prefetch_factor=2,        # batches queued per worker
    pin_memory=True,          # faster host-to-GPU copies
    persistent_workers=True,  # avoids respawning workers every epoch
)

Four workers is a reasonable default. Going above the number of physical CPU cores rarely helps and always costs memory. persistent_workers=True avoids the memory spike of tearing down and recreating workers at every epoch boundary, which matters when epochs are short.

5. Stop the per-worker memory growth at its source. If your dataset holds a big Python list of paths, labels or dictionaries, store it in a form that does not trigger reference counting on every access:

python
import numpy as np

class Files(torch.utils.data.Dataset):
    def __init__(self, paths, labels):
        # one numpy array instead of a large Python list of str objects
        self.paths = np.array(paths, dtype="S")   # fixed-width bytes, no per-item objects
        self.labels = np.array(labels, dtype=np.int64)

    def __len__(self):
        return len(self.labels)

    def __getitem__(self, i):
        path = self.paths[i].decode()
        return load_image(path), self.labels[i]

The important principle: load file paths in __init__, load actual file contents in __getitem__. A dataset that reads every image into memory up front will be copied into every worker.

6. On Windows and macOS, guard the entry point. These platforms spawn workers rather than forking, which re-imports your script in each new process. Without a guard, the script starts training again inside every worker.

python
if __name__ == "__main__":
    main()

Missing this produces recursive process creation, a flood of errors, and sometimes this same message.

7. If workers die on one specific file, find it. Run with num_workers=0 and iterate the dataset directly:

python
for i in range(len(dataset)):
    try:
        dataset[i]
    except Exception as exc:
        print(f"index {i} failed: {type(exc).__name__}: {exc}")

One truncated JPEG in a scraped folder is a common culprit, and PIL.ImageFile.LOAD_TRUNCATED_IMAGES = True handles the mild cases.

How to prevent it

Set the shared memory size in your Dockerfile documentation, compose file or run script once, so it is never a surprise on a new machine. --shm-size=8g costs nothing when it is not used.

Keep num_workers at a moderate value tied to the hardware rather than a number copied from a tutorial:

python
import os
num_workers = min(8, (os.cpu_count() or 2))

Validate the dataset before the first long run — iterate all of it once with num_workers=0 and catch the broken files early. And when you scale up a job, watch memory for the first two minutes. RAM that climbs steadily rather than settling means the per-worker copying problem, and it will kill the run at hour three, not minute three.