RuntimeError: An attempt has been made to start a new process before the current process has finished its bootstrapping phase
On Windows and macOS, worker processes re-import your script, and unguarded top-level code re-runs forever. Wrap the entry point in if __name__ == "__main__".
Updated
The error
RuntimeError:
An attempt has been made to start a new process before the
current process has finished its bootstrapping phase.
This probably means that you are not using fork to start your
child processes and you have forgotten to use the proper idiom
in the main module:
if __name__ == '__main__':
freeze_support()
...
The "freeze_support()" line can be omitted if the program
is not going to be frozen to produce an executable.Sometimes it loops endlessly, spawning process after process, instead of stopping cleanly.
What it means
On Windows (and macOS with Python 3.8+), new processes start with the spawn method: each worker launches a fresh Python and imports your script to reconstruct what it needs. Any code at the top level of the script runs again during that import — including the line that creates workers. Each worker then tries to create its own workers, and Python detects the runaway before your machine does.
In deep learning, the process creation is usually hidden inside DataLoader(num_workers=4).
Why it happens
The script has training code at module level:
# train.py — broken on Windows
loader = DataLoader(ds, batch_size=32, num_workers=4)
for x, y in loader: # first iteration spawns workers -> workers import this file
...Linux defaults to fork, which clones the process without re-importing, so the same script works there. Code moves from a Linux tutorial to a Windows laptop and breaks.
How to fix it
1. Put the entry point behind the main guard.
from torch.utils.data import DataLoader
def main():
loader = DataLoader(ds, batch_size=32, num_workers=4)
for x, y in loader:
...
if __name__ == "__main__":
main()__name__ equals "__main__" only in the process you launched. In workers, the file is imported under a different name, the guard is False, and nothing re-runs. Definitions (classes, functions, imports) can stay at top level — only actions need the guard.
2. freeze_support() is only for frozen .exe builds. Scripts do not need it; the guard alone fixes the error.
3. In Jupyter on Windows, keep the Dataset importable. Notebook cells complicate spawn. If workers fail in a notebook, either set num_workers=0 (fine for small data) or move the Dataset class into a .py file the notebook imports.
loader = DataLoader(ds, batch_size=32, num_workers=0)4. Avoid lambdas and nested functions in Dataset/collate_fn. Spawn must pickle what workers need, and lambdas do not pickle. Use module-level functions.
How to prevent it
Structure every training script as def main() plus the guard, from the first line you write — it costs nothing and works identically on Linux, Windows and macOS. Treat "runs on Linux, spawns forever on Windows" as this exact bug until proven otherwise.
Related errors
- DataLoader worker is killed by signal — other ways workers die
- OSError: WinError 1455 — The paging file is too small — many spawned workers exhausting Windows memory commit
- The kernel appears to have died