GitHub Actions for ML projects
GitHub Actions is the machine that actually runs your checks on every push — and an ML pipeline needs a few habits ordinary web-app CI does not, like skipping an expensive retrain when only the docs changed.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
GitHub Actions is the service that runs your tests and checks automatically, every time code is pushed, on GitHub's own machines.
The analogy you have already lived
You have used a washing machine's programme dial. You pick "cottons, 40 degrees" once, and close the lid. Every load after that runs the same steps in the same order, without you standing there managing each stage by hand.
A GitHub Actions workflow file is that programme dial for your code. Write the steps once — install, train, test, build. Every push then runs the exact same sequence, on a fresh machine, without anyone remembering to do it by hand.
Why it exists
Every earlier lesson in this section built a real check: unit tests, behavioural tests, an integration test, a gate. All of that is worthless if it only runs when someone remembers to type the command.
GitHub Actions is the piece that removes "remembering" from the process. A push happens, a fresh virtual machine spins up, and your steps run in order. The result — pass or fail — shows up directly on the change itself, where nobody can miss it.
How it works
you push a change to GitHub
|
v
GitHub starts a fresh, empty machine
|
v
runs the steps you wrote, in order:
- check out the code
- install dependencies
- train (or skip, if nothing training-relevant changed)
- run the test suite from earlier in this section
- run the eval gate
|
all steps succeeded?
/ \
yes no
| |
green check red X, right on the changeA real example you have seen
Every green tick or red cross you have seen next to a pull request on GitHub was produced by exactly this. Nobody ran those checks by hand — a workflow file did, the moment the change was pushed.
Remember this
- A workflow file describes steps that run automatically on a fresh machine, on every push.
- The machine is thrown away after each run — nothing "sticks" between runs unless you explicitly cache or save it.
- ML pipelines add one extra habit ordinary web CI rarely needs: skip the expensive parts when nothing training-relevant changed.
What to learn next
- Automated retraining pipelines — what happens when the "train" step above is triggered on a schedule instead of by a human push.
- Model registries — a more durable home for the artifact this workflow uploads than a CI run's temporary storage.
- Canary releases for models — what a deploy job downstream of this workflow should actually do with a newly built model.
Developer — Code and libraries.
Setup
No installation — this runs on GitHub's own machines, triggered by .github/workflows/*.yml files in your repository.
Deciding whether to skip the expensive part
Training can take minutes even on a small model. Running it on every push, including a typo fix in a README, wastes real machine time. This small script decides, from the files a commit actually touched, whether training is worth it.
"""Decides whether a change is worth the cost of a full retrain-and-test job."""
import subprocess
import sys
TRAINING_RELEVANT_PREFIXES = ("src/", "requirements.txt")
def changed_files(base_ref: str = "HEAD~1") -> list[str]:
out = subprocess.run(
["git", "diff", "--name-only", base_ref, "HEAD"],
capture_output=True, text=True, check=True,
)
return [line for line in out.stdout.splitlines() if line]
def should_train(files: list[str]) -> bool:
return any(f.startswith(TRAINING_RELEVANT_PREFIXES) for f in files)
if __name__ == "__main__":
files = changed_files()
print("changed files:", files)
decision = should_train(files)
print("should_train:", decision)
sys.exit(0 if decision else 1)Run against two real commits in a small test repository — first a documentation-only change, then a dependency change:
git init -q && git add should_train.py && git commit -q -m "add should_train.py"
echo "# notes" > README.md
git add README.md && git commit -q -m "update readme"
python should_train.py # after a README-only commitchanged files: ['README.md'] should_train: False
echo "requests==2.31.0" > requirements.txt
git add requirements.txt && git commit -q -m "add dependency"
python should_train.py # after a commit that touches requirements.txtchanged files: ['requirements.txt'] should_train: True
Both are real runs against a real, tiny git repository with those exact two commits — git diff --name-only is reading actual commit history, not a simulation.
Testing the decision logic on its own
from should_train import should_train
def test_docs_only_change_skips_training():
assert should_train(["README.md", "docs/guide.md"]) is False
def test_a_change_under_src_triggers_training():
assert should_train(["src/model.py"]) is True
def test_a_requirements_change_triggers_training():
assert should_train(["requirements.txt"]) is True
def test_a_mix_of_relevant_and_irrelevant_files_triggers_training():
assert should_train(["README.md", "src/features.py"]) is Truepytest test_should_train.py -q.... [100%] 4 passed in 0.02s
The workflow file
name: ml checks
on:
push:
pull_request:
jobs:
train-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 2 # need the previous commit for should_train.py's diff
- uses: actions/setup-python@v5
with:
python-version: "3.11"
cache: pip # caches installed packages between runs
- name: Install dependencies
run: pip install -r requirements.txt
- name: Decide whether to train
id: decide
run: python should_train.py || echo "skip=true" >> "$GITHUB_OUTPUT"
- name: Train
if: steps.decide.outputs.skip != 'true'
run: python src/train.py
- name: Run the test suite
run: pytest -q
- name: Run the evaluation gate
run: python gate.py eval_report.json
# Shares the trained model with a later job, without retraining it there.
- name: Upload the model artifact
uses: actions/upload-artifact@v4
with:
name: model
path: model.joblibNo output block here — the real log is a live web page with collapsible steps and per-line timings, and reproducing it as static text would misrepresent what you actually see.
Line-by-line: the parts an ML workflow needs that a plain web app rarely does
cache: pip on setup-python caches installed packages by hashing requirements.txt — an unchanged file reuses the cache and skips re-downloading everything, which matters more here than in most web projects because ML dependencies (torch, scikit-learn) are large.
fetch-depth: 2 — GitHub's default checkout only fetches the latest commit, with no history. should_train.py needs the previous commit to diff against, so the checkout step has to ask for at least two.
if: steps.decide.outputs.skip != 'true' — this is what actually skips the training step. should_train.py exiting non-zero on a docs-only change sets an output the if condition on the next step reads.
upload-artifact — the trained model produced in this job does not automatically exist in any other job; each job starts on its own fresh machine. Uploading it as an artifact is how a later deployment job could download the exact model this run produced, rather than retraining or trusting a different copy.
Common mistakes
Printing a secret to the log "only to check it's set". GitHub Actions logs are visible to anyone with read access to the repository (and to the world, on a public repo). Never echo an API key or token — reference it only as ${{ secrets.NAME }} inside a step that uses it directly.
A cache key that never changes. If the pip cache key is not tied to requirements.txt's contents, a dependency bump silently keeps using stale cached packages. setup-python's built-in cache: pip handles this correctly by hashing the lock file for you — a hand-rolled cache step easily gets this wrong.
A matrix that multiplies cost without adding safety. Testing against Python 3.9 through 3.13, on three operating systems, is nine full runs for every push. Match the matrix to what you actually support in production, not to everything you could possibly test.
Running the expensive job on every single push during active development. should_train.py above is one answer; another common one is a [skip ci] marker in a commit message for genuinely trivial changes, used sparingly and on purpose.
Trusting pull_request_target with untrusted code. This trigger runs with access to secrets even for pull requests from forks — appropriate for a few narrow, careful use cases, and a real security risk if it checks out and runs a fork's code with that same access. Use plain pull_request unless you specifically need the elevated trigger and understand why.
Try it yourself
Add a second job, deploy, that depends on train-and-test finishing successfully (needs: train-and-test), downloads the uploaded model artifact, and prints its file size. This is the same artifact hand-off pattern real deployment jobs use, without needing any real deployment target to try it.
What to learn next
- Automated retraining pipelines — what happens when the "train" step above is triggered on a schedule instead of by a human push.
- Model registries — a more durable home for the artifact this workflow uploads than a CI run's temporary storage.
- Canary releases for models — what a deploy job downstream of this workflow should actually do with a newly built model.
Researcher — Mathematics and papers.
Fan-out cost and the matrix trap
A build matrix over $k$ independent dimensions (Python versions, operating systems, hardware backends) multiplies job count combinatorially: $n_1 \times n_2 \times \cdots \times n_k$. For CPU-only jobs this is usually a cost problem; for anything requesting GPU runners it is frequently a queueing problem as well, since GPU-backed CI capacity is far more constrained than CPU capacity across every major provider. The disciplined response is to test the full matrix on a slow cadence (nightly, or pre-release) and a single representative configuration on every push — the same fast/slow split argued for training itself in CI/CD for machine learning.
Caching as a content-addressed problem
setup-python's cache: pip and the general-purpose actions/cache action both key their cache on a hash of a declared file (typically requirements.txt or a lock file), not on a version number or a timestamp. This is a content-addressed cache: the key is a function of the exact dependency specification, so two commits with byte-identical dependency files always share a cache entry regardless of when they ran, and any change to that file — even reordering — is guaranteed to produce a new key. The correctness property this buys is real: a stale cache silently serving old dependencies is structurally impossible as long as the hashed file genuinely captures everything that affects the install.
Provenance and supply-chain considerations
An uploaded model artifact, as in the workflow above, is only as trustworthy as the chain that produced it. actions/upload-artifact and download-artifact do not by default provide cryptographic attestation that a downloaded artifact came from the specific run that claims to have produced it, within the same workflow run they are bound tightly enough for practical purposes, but cross-workflow or cross-repository artifact sharing has historically been an actively exploited weakness. For anything higher-stakes than an internal CI hand-off, sign the artifact (cosign, or GitHub's own artifact attestation feature) rather than trusting the filename alone — the same discipline covered for shipped models in model lineage and traceability.
Papers and further reading
- Kim et al., Containerized Continuous Integration: An Empirical Study of Adoption and Cost, various empirical CI studies converge on caching and job-skipping as the dominant levers for CI cost at scale — the same two levers this lesson focuses on.
- GitHub, Security hardening for GitHub Actions — the canonical treatment of the
pull_request_targetrisk and secrets handling flagged above.
What to learn next
- Automated retraining pipelines — what happens when the "train" step above is triggered on a schedule instead of by a human push.
- Model registries — a more durable home for the artifact this workflow uploads than a CI run's temporary storage.
- Canary releases for models — what a deploy job downstream of this workflow should actually do with a newly built model.