Stop rebuilding your CI runner on every push
Stateless CI runners made builds clean, and then made us write caching logic into every workflow. Warm microVM snapshots give you a disposable runner that already remembers everything.
CI used to run on long-lived build servers (remember Jenkins?). They were a pain to maintain: state piled up and builds influenced each other. But they had one great property. They were always warm. Yesterday’s Docker layers and packages were still on disk.
So we moved to ephemeral runners. Partly for isolation and reproducibility, but mostly because it was simple: GitHub Actions and GitLab handed them to us managed, out of the box. No build servers to babysit. The catch is that every job starts from a clean machine and pays again to rebuild its world from scratch.
Our fix was caching. Lots of it, spelled out in every workflow.
The mess we got used to
Take something everyone has built: a Node app that runs npm ci, runs its tests, and builds a Docker image. A properly cached GitHub Actions workflow looks something like this:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm # cache #1: ~/.npm
- run: npm ci
- run: npm test
- uses: docker/setup-buildx-action@v3
- uses: actions/cache@v4 # cache #2: BuildKit cache mounts
with:
path: buildkit-cache
key: buildkit-${{ hashFiles('package-lock.json') }}
- uses: reproducible-containers/buildkit-cache-dance@v3
with:
cache-map: '{"buildkit-cache": "/root/.npm"}'
- uses: docker/build-push-action@v6
with:
push: true
tags: ghcr.io/acme/app:${{ github.sha }}
cache-from: type=gha # cache #3: image layers
cache-to: type=gha,mode=max
Three different caches, each with its own keys, limits and quirks. There’s even a third-party action whose only job is to smuggle npm’s cache in and out of BuildKit, because RUN --mount=type=cache doesn’t survive a fresh runner on its own.
And after all that, every cache still has to be downloaded at the start of the job and uploaded at the end, over the network, on every run.
The idea: make the whole machine the cache
Modern microVMs can start from a memory and disk snapshot in a fraction of a second. That changes the maths. Instead of starting from an empty image and restoring caches into it, you start a fresh runner from a warm snapshot: an OS image that already has Docker, the image layers, ~/.npm, pip caches and your build tools on its local disk.
warm snapshot ──(< 1 s)──▶ ephemeral runner ──▶ your normal workflow
▲ │
└──────── successful main build: new snapshot ◀────┘
Each runner is still disposable. It never gets reused. But the baseline it boots from carries state forward. Every successful build on main becomes the next snapshot, so the environment keeps warming itself up as CI runs. Pull requests boot from the latest main snapshot but never write back to it, so one bad branch can’t poison everybody’s cache.
The workflow from above becomes:
runs-on: warm-runner
steps:
- uses: actions/checkout@v4
- run: npm ci # ~/.npm is already on disk: almost nothing to download
- run: npm test
- run: docker build -t ghcr.io/acme/app:${{ github.sha }} .
# base image, layers and cache mounts are already local:
# only the layers your change touched get rebuilt
- run: docker push ghcr.io/acme/app:${{ github.sha }}
No cache keys, no buildx plumbing, no cache dance. These are the plain commands you’d run on your laptop, and they’re fast for the same reason they’re fast on your laptop the second time: the cache is already there. If something isn’t there yet, it simply gets downloaded or built, like on any normal machine.
Caching becomes infrastructure instead of workflow logic. The workflow describes what to build, not how to babysit a temporary machine.
Where the time goes
Here’s what to expect for a typical Node service. These are illustrative numbers, not a benchmark:
| Step | Clean runner + caches | Warm microVM |
|---|---|---|
| Runner start | 10–30 s | < 1 s |
| Restore caches over the network | 20–60 s | 0 s, already on disk |
npm ci |
30 s | 10 s |
docker build (only app code changed) |
60–90 s | 10–15 s |
| Save caches | 15–40 s | 0 s, snapshot happens after the job |
| Total | ~2.5–4 min | ~30–40 s |
The biggest win isn’t any single step. It’s that the restore-and-save overhead, the part that grows with every cache you add, disappears entirely.
And you get a real VM
Speed is only half of it. A microVM is a real virtual machine with its own kernel, which fixes a list of long-standing CI annoyances:
- Real isolation. Each job runs behind a hardware virtualisation boundary, not a shared kernel. That matters more and more when the code being tested was written by an AI agent.
- No Docker-in-Docker. The runner has its own Docker daemon. No privileged containers, no mounted sockets, no DinD sidecars on your Kubernetes runners.
- Normal Linux. systemd works, and so do services like Postgres or Redis started the normal way. So does anything else that assumes it owns the machine.
- Persistence when you want it. State lives in the snapshot, not in a pile of cache keys.
Why this matters now
AI tools generate more code, in smaller and more frequent changes, than teams ever produced by hand. Every one of those changes goes through CI. Coding agents even iterate against CI: push, wait for the result, fix, push again. Slow CI used to cost a developer a coffee break. Now it’s a bottleneck for the whole loop.
We need CI to be faster, but we also need it to do more: more tests, more security scanning, more end-to-end checks, because there’s more code that no human wrote line by line. You can only afford those extra checks if the baseline cost of a run gets close to zero. Warm runners give you that budget back.
What’s next
I’m testing this with a simple experiment: take one real workflow, run it on a microVM runner booted from a warm snapshot, delete every caching step, and measure. I’ll publish the real numbers here.
The bigger idea is that this shouldn’t be yet another CI system. GitHub Actions, GitLab CI and Buildkite can keep running the workflow. What changes is the layer underneath.
GitHub Actions already has the hooks for this. When a job is queued, a provider can boot a microVM from the latest warm snapshot, register it as a just-in-time runner for exactly that one job, and throw it away afterwards. From the developer’s side, adopting it is a one-line change:
runs-on: ubuntu-latest → runs-on: warm-runner
Then you delete the caching steps you no longer need.
What I don’t want is to build and run that microVM infrastructure myself, and I doubt you do either. Providers like Boxd, Daytona and Modal already solve the hard parts: fast boot, snapshots and isolation at scale. The missing piece is the integration: connect your GitHub organisation, choose when snapshots get promoted, and set runs-on. If you’re building one of those platforms, this is the integration I’d love to see.