New: Debug encrypted microservice traffic with Speedscale's eBPF collector Read the announcement

GitLab CI job pods use C3D, N2D, and N2 Spot pools with an on-demand fallback; image builds use a separate BuildKit pool.

How We Rebuilt GitLab CI on GKE Spot VMs


Here at Speedscale, our monorepo has 19 Go projects in CI. Fifteen produce container images for both amd64 and arm64; a few of those need CGO. We also publish cross-compiled CLI binaries for macOS, Linux, and Windows.

We host our repos on GitLab and use our own GitLab CI runners, which for over 5 years were a fleet of EC2 instances that scaled up and down on a schedule - basically the time blocks we knew we were probably going to need them. When scaled up, these instances ran whether they were executing jobs or not, otherwise they sat idle. Is it sub-optimal? Yes. Was it Good Enough(TM) for a small engineering team that has more important things to work on? Also yes. And (shocker) when the latter is part of the picture, you tend to just stick with the duct tape and baling wire solution, even though you detest it.

That’s bad in an obvious way (spoiler: it’s needlessly expensive), but it’s also bad in a less obvious way: optimization. Every attempt at optimization with this kind of scenario ends up reaching for an isolated, local one. Some of the things we did included trying to cache things on the instance’s disk and conditionally running project jobs based on whether or not they were impacted (directly or indirectly). Does it work? Sure, but it’s slow as hell and doesn’t really fix the real problems.

So what did we do?

We moved the runners to a GKE Standard cluster. Most job capacity uses Spot VMs and scales to zero when idle; a small on-demand node keeps the runner manager available. That attacks the old system’s idle-compute bill, but Spot capacity comes with a reliability problem we had to solve. The engineering challenge was making ephemeral Kubernetes jobs useful even when a particular machine type is unavailable. AI helped me build it, but the cluster taught us which assumptions were wrong.

What five years of duct tape looked like

Before the teardown, a quick nickel tour of what we were living with.

A merge request didn’t run a pipeline of the same statically defined jobs - it would more or less generate one. A prepare stage did diff analysis across projects to figure out which ones were, or could be, impacted by a change. The exception was changing a common shared library or the pipeline configuration itself, which would force everything to run. Regardless, this prepare stage would run a bash script to produce a pipeline YAML artifact that it would then launch as a dynamic child pipeline. If we had a problem with a job configuration, we’d have to download this generated YAML artifact and determine what went wrong.

Objectively gross, but that was one of our micro-focused optimization attempts at making merge request builds “faster”. Which it did and it was acceptable enough that we called it done and moved on.

Every job that packaged a container image also spun up its own throwaway docker:dind builder, which removed any possibility of shared layer caching between builds.

No shared layer cache.

Every. Single. Time.

Combine this with a whole slew of other problematic things, e.g. job preparation stages installing tools that should be already available, relying on jobs to pass along tarballs of cache for the next job to use, and per-project Makefiles where everything was mostly similar but somehow just a little bit different, and you start to get the idea.

This system grew through five years of decisions that made sense at the time. As they piled up, they compounded the problem.

Building the replacement

The first thing we needed to do was to ditch the idea that we needed persistent available runners and instead consider everything to be ephemeral and disposable. Arguably something we should have done at the beginning, but like I mentioned, you tend to just accept Good Enough(TM) when you have other priorities and no dedicated team of engineers to handle fixing the build pipeline.

So where did we start? Deciding that we wanted to run CI jobs on GitLab Runners in GKE. This was the system we had originally wanted to have, but using this also made a lot of sense from an infrastructure standpoint, as we’ll see later.

The GKE cluster by itself is just an empty playground. To really make this work, we had to actually think through each design decision about how our software was built, tested, and published. Probably the biggest choice was that not a single job would require docker:dind. For us, that was a big lift: years of technical build debt buried across Dockerfiles with QEMU binfmt sprinkled in to cross-compile CGO binaries.

Our image-build jobs are intentionally thin and disposable. They check out the target commit and send the build context to one shared buildkitd pod. BuildKit does the heavy image build and keeps a layer cache on persistent storage. Go verification, tests, and CLI compilation still run in job pods, where their CPU and memory requests can be sized for the work.

We also made the decision that everything a job needs to build or test something is baked into the job’s image (e.g. Go, Docker, gobuildcache, libpcap headers, etc). We had attempted to have something like this before but we were too aggressive in wanting a single, monolithic builder image. Good idea, but poor execution. We ended up with a bloated, unnecessarily large base image that was a junk drawer of tools or libraries that maybe one or two apps needed.

What we have now is a system where jobs run with no prepare stage and no artifact hand-off between build and package, so there’s no job dependency graph to screw up. Simple.

Infrastructure

The cluster has a small on-demand system pool, a dedicated Spot pool for BuildKit, and several pools that can run CI job pods. The system pool uses one e2-standard-2 node for the GitLab Runner manager and cluster services. It stays up so GitLab can hand the runner work even after every job pool has scaled to zero.

BuildKit runs as a single StatefulSet on a dedicated c3d-standard-16 Spot pool. Its cache lives on an attached high-availability persistent disk across two zones, so a node preemption does not erase the cache. It still kills in-flight image builds and causes a stall while BuildKit is rescheduled. We isolate it from ordinary job pods with a separate label and taint: job bursts cannot fill the node that holds the shared builder.

Why we added more than one job pool

We first used c3d-standard-16 Spot nodes for CI jobs. They are a good fit for our mix of image-build clients, Go verification, tests, and CLI builds, and the pool scales to zero outside pipeline hours. Then we learned the part that a machine-size calculation misses: Spot capacity is a market, not a reservation. When GKE could not provision C3D in the zones we had chosen, runner pods stayed Pending even though the pool’s maximum had room for them. A higher node limit cannot create capacity that Google does not have available.

Our next move was to give the scheduler more eligible capacity. We added n2d-standard-16 and n2-standard-16 Spot pools with the same job-pod label and taint as C3D, across the zones that support those machine types. The runner’s node selector and toleration match all three. GKE can provision from a different machine family or zone when one Spot market is short. We also added an n2d-standard-16 on-demand pool with the same scheduling contract. It is an escape hatch for times when none of the configured Spot pools can supply nodes, not a permanently warm second fleet. Its minimum is zero, so it can scale back down after the jobs finish.

Each job pool has its own eight-node ceiling; that is not a promise of 32 active nodes. The runner’s admission limit is derived from the configured builder-node limit: currently 48 jobs for the eight-node planning target, shared across all job pools. Jobs beyond that wait in GitLab rather than becoming more Pending pods. We still watch pending pods, node count, and job failures: more machine types improve the odds of getting capacity, but they do not guarantee immediate scheduling or remove Spot preemption.

A preempted job may be retried when its GitLab job defines the matching retry reasons. Our main build jobs do; we do not assume every pipeline job does. BuildKit failures during an image solve have a separate retry path. Keeping BuildKit and job nodes in separately tainted pools limits the blast radius, but a preemption can still waste work.

Pipeline improvements and optimizations

Go module proxying and caching

When our build infrastructure was in AWS, cache was just more duct tape: we would use GitLab’s object-storage cache to grab the entire Go build cache, zip it up, and ship it to S3. Each job would pull that down, extract it, and attempt to use it. It did pretty much nothing useful, as you might expect.

On the new CI platform, Go modules come from an Artifact Registry repository set up to proxy proxy.golang.org. Keeping the repository in the same region as GKE avoids Artifact Registry data-transfer charges for those pulls; storage still costs money. Granted, you have to make sure to set GOPROXY but it’s pretty straightforward:

GOPROXY: "${GOPROXY_URL}|https://proxy.golang.org,direct"

There’s some subtlety here that I should address. The pipe between ${GOPROXY_URL} and proxy.golang.org means that on any error, Go falls through to the next endpoint. If separated by a comma, fallthrough only happens on a 404 or 410 error. Artifact Registry has an upstream fetch quota, and a swarm of jobs starting simultaneously against a cold cache will start getting 429s. Go happens to treat that 429 as an authoritative answer, so the module fetch fails and the job dies due to rate limiting. The pipe changes that behavior so we’ll always fall back to proxy.golang.org if we can’t get it directly from Artifact Registry.

Private modules, i.e. our internal private repos, never touch the Artifact Registry proxy at all. Setting GOPRIVATE routes them direct over git-https using an insteadOf token config, which also means they skip the checksum database.

Go build cache

We could have stopped here, called it a day, and enjoyed all of the speed gains. But doing so would have prevented us from really speeding things up using a custom GOCACHEPROG. This is the piece I’d recommend to anyone running Go CI pipelines. It was released in Go 1.24 in early 2025 which, all things considered, isn’t that old at all, so it makes sense if you’ve never heard of it.

In a nutshell, GOCACHEPROG lets you specify an external program to manage build caches for you. Rather than relying on the standard behavior with a GOCACHE directory (which you’d have to share between jobs), Go will instead run the program you specify and communicate with it via a pipe, allowing the helper program to fully manage build caches. If you’re curious about the technical details, Depot’s remote caching writeup is a pretty good deep dive into the topic.

What this meant for us was that we could use something like gobuildcache and have it manage build cache directly from a GCS bucket:

GOBUILDCACHE_BACKEND_TYPE: "gcs"
GOBUILDCACHE_GCS_BUCKET: "${GOCACHE_BUCKET}"

In practice this gets us a cache with per-action granularity (individual compile and test steps), shared across every job and every project, and reachable from a pod with no volume attached to it at all. That’s a huge win. Cache fetches still take a network round trip, but that was a better trade for us than copying one large tarball between jobs.

If you decide to use gobuildcache, be aware that it holds a per-key flock. So if you configure it to use a shared host volume mounted to CI pods, a sudden burst of build jobs will likely fail because it cannot acquire the lock. To make sure you don’t run into this, configure gobuildcache to use pod-local directories for its own local cache and just use GCS as the shared layer across all running jobs.

# Pod-local on purpose, NOT the shared hostPath: gobuildcache holds a per-key flock
# across the GCS fetch with a hardcoded 1s acquire timeout, so pods stampeding the same
# keys die with "failed to acquire lock."
GOBUILDCACHE_CACHE_DIR: "/tmp/gobuildcache/cache"
GOBUILDCACHE_LOCK_DIR: "/tmp/gobuildcache/locks"

Keep in mind that Go’s build cache is write-heavy and entirely regenerable, so a more aggressive expiration/retention window for the GCS bucket is probably a good idea. We set ours to expire at 14 days (really more of a guess since we didn’t have data to back it up). There’s really no sense in spending money for GCS storage for old build cache artifacts that you’re never going to use again.

What about CGO cross-compilation?

Most of our services are pure Go, so building and packaging them for multiple platforms is a non-issue. It’s an extremely simple setup: let Go do the cross-compilation, then use a COPY-only Dockerfile to package it.

However, a few projects aren’t that simple and require CGO to link against C libraries. That means builds need actual C toolchains per compilation target (see musl.cc). On our old AWS system, we dealt with this in the laziest way possible: builders were x86_64, so the monolithic build image I mentioned earlier included an arm64 musl cross toolchain in every job, whether it was needed or not. On top of that, every image build job registered QEMU binfmt handlers (via a privileged qemu-user-static container, naturally) so the arm64 image stages could run under emulation. Wasted time, wasted bandwidth, and a privileged container per job. Not exactly ideal.

Now, these services that require CGO compile completely on the buildkitd instance using a fantastic little setup called xx. I wrote more about the technique in Easy Cross-Platform cgo Builds:

FROM --platform=$BUILDPLATFORM ${BASE_REG}/tonistiigi/xx:1.9.0 AS xx

FROM --platform=$BUILDPLATFORM ${BASE_REG}/golang:${GO_VERSION}-alpine AS build
COPY --from=xx / /
RUN apk add --no-cache git clang lld
ARG TARGETPLATFORM
RUN xx-apk add --no-cache gcc musl-dev libpcap-dev

RUN --mount=type=cache,target=/root/.cache/go-build \
    --mount=type=cache,target=/go/pkg \
    --mount=type=secret,id=netrc,target=/root/.netrc \
    xx-go build -trimpath \
      -tags "${TAGS:+${TAGS},}netgo,osusergo" \
      -ldflags "-s -w -linkmode external -extldflags -static" \
      -o /out/goproxy ./goproxy && \
    xx-verify /out/goproxy

For context, the tonistiigi/xx image comes with xx- prefixed wrappers around the usual tools that all key off TARGETPLATFORM. So in the snippet above, xx-apk installs the target architecture’s headers and libraries, xx-go wires up clang with the correct toolchain sysroot, and xx-verify confirms the output binary is actually the architecture you asked for. The compile stage runs on the buildkitd instance without emulating the target CPU. We still install an ARM binfmt handler on that node for runtime-stage commands that execute inside an arm64 image. The improvement is that we no longer register QEMU in every CI job.

Cloud NAT costs money

Our CI nodes are private and egress through Cloud NAT, which bills a data processing fee per GiB in each direction. Every FROM golang:1.26 or FROM debian:trixie-slim that hits Docker Hub is billed traffic. That really starts to add up to something problematic when you’re building many different projects, each for two architectures, across a few dozen pipelines a day.

Artifact Registry, on the other hand, is reachable over Private Google Access, which never touches the NAT gateway. Those same-region pulls have no Artifact Registry data-transfer charge, though the repository still has storage costs. Every base and toolchain image our Dockerfiles reference is mirrored into our own AR repo, and that repo is the default everywhere:

IMAGE_BASE_REGISTRY="${IMAGE_BASE_REGISTRY:-us-central1-docker.pkg.dev/<gcp-project-id>/<image-repo>}"

At some point, you have to pull the bandaid off

It’s worth being honest about how this actually got built, because the answer is “almost entirely not by hand”.

This was deliberately not an incremental migration. The old system was years of accumulated technical debt. Porting that piece by piece would have meant carefully preserving decisions that only ever made sense for how the old system operated.

So we just decided to pull the bandaid off: start with no assumptions and rebuild the entire CI and build system based on where we wanted to be, not what was already there.

And you know what? I did the work, soup to nuts, with Claude Code running Fable 5. The first version was not the last. Spot shortages, scheduling limits, and build failures forced more changes after we started using it.

I reviewed every line, and went back and forth with the agent plenty of times before anything merged. But the honest summary of the whole experience?

Fable 5 might be my new favorite DevOps coworker.

If you want to see the same idea applied to testing rather than building: Speedscale captures real production traffic and replays it against your service, so your CI can run integration tests against the actual shapes your API sees instead of the ones somebody imagined at fixture-writing time. Try proxymock.

Stop writing API mocks by hand

proxymock records real traffic from your running app and replays it as mocks — HTTP, gRPC, Postgres, Kafka, and more. Install in 30 seconds, no account required.