Building images with BuildKit¶
Every image rendimiento deploys, including rendimiento itself and this book, is built by BuildKit. This chapter explains what that means from the ground up.
Images in two minutes¶
A container image is a stack of layers (tar archives of files) plus a small JSON config (the command to run, environment, user, exposed ports) and a manifest listing them. Everything is content-addressed: each piece is named by the SHA-256 of its bytes, and the image as a whole by the SHA-256 of its manifest: the digest, like sha256:5fa9dd15….
- A tag (
:1.2.3,:latest) is a movable label; it can point at a different image tomorrow. - A digest can never change meaning: the same digest is always the same bytes.
That is why every rendimiento release pins images by digest (registry.example.lan:5000/shop-api@sha256:…): what runs is exactly what was built and tested, and a rollback brings back exactly what ran before.
What BuildKit is¶
BuildKit is the build engine behind docker build, available on its own. It has two parts:
buildkitd, a daemon that does the work: runs build steps in isolated containers, keeps a content-addressed cache, pushes results to a registry. It needs privileges (it creates containers), so it runs in its own dedicated pods;buildctl, a thin client that sends it a build and streams the progress back. Build pods run only this client, unprivileged.
Inside, BuildKit does not run a Dockerfile line by line. A frontend first translates the build definition into LLB, a graph of low-level operations (fetch an image, run a command, copy files). BuildKit then:
- runs independent parts of the graph in parallel (for example two stages that do not depend on each other);
- skips any operation whose inputs have not changed (the cache key is the hash of the operation and its inputs);
- pulls only what it needs.
Two frontends are used here:
| Frontend | Input | Used for |
|---|---|---|
dockerfile.v0 (built in) |
a Dockerfile |
services with a Dockerfile |
gateway.v0 + ghcr.io/railwayapp/railpack-frontend |
a Railpack build plan (JSON) | services without one (Railpack) |
There is no Docker daemon anywhere in the cluster; nothing needs one.
Multi-stage Dockerfiles¶
A multi-stage Dockerfile has several FROM lines. Each starts a stage with its own base image; later stages copy only what they need from earlier ones with COPY --from=<stage>. Only the last stage becomes the image. Compilers, package caches and sources stay in the build stages and are thrown away, so the final image is small and has less to attack.
rendimiento's own Dockerfile¶
FROM node:20-alpine AS ui
WORKDIR /web
COPY web/package.json web/package-lock.json ./
RUN npm ci
COPY web/ ./
# Types and translations are checked here, so every image build checks the UI.
RUN npm run typecheck && npm run build
FROM golang:1.27 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
COPY --from=ui /web/dist ./web/dist
ARG GIT_SHA=""
RUN CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /out/rendimiento ./cmd/rendimiento
FROM gcr.io/distroless/static-debian12:nonroot
COPY --from=build /out/rendimiento /rendimiento
USER nonroot:nonroot
EXPOSE 8080 9090
ENTRYPOINT ["/rendimiento"]
| Stage | Base | Does | Kept in the image |
|---|---|---|---|
ui |
node:20-alpine |
npm ci, then npm run build (Vite) → web/dist |
nothing directly |
build |
golang:1.27 |
go mod download, copies the sources and the built UI, compiles |
nothing directly |
| final | distroless/static-debian12:nonroot |
— | one file: the /rendimiento binary (with the UI embedded) |
Details worth knowing:
- Order for caching.
go.mod/go.sum(andpackage.json/package-lock.json) are copied and their dependencies downloaded before the sources. A code change then reuses the dependency layer; only a dependency change downloads again. CGO_ENABLED=0builds a pure-Go, statically linked binary: no C library needed, so it runs ondistroless/static, an image with no shell, no package manager, nothing but CA certificates and a non-root user.-trimpath -ldflags="-s -w"removes local paths and debug symbols (a smaller binary).- The UI is compiled in its own stage and copied in, and Go's
//go:embedputs it inside the binary. - The result is about 55 MB and runs as a non-root user, with a read-only root filesystem (see
deploy/rendimiento.yaml).
This book's Dockerfile¶
# syntax=docker/dockerfile:1
# The book, built in three stages from the whole repository (build context
# = repository root, so the book can include real source files):
# 1. codemap: generate the code reference from the Go sources
# 2. site: build the static site with MkDocs Material
# 3. runtime: serve it with an unprivileged nginx on port 8080
# Only the last stage becomes the image; the others are thrown away.
FROM golang:1.27-alpine AS codemap
WORKDIR /src
COPY . .
RUN go run ./hack/codemap > /code-map.md
FROM python:3.12-slim AS site
WORKDIR /src
# Dependencies first: this layer is reused until requirements.txt changes.
COPY docs/requirements.txt docs/requirements.txt
RUN pip install --no-cache-dir -r docs/requirements.txt
COPY . .
COPY --from=codemap /code-map.md docs/content/reference/code-map.md
RUN cd docs && mkdocs build --site-dir /site
FROM nginxinc/nginx-unprivileged:1.27-alpine
COPY --from=site /site /usr/share/nginx/html
COPY docs/nginx.conf /etc/nginx/conf.d/default.conf
EXPOSE 8080
Three stages with three different toolchains (Go, Python, nginx). BuildKit runs codemap and the pip install of site in parallel, since neither depends on the other until the final COPY --from=codemap.
The templates for detected apps¶
When a repository has no Dockerfile and you choose to add one, rendimiento generates it from templates/dockerfiles/. For example, Next.js with output: 'standalone':
FROM node:{{.Version}}-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
ARG GIT_SHA=""
ENV GIT_SHA=$GIT_SHA NEXT_TELEMETRY_DISABLED=1
RUN npm run build
FROM node:{{.Version}}-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production NEXT_TELEMETRY_DISABLED=1
RUN addgroup --system --gid 1001 nodejs && adduser --system --uid 1001 nextjs
COPY --from=builder /app/public ./public
COPY --from=builder --chown=nextjs:nodejs /app/.next/standalone ./
COPY --from=builder --chown=nextjs:nodejs /app/.next/static ./.next/static
USER nextjs
EXPOSE {{.Port}}
ENV PORT={{.Port}} HOSTNAME="0.0.0.0"
CMD ["node", "server.js"]
How a build runs here¶
The build pod's step container runs buildctl against a BuildKit daemon (the scripts):
buildctl --addr tcp://buildkitd-2.buildkitd-pool.devops-tools.svc.cluster.local:1234 build \
--frontend dockerfile.v0 --local context=/workspace/src/api --local dockerfile=/workspace/src/api \
--opt filename=Dockerfile --opt build-arg:GIT_SHA=… \
--output type=image,name=registry.example.lan:5000/shop-api:4f2a9c1e0b7d,push=true,registry.insecure=true \
--import-cache type=registry,ref=registry.example.lan:5000/shop-api:buildcache,registry.insecure=true \
--export-cache type=registry,ref=registry.example.lan:5000/shop-api:buildcache,mode=max,registry.insecure=true \
--metadata-file /tmp/metadata.json
--local context=…sends the cloned folder to the daemon.--output type=image,…,push=truepushes the result, tagged with the commit's short SHA; the digest comes back inmetadata.json.- The registry is plain HTTP (
registry.insecure=true); the daemons know it from their config (buildkitd.toml).
Caching, twice¶
-
Each daemon's local cache, on a node-local volume (
local-path, up to about 15 GB, garbage-collected). Fastest, but only on that node.Mind the units in
buildkitd.tomlSince BuildKit 0.17, a bare number in the GC settings is bytes. The pool first had
gckeepstorage = 15000, meant as 15 GB, which BuildKit read as 15 kB. Every daemon threw its cache away after each build, and every build re-downloaded its base images and cached layers: a tiny Go app went from 85 s to over 7 minutes. The config inp0dxD/gitops/buildkit/config.yamlnow uses explicit units:reservedSpace = "5GB",maxUsedSpace = "15GB",minFreeSpace = "15%". To check what a daemon really applies:kubectl -n devops-tools exec buildkitd-0 -- buildctl --addr unix:///run/buildkit/buildkitd.sock debug workers --verbose.- A registry cache per image:
--export-cache …:buildcache,mode=maxstores every intermediate layer (modemax, not just the final ones) under the image'sbuildcachetag;--import-cachereads it. Any daemon can use it, so a service moved to another daemon is still reasonably warm.
- A registry cache per image:
The BuildKit pool¶
One daemon cannot use more than one node's CPUs. The pool runs one daemon per worker node (not on the control plane, nor on nodes excluded from builds such as a GPU node) as a StatefulSet with pod anti-affinity. It is an add-on defined in p0dxD/gitops/buildkit/.
The platform finds the daemons through a headless Service (buildkitd-pool): its DNS SRV records list only daemons that are ready. For each build, KubeExecutor.buildkitFor chooses one by rendezvous hashing on the image name:
flowchart LR
key["image: shop-api"] --> h0["hash(buildkitd-0 | shop-api) = 812…"]
key --> h1["hash(buildkitd-1 | shop-api) = 403…"]
key --> h2["hash(buildkitd-2 | shop-api) = 977…"]
key --> h3["hash(buildkitd-3 | shop-api) = 150…"]
h2 --> win["highest wins → buildkitd-2"]
Each image has a stable favourite, so it always finds its warm cache there, while different images spread across the pool and build in parallel. If a daemon disappears, only the images that favoured it move; everything else stays put. If no daemon is ready, builds fall back to BUILDKIT_ADDR. The build log says which daemon built each image.
// buildkitFor picks the pool daemon for key by rendezvous hashing: each key
// has a stable favourite among the ready daemons, and only the keys of a
// daemon that goes away move elsewhere. It returns the address and the
// daemon's name, or the single BuildkitAddr (and "") without a pool.
func (k *KubeExecutor) buildkitFor(ctx context.Context, key string) (string, string) {
if k.BuildkitPool == "" {
return k.BuildkitAddr, ""
}
res := k.Resolver
if res == nil {
res = net.DefaultResolver
}
lctx, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
_, srvs, err := res.LookupSRV(lctx, "buildkit", "tcp", k.BuildkitPool)
if err != nil || len(srvs) == 0 {
return k.BuildkitAddr, ""
}
var best *net.SRV
var bestScore uint64
for _, s := range srvs {
h := fnv.New64a()
h.Write([]byte(strings.TrimSuffix(s.Target, ".") + "|" + key))
if score := h.Sum64(); best == nil || score > bestScore {
best, bestScore = s, score
}
}
host := strings.TrimSuffix(best.Target, ".")
daemon, _, _ := strings.Cut(host, ".")
return fmt.Sprintf("tcp://%s:%d", host, best.Port), daemon
}
Railpack: no Dockerfile needed¶
For a service without a Dockerfile, the build pod's plan container runs the Railpack CLI (railpack prepare). It detects the language (Node, Python, Go, Java, static sites…), picks versions, and writes a build plan: which packages to install (through mise), which commands to run, and what the image should start. The step container hands that plan to BuildKit's gateway.v0 frontend with Railpack's frontend image, which turns it into LLB.
| Dockerfile | Railpack | |
|---|---|---|
| You write | a Dockerfile | nothing (build.start to override the start command) |
| Control | total | through railpack.json or build args |
| Image size | small with multi-stage (e.g. 4 MB for a Go app on distroless) | larger (about 40 MB for the same Go app: a general runtime base) |
| Runs as | whatever you choose | root by default |
Use Railpack to get started quickly; add a Dockerfile when size, hardening or system packages matter. A service with a Dockerfile always uses it; build.builder can force either.
Building rendimiento itself: make image¶
# Build on the cluster's BuildKit pool (the buildkitd Service reaches one of
# its daemons), not on this node: main is also the k3s control plane, and a
# local compile starves its SQLite datastore.
BUILDKIT_PORT ?= 12345
image:
@kubectl port-forward -n devops-tools svc/buildkitd $(BUILDKIT_PORT):1234 >/dev/null 2>&1 & pf=$$!; \
trap "kill $$pf 2>/dev/null" EXIT; sleep 3; \
docker run --rm --network host -v $(CURDIR):/src:ro --entrypoint buildctl moby/buildkit:v0.18.2 \
--addr tcp://127.0.0.1:$(BUILDKIT_PORT) build --frontend dockerfile.v0 \
--local context=/src --local dockerfile=/src \
--output type=image,name=$(IMAGE):latest,push=true,registry.insecure=true \
--import-cache type=registry,ref=$(IMAGE):buildcache,registry.insecure=true \
--export-cache type=registry,ref=$(IMAGE):buildcache,mode=max,registry.insecure=true
make image port-forwards to the buildkitd Service and runs the buildctl client in a throwaway container with the repository mounted read-only. The work happens on a pool daemon, not on main.
Architectures¶
Every node is arm64, so every image is built for linux/arm64. A cloud target on x86 (linux/amd64) would need either:
- multi-platform builds:
--opt platform=linux/amd64,linux/arm64produces one image index containing both; an arm64 daemon builds the amd64 half under QEMU emulation, which is slow; or - native builders for each architecture: an amd64 daemon in the pool, with the executor choosing by platform (the rendezvous key could include it).
See Scaling and open source.