Skip to content

How the platform is deployed

rendimiento deploys apps, but it does not deploy itself. The platform is installed with plain Kubernetes manifests in deploy/, applied with kubectl apply -k deploy. Keeping the platform outside its own control means a broken rendimiento can always be fixed with kubectl.

Your own settings

The manifests in deploy/ carry example values (example.com, registry.example.lan, [email protected]). A real cluster's settings belong in a private overlay, a kustomization in a private repository that builds on deploy/ and replaces what differs:

# kustomization.yaml in a private repository, next to a clone of this one
resources:
  - ../../rendimiento.ai/deploy
images:
  - name: registry.example.lan:5000/rendimiento
    newName: registry.home.lan:5000/rendimiento   # your registry
patches:
  - path: configmap.yaml          # the rendimiento ConfigMap with your settings
  - target: { kind: Ingress, name: rendimiento }
    patch: |-
      - { op: replace, path: /spec/rules/0/host, value: rendimiento.your-domain.com }
      - { op: replace, path: /spec/tls/0/hosts/0, value: rendimiento.your-domain.com }

Then point the Makefile at it from a local.mk (git-ignored) at the repository root:

IMAGE := registry.home.lan:5000/rendimiento
DEPLOY_DIR := $(HOME)/private-repo/rendimiento-platform
TEST_EXCLUDE_NODES := my-gpu-node     # nodes remote tests must avoid
ITEST_REGISTRY := registry.home.lan:5000

make deploy renders DEPLOY_DIR (deploy/ by default) and applies it. Your hostnames, node names and email then never enter the public repository.

What gets installed

flowchart TB
    subgraph rs[namespace rendimiento-system]
        dep[Deployment rendimiento<br/>1 replica, Recreate]
        svc[Service rendimiento :80]
        ing[Ingress rendimiento.joserod.space<br/>TLS by cert-manager]
        db[(StatefulSet rendimiento-db<br/>Postgres 17, 5Gi Longhorn)]
        cm[ConfigMap rendimiento<br/>settings]
        sa[ServiceAccount rendimiento]
        sa2[ServiceAccount rendimiento-addons<br/>no pods, impersonated]
    end
    subgraph rb[namespace rendimiento-builds]
        np[NetworkPolicy isolate-builds]
    end
    crds[CRDs apps.rendimiento.ai<br/>addons.rendimiento.ai]
    ing --> svc --> dep --> db
    dep -. reads .-> cm
File What it contains
kustomization.yaml The list below, applied together.
namespace.yaml rendimiento-system (the platform) and rendimiento-builds (CI pods).
crds/rendimiento.ai_apps.yaml, crds/rendimiento.ai_addons.yaml The two custom resource definitions, generated from api/v1alpha1 by make generate. Never edit them by hand.
rbac.yaml What the platform may do (see below), the rendimiento-addons identity bound to cluster-admin, and the build namespace's Role.
postgres.yaml The platform database: a one-replica StatefulSet on a 5 Gi Longhorn volume.
rendimiento.yaml The settings ConfigMap, the Deployment, its Service and Ingress.
networkpolicy.yaml Isolation for CI pods: DNS, BuildKit and the public internet only.
railpack/Dockerfile Not applied: the image with the Railpack CLI that build pods use (make railpack-image).
deploy/rendimiento.yaml (settings, Deployment, Service, Ingress)
apiVersion: v1
kind: ConfigMap
metadata:
  name: rendimiento
  namespace: rendimiento-system
data:
  # Example values: a real cluster's settings live outside this public
  # repository (an overlay applied over these manifests; see the book).
  BASE_URL: https://rendimiento.example.com
  ALLOWED_USERS: your-github-login
  GITHUB_APP_NAME: rendimiento
  DNS_ZONE: example.com
  DNS_PROXIED: "true"
  REGISTRY: registry.example.lan:5000
  REGISTRY_INSECURE: "true"
  BUILDKIT_ADDR: tcp://buildkitd.devops-tools.svc.cluster.local:1234
  # One BuildKit per worker (the "buildkit" add-on); each image builds on
  # the daemon its name hashes to. BUILDKIT_ADDR is the fallback.
  BUILDKIT_POOL: buildkitd-pool.devops-tools.svc.cluster.local
  MAX_PARALLEL_STEPS: "4"
  # Longest a single test or build step may run. First builds on the Pis
  # start with a cold cache (e.g. compiling a Python wheel from source).
  STEP_TIMEOUT: 45m
  # Nodes that must never run builds: e.g. a node whose kernel cannot
  # enforce the build NetworkPolicy, or a GPU node reserved for inference.
  BUILD_EXCLUDE_NODES: gpu-node
  # One replica with the Recreate strategy never overlaps, so leader
  # election only adds a way to crash when the API server is slow.
  LEADER_ELECTION: "false"
  CLUSTER_ISSUER: letsencrypt-prod
  INGRESS_CLASS: nginx
  STORAGE_CLASS: longhorn
  # Labels added to every app volume claim; these put it in a Longhorn backup job.
  # VOLUME_LABELS: recurring-job.longhorn.io/source=enabled,recurring-job-group.longhorn.io/backup-nightly=enabled
  # Services without a Dockerfile are built from source with Railpack. The
  # CLI image comes from `make railpack-image`; keep both versions in step.
  RAILPACK_IMAGE: registry.example.lan:5000/rendimiento-railpack:0.40.0
  RAILPACK_FRONTEND: ghcr.io/railwayapp/railpack-frontend:v0.40.0
  # How a service with `gpu: 1` gets a GPU (here: an NVIDIA device plugin): its
  # runtime class, the host's driver libraries (read-only) and a larger /dev/shm.
  GPU_RESOURCE: nvidia.com/gpu
  GPU_RUNTIME_CLASS: nvidia
  GPU_HOST_PATHS: /usr/lib/aarch64-linux-gnu/nvidia
  GPU_ENV: NVIDIA_VISIBLE_DEVICES=all;NVIDIA_DRIVER_CAPABILITIES=all;LD_LIBRARY_PATH=/usr/lib/aarch64-linux-gnu/nvidia:/usr/local/cuda/lib64
  GPU_SHARED_MEMORY: 1Gi
  # Failed builds, rolled-back releases, outages and recoveries are emailed
  # here, through Resend (RESEND_API_KEY in the rendimiento-notify secret).
  NOTIFY_EMAIL_TO: [email protected]
  NOTIFY_LANG: en                # en | es (Mexican Spanish)
  NOTIFY_EMAIL_FROM: rendimiento <[email protected]>
  # GET /api/public/stats: aggregate numbers for a public page (a portfolio,
  # a status page), served only on the internal port 8081 (the Service's
  # "stats" port, which the public ingress does not route), so it is
  # reachable from inside the cluster but not from the internet.
  PUBLIC_STATS: "true"
  STATS_LISTEN: ":8081"
  PUBLIC_STATS_TZ: America/New_York
  # Finished runs' step logs move to S3-compatible storage (gzip; Garage
  # here), which deletes them after LOG_RETENTION_DAYS. Credentials: the
  # rendimiento-logs secret (a key that can only use this bucket).
  LOG_ARCHIVE_ENDPOINT: garage.garage.svc.cluster.local:3900
  LOG_ARCHIVE_BUCKET: rendimiento-logs
  LOG_RETENTION_DAYS: "365"
  # Visitor numbers from Umami (each app's Visits tab), read inside the
  # cluster with a view-only user (the rendimiento-umami secret).
  # UMAMI_URL: http://umami.umami.svc.cluster.local:3000
  # UMAMI_PUBLIC_URL: https://umami.example.com
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: rendimiento
  namespace: rendimiento-system
spec:
  replicas: 1
  strategy: { type: Recreate }
  selector:
    matchLabels: { app: rendimiento }
  template:
    metadata:
      labels: { app: rendimiento }
    spec:
      serviceAccountName: rendimiento
      containers:
        - name: rendimiento
          image: registry.example.lan:5000/rendimiento:latest
          imagePullPolicy: Always
          envFrom:
            - configMapRef: { name: rendimiento }
            # CLOUDFLARE_API_TOKEN and DNS_TARGET; optional until DNS automation is wanted.
            - secretRef: { name: rendimiento-dns, optional: true }
            # RESEND_API_KEY for notification emails; optional (no email without it).
            - secretRef: { name: rendimiento-notify, optional: true }
            # LOG_ARCHIVE_ACCESS_KEY / LOG_ARCHIVE_SECRET_KEY for the log archive.
            - secretRef: { name: rendimiento-logs, optional: true }
            # UMAMI_USERNAME / UMAMI_PASSWORD (a view-only Umami user) for visitor numbers.
            - secretRef: { name: rendimiento-umami, optional: true }
          env:
            - name: DB_PASSWORD
              valueFrom: { secretKeyRef: { name: rendimiento-db, key: password } }
            - name: DATABASE_URL
              value: postgres://rendimiento:$(DB_PASSWORD)@rendimiento-db:5432/rendimiento?sslmode=disable
            - name: SETUP_TOKEN
              valueFrom: { secretKeyRef: { name: rendimiento-setup, key: token } }
          ports:
            - { name: http, containerPort: 8080 }
            - { name: metrics, containerPort: 9090 }
            - { name: stats, containerPort: 8081 }
          readinessProbe:
            httpGet: { path: /healthz, port: http }
            periodSeconds: 10
          livenessProbe:
            httpGet: { path: /healthz, port: http }
            initialDelaySeconds: 20
            periodSeconds: 20
          resources:
            requests: { cpu: 100m, memory: 128Mi }
            limits: { memory: 512Mi }
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            runAsNonRoot: true
            runAsUser: 65532
            runAsGroup: 65532
            capabilities: { drop: [ALL] }
---
apiVersion: v1
kind: Service
metadata:
  name: rendimiento
  namespace: rendimiento-system
spec:
  selector: { app: rendimiento }
  ports:
    - { name: http, port: 80, targetPort: http }
    # In-cluster only: the public stats (the Ingress routes only "http").
    - { name: stats, port: 8081, targetPort: stats }
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: rendimiento
  namespace: rendimiento-system
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt-prod
    nginx.ingress.kubernetes.io/ssl-redirect: "true"
    nginx.ingress.kubernetes.io/force-ssl-redirect: "true"
    # Live logs use server-sent events.
    nginx.ingress.kubernetes.io/proxy-buffering: "off"
    nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
    nginx.ingress.kubernetes.io/proxy-body-size: 25m
spec:
  ingressClassName: nginx
  tls:
    - hosts: [rendimiento.example.com]
      secretName: rendimiento-tls
  rules:
    - host: rendimiento.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service: { name: rendimiento, port: { name: http } }
deploy/rbac.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: rendimiento
  namespace: rendimiento-system
---
# Cluster-wide because each app gets its own namespace. The controller
# refuses namespaces and hostnames it does not own (see checkOwnership).
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: rendimiento
rules:
  - apiGroups: [rendimiento.ai]
    resources: [apps, apps/status, apps/finalizers, addons, addons/status, addons/finalizers]
    verbs: ["*"]
  # Add-ons are applied as the rendimiento-addons identity (below); the
  # platform itself may only act as it, not hold its rights.
  - apiGroups: [""]
    resources: [serviceaccounts]
    resourceNames: [rendimiento-addons]
    verbs: [impersonate]
  - apiGroups: [""]
    resources: [namespaces, services, persistentvolumeclaims, secrets]
    verbs: [get, list, watch, create, update, patch, delete]
  - apiGroups: [""]
    resources: [pods]
    verbs: [get, list, watch]
  - apiGroups: [apps]
    resources: [deployments]
    verbs: [get, list, watch, create, update, patch, delete]
  - apiGroups: [networking.k8s.io]
    resources: [ingresses]
    verbs: [get, list, watch, create, update, patch, delete]
  - apiGroups: [batch]
    resources: [cronjobs]
    verbs: [get, list, watch, create, update, patch, delete]
  # Post-deploy tasks run as Jobs in the app's namespace; their logs are kept.
  - apiGroups: [batch]
    resources: [jobs]
    verbs: [get, list, watch, create, delete]
  - apiGroups: [""]
    resources: [pods/log]
    verbs: [get]
  - apiGroups: [cert-manager.io]
    resources: [certificates]
    verbs: [get, list, watch]
  - apiGroups: [""]
    resources: [events]
    verbs: [create, patch]
  # Read-only, for the environment page.
  - apiGroups: [""]
    resources: [nodes]
    verbs: [get, list]
  - apiGroups: [metrics.k8s.io]
    resources: [nodes]
    verbs: [get, list]
  - apiGroups: [networking.k8s.io]
    resources: [ingressclasses, networkpolicies]
    verbs: [get, list]
  - apiGroups: [storage.k8s.io]
    resources: [storageclasses]
    verbs: [get, list]
  - apiGroups: [apiextensions.k8s.io]
    resources: [customresourcedefinitions]
    verbs: [get, list]
  - apiGroups: [cert-manager.io]
    resources: [clusterissuers]
    verbs: [get, list]
  - apiGroups: [longhorn.io]
    resources: [volumes, recurringjobs, backuptargets]
    verbs: [get, list]
  # Read-only, for the services catalog (who calls what).
  - apiGroups: [apps]
    resources: [statefulsets]
    verbs: [get, list]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: rendimiento
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: rendimiento
subjects:
  - kind: ServiceAccount
    name: rendimiento
    namespace: rendimiento-system
---
# CI pods: create, watch, read logs, clean up.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: rendimiento-builds
  namespace: rendimiento-builds
rules:
  - apiGroups: [""]
    resources: [pods, secrets]
    verbs: [get, list, watch, create, delete]
  - apiGroups: [""]
    resources: [pods/log]
    verbs: [get]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: rendimiento-builds
  namespace: rendimiento-builds
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: rendimiento-builds
subjects:
  - kind: ServiceAccount
    name: rendimiento
    namespace: rendimiento-system
---
# Leader election.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: rendimiento-leader
  namespace: rendimiento-system
rules:
  - apiGroups: [coordination.k8s.io]
    resources: [leases]
    verbs: [get, list, watch, create, update, patch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: rendimiento-leader
  namespace: rendimiento-system
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: rendimiento-leader
subjects:
  - kind: ServiceAccount
    name: rendimiento
    namespace: rendimiento-system
---
# Add-ons install cluster software (Helm charts such as Longhorn: CRDs,
# ClusterRoles, DaemonSets), which needs cluster-admin, as ArgoCD has had.
# No pod runs as this account: the platform impersonates it only while
# applying add-ons, so the audit log shows exactly what add-ons changed.
apiVersion: v1
kind: ServiceAccount
metadata:
  name: rendimiento-addons
  namespace: rendimiento-system
automountServiceAccountToken: false
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: rendimiento-addons
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: cluster-admin
subjects:
  - kind: ServiceAccount
    name: rendimiento-addons
    namespace: rendimiento-system
deploy/networkpolicy.yaml
# CI steps run code from the repos being built (tests, Dockerfile RUN lines).
# Keep them away from everything else: they may reach DNS, the shared
# BuildKit daemon and the public internet (git, package registries), but
# not other apps, databases, the Kubernetes API or the home network.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: isolate-builds
  namespace: rendimiento-builds
  labels:
    app.kubernetes.io/managed-by: rendimiento
spec:
  podSelector: {}
  policyTypes: [Ingress, Egress]
  ingress: []   # nothing connects to build pods
  egress:
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: kube-system }
          podSelector:
            matchLabels: { k8s-app: kube-dns }
      ports:
        - { protocol: UDP, port: 53 }
        - { protocol: TCP, port: 53 }
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: devops-tools }
          podSelector:
            matchLabels: { app: buildkitd }
      ports:
        - { protocol: TCP, port: 1234 }
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
            except:
              - 10.0.0.0/8       # cluster pods and services (10.42/16, 10.43/16)
              - 172.16.0.0/12
              - 192.168.0.0/16   # home network, including the router
              - 169.254.0.0/16   # link-local / cloud metadata
              - 100.64.0.0/10    # carrier-grade NAT

Permissions, and why they are shaped this way

The platform's own service account (rendimiento) can:

  • manage the App and Addon resources;
  • manage namespaces, Deployments, Services, Ingresses, CronJobs, PVCs and Secrets cluster-wide, because every app gets its own namespace. The controller refuses namespaces and hostnames it does not own (ownership checks);
  • read nodes, metrics, ingress classes, storage classes, CRDs, cluster issuers and StatefulSets, for the Environment and Services pages;
  • create and delete pods and secrets in rendimiento-builds (CI);
  • impersonate the rendimiento-addons service account, and nothing more.

rendimiento-addons is bound to cluster-admin, because installing software like Longhorn means creating CRDs, ClusterRoles and DaemonSets, exactly what ArgoCD needed. No pod runs as it: only the add-on controller acts as it, so the Kubernetes audit log shows every add-on change under that name. See Security model.

Secrets you create once

Secret (namespace rendimiento-system) Keys Used for
rendimiento-db password Postgres password (the Deployment builds DATABASE_URL from it)
rendimiento-setup token One-time token that protects the GitHub App setup page
rendimiento-dns (optional) CLOUDFLARE_API_TOKEN, DNS_TARGET DNS automation. DNS_TARGET=auto follows the home network's public IP.
rendimiento-github created by the setup flow The GitHub App's ID, private key, webhook secret and OAuth client
kubectl -n rendimiento-system create secret generic rendimiento-db --from-literal=password="$(openssl rand -hex 24)"
kubectl -n rendimiento-system create secret generic rendimiento-setup --from-literal=token="$(openssl rand -hex 16)"
# a Cloudflare token with Zone:DNS:Edit on your zones, entered without echoing it:
read -rs CF && kubectl -n rendimiento-system create secret generic rendimiento-dns \
  --from-literal=CLOUDFLARE_API_TOKEN="$CF" --from-literal=DNS_TARGET=auto; unset CF

The GitHub App

rendimiento creates its own GitHub App with the manifest flow, so nobody fills in GitHub's forms by hand:

sequenceDiagram
    actor You
    participant R as rendimiento
    participant G as GitHub
    You->>R: open /api/setup/github?token=… (the rendimiento-setup token)
    R->>You: a page that posts the App manifest to GitHub
    You->>G: create the App (name, permissions, webhook URL)
    G->>R: redirect to /api/setup/github/callback?code=…
    R->>G: exchange the code for the App's credentials
    R->>R: store them in the rendimiento-github secret
    You->>G: install the App on your account (all or chosen repos)

The manifest asks for these repository permissions: Contents read & write (read code, push onboarding branches), Pull requests read & write (open onboarding PRs), Checks read & write (report CI status), Metadata read, Issues read & write and Commit statuses read (both for the Renovate add-on). It subscribes to push events. Login to the dashboard uses the same App's OAuth, and only users in ALLOWED_USERS get a session.

Changing the App's permissions later

Change them at https://github.com/settings/apps/<app-name>/permissions, then accept the new permissions on the installation (https://github.com/settings/installations → Configure). Until accepted, installations keep the old ones; the Add-ons page warns when Renovate lacks what it needs.

Shipping a new version of rendimiento

make test-remote     # the full test suite, on a worker node
make image           # build and push registry.example.lan:5000/rendimiento:latest on the BuildKit pool
make deploy          # kubectl apply -k deploy (only needed when deploy/ or the CRDs changed)
kubectl -n rendimiento-system rollout restart deploy/rendimiento
kubectl -n rendimiento-system rollout status deploy/rendimiento

The Deployment pulls :latest with imagePullPolicy: Always and uses the Recreate strategy (one replica, never two at once), so a restart takes about 20 seconds during which the UI and webhooks are unavailable. Apps keep running: they do not depend on the platform being up. GitHub retries webhooks that fail.

On restart the platform:

  1. applies database migrations (internal/store/migrations), each once;
  2. requeues CI runs a previous process left running (up to two retries) and deletes leftover build pods;
  3. closes Renovate runs that were interrupted;
  4. starts the controllers, the CI worker, the Renovate scheduler, the add-on sync from git, and dynamic DNS.

:latest has no history

Rolling back the platform means rebuilding the previous commit. Tagging each build with its commit is on the roadmap; until then, git checkout <good commit> && make image is the rollback.

rendimiento builds itself

rendimiento's repository is onboarded like any app. Its rendimiento.yaml has the book (docs, a service) and the platform's own image (builds: platform), so every push runs:

  • platform:test: hack/ci-test.sh (generate, vet and the Go tests, with envtest and a throwaway Postgres) in golang, with a kept cache;
  • platform:build: the Dockerfile, which also checks the UI's types and translations, pushed as <registry>/rendimiento-ai-platform:<commit>.

Pull requests get the same check without a release. On the default branch the image's digest is kept with the release, next to the book's. It is not deployed by itself yet: shipping is still the steps above, which build the same Dockerfile. Letting a release update the platform (a rolling update that keeps the old version until the new one is ready, so a broken image cannot take the platform down) is the next step.

A field the running platform doesn't know

The platform reads rendimiento.yaml strictly. A change that adds a field to it (as builds: did) must be shipped before the commit that uses the field is pushed, or that push's run fails to read its own spec.

The build side

What Where it comes from
BuildKit daemons the buildkit add-on (p0dxD/gitops/buildkit/): a StatefulSet with one daemon per worker node
Railpack CLI image make railpack-image from deploy/railpack/Dockerfile, checksum-pinned
Railpack frontend pulled by BuildKit from ghcr.io/railwayapp/railpack-frontend at the pinned version
Clone and build client images alpine/git, moby/buildkit (the client must match the daemon's version)