lukewilkinson.io
← Projects

case study

A homelab run like production, including the outage

Twenty-two Proxmox guests declared in Terraform and deployed by GitHub Actions from encrypted, saved plans. One apply from an open pull request destroyed a production VM, and the pipeline was rebuilt around that.

Visit github.com ↗

What runs

One Proxmox node runs twenty-two VMs and containers, split across separate VLANs for lab, application, service, and client workloads:

  • a private gateway (Caddy with a Homepage portal)
  • two private web app hosts
  • a small multiplayer app with separate dev and prod environments
  • Plex and its NFS storage, each in dev and prod
  • two self-hosted AI assistant VMs, a CTF workstation, and a Windows 11 VM with nested virtualisation
  • a pool of nine self-hosted GitHub Actions runners (one controller, eight workers) that deploy everything else

Each workload is its own Terraform deployment with its own state, so a media-server change never shares a plan with a lab VM. The few manual steps, such as a Windows install and one hand-built tunnel, are called out in the code as exceptions.

How a VM gets built

  1. Stateless hosts use a versioned module, pm-cloudinit-vm, pinned by git tag. It renders a cloud-init template for the host, uploads it to Proxmox as a snippet, and full-clones the VM from an Ubuntu 24.04 template that a script keeps current.
  2. On first boot, cloud-init sets the host up and starts the workload with Docker Compose.
  3. Hosts whose lifecycle needs care, such as the multiplayer app, the app servers, and the runners, use raw resources instead, so their lifecycle rules are visible right where they’re declared.

The module has seventeen terraform test cases that run against a mock Proxmox provider, so a module change is tested before any deployment moves to the new tag. State lives in HCP Terraform, and the HCP projects (Lab, Dev, Prd, Platform) and their workspaces are themselves managed as code.

The pipeline

One deploy workflow works out which deployments a change touches and fans out a matrix over two reusable workflows:

  • Plan on the pull request. It runs init, validate, and plan, then posts a comment listing every resource to be created, updated, or destroyed. Because the repository is public, the saved plan is encrypted with AES-256 before it is uploaded as an artifact.
  • Apply on merge. Merging applies the pull request’s own saved plan, not a fresh plan from main. If the plan has gone stale, the apply fails instead of silently doing something nobody reviewed.

Secrets never sit in the repository. Config files hold only Bitwarden Secrets Manager IDs, and CI resolves the values at run time. Pre-commit runs terraform fmt and validate, tflint, Trivy, shellcheck, shfmt, prettier, and Python formatters on every change.

Public services sit behind a Cloudflare Zero Trust tunnel with TLS terminated at Cloudflare, so nothing needs an inbound port. The private gateway uses a wildcard certificate from KrakenKey, my own certificate automation, and a systemd timer renews it.

The outage

The apply job originally had one guard: it treated environments with prd in the name as production. A production VM lived in a workspace that didn’t match, so on 2026-09-10 an apply from an open pull request destroyed it along with its data disk.

What changed:

  • Apply only from main. Pull requests plan and comment, and nothing more. A manual run of the deploy workflow still exists, but the runbook documents it as break-glass.
  • Production refuses to be destroyed. The rebuilt VM carries prevent_destroy, so a plan that would delete it fails.
  • Replacement is deliberate. Rebuilding a single resource has its own manual workflow. Its inputs reach the shell through environment variables, never inline interpolation, so a crafted input can’t inject commands.
  • Dead workflows are deleted, not disabled. A disabled workflow is one click away from an ungated apply.
  • Data has a copy off the box. The app’s data is backed up with restic to Cloudflare R2, one bucket per environment. Retention is restic’s own forget --prune, not a bucket expiry rule that could delete data restic still references.

Deploying the runners that deploy everything

The runner pool can’t safely deploy itself, so its workflow is manual and defaults to a dry run. It mints and masks registration tokens. It refuses any plan that deletes or replaces a runner, or that reuses a name already registered. It always runs on the opposite role: workers are deployed from the controller, and the controller from a worker.

Tradeoffs worth naming

  • Stateful VMs ignore changes to their cloud-init. Editing a template used to force Terraform to replace the VM, so stateful hosts now carry ignore_changes on initialization. Template edits no longer reach running hosts. That’s the price of not rebuilding a database to fix a typo.
  • Modules remove repetition, not judgment. When a guest’s lifecycle needs to be explicit, it gets a raw resource rather than another module flag.
  • The app’s dev environment is sized like prod, so a rehearsal tells you something.

Not done yet

A Kubernetes cluster used to run here and was retired; its bootstrap scripts are kept for reference. Configuration management beyond cloud-init, monitoring, and backups for anything beyond the app’s data are listed in the repository as planned, not presented as built.