Where I forge my DevOps and infrastructure skills through hands-on projects.
A personal lab for building, breaking, and learning infrastructure — starting on a single real VPS.
OpsForge has two layers:
| Layer | Where | What it is |
|---|---|---|
| Journal | labs/ |
One folder per lab: knowledge.md (concepts to learn first) and README.md (goal, steps, what broke, what I learned). |
| Infrastructure | ansible/, terraform/, stacks/, k8s/, apps/, scripts/ |
The real, reusable code that runs my server. Labs produce this code; later labs replace earlier manual work with it. |
Rule of thumb: if I'd want to run it again next month, it belongs in an infrastructure folder, and the lab links to it. See docs/conventions.md.
Learning path: docs/learning-roadmap.md is the curriculum behind the labs: 14 skill domains (plus service mesh and data tracks) in levels, with checkpoints, resources, certifications, fundamentals and career skills, and a 21-month plan.
Status: ⬜ todo · 🟨 in progress · ✅ done
| # | Lab | Status |
|---|---|---|
| 01 | VPS hardening | 🟨 |
| 02 | Nginx, SSL & reverse proxy | 🟨 |
| 03 | Automated backup & restore | ⬜ |
| # | Lab | Status |
|---|---|---|
| 04 | Dockerize an app | 🟨 |
| 05 | Multi-service reverse proxy | 🟨 |
| 06 | Self-hosted tools | 🟨 |
| # | Lab | Status |
|---|---|---|
| 07 | CI/CD pipeline | 🟨 |
| 08 | Zero-downtime deployment | 🟨 |
| # | Lab | Status |
|---|---|---|
| 09 | Ansible: rebuild the server in one command | 🟨 |
| 10 | Terraform: provision VPS, DNS, firewall | 🟨 |
| # | Lab | Status |
|---|---|---|
| 11 | Observability stack | 🟨 |
| 12 | K3s, Helm & GitOps | 🟨 |
| # | Lab | Status |
|---|---|---|
| 13 | Network fundamentals | ⬜ |
| 14 | Container networking by hand | ⬜ |
| 15 | DNS deep dive | ⬜ |
| 16 | WireGuard VPN and private access | ⬜ |
| # | Lab | Status |
|---|---|---|
| 17 | Processes, signals, and systemd | ⬜ |
| 18 | Containers from scratch | ⬜ |
| 19 | Storage and filesystems | ⬜ |
| 20 | Performance troubleshooting | ⬜ |
| 21 | Linux security hardening | ⬜ |
| # | Lab | Status |
|---|---|---|
| 22 | Load testing and capacity planning | ⬜ |
| 23 | Horizontal scaling and autoscaling | ⬜ |
| 24 | Caching and database scaling | ⬜ |
| 25 | Resilience and chaos engineering | ⬜ |
Phases 6–8 deepen the foundations: each lab lists what it builds on, and all of them run locally (fake VPS, Docker, or k3d).
How the GCP projects are structured: ADR 0004.
Labs 33–38 run free on k3d; 39–40 need real VMs (GCE, lab 29); 31 and 42 use GKE.
| # | Lab | Status |
|---|---|---|
| 43 | Why a service mesh? | ⬜ |
| 44 | Envoy by hand | ⬜ |
| 45 | Linkerd: a mesh in 15 minutes | ⬜ |
| 46 | Istio basics | ⬜ |
| 47 | Traffic management | ⬜ |
| 48 | Zero-trust security with a mesh | ⬜ |
| 49 | Mesh observability and distributed tracing | ⬜ |
| 50 | Sidecarless meshes and running in production | ⬜ |
Demo app: frontend → hello-api + quotes (v1/v2), all on k3d. Learn the concepts with Linkerd, go deep with Istio, then compare sidecarless options (Istio ambient, Cilium).
| # | Lab | Status |
|---|---|---|
| 51 | Operating data services: the playbook | ⬜ |
| 52 | PostgreSQL in depth | ⬜ |
| 53 | MySQL | ⬜ |
| 54 | Redis and Valkey | ⬜ |
| 55 | MongoDB | ⬜ |
| 56 | RabbitMQ | ⬜ |
| 57 | Apache Kafka | ⬜ |
| 58 | Event-driven OpsForge | ⬜ |
| 59 | Schema migrations and disaster recovery | ⬜ |
Each data service is run with Docker Compose first, then on k3d with its operator, through the same checklist: deploy, HA, backup + restore, monitoring, upgrades, security. Run one at a time: Kafka and MongoDB clusters need a few GB of RAM each.
New labs get the next number (60, 61, …) and a new phase heading if needed: DevSecOps and SRE (phases 13–14, planned in the learning roadmap), or whatever comes next.
OpsForge/
├── labs/ # one folder per lab: knowledge.md (learn) + README.md (practice)
├── apps/ # sample applications deployed during the labs
├── stacks/ # Docker Compose stacks (proxy, self-hosted tools, monitoring)
├── scripts/ # standalone shell scripts (backup, helpers)
├── ansible/ # server configuration as code
├── terraform/ # cloud resources as code
├── k8s/ # Kubernetes manifests, Helm charts, GitOps apps
├── .github/workflows/ # CI/CD pipelines
└── docs/
├── learning-roadmap.md # curriculum: domains, levels, checkpoints, plan
├── architecture.md # what currently runs on the server
├── conventions.md # naming, secrets, workflow rules
└── decisions/ # why I chose X over Y (ADRs)
- Server: 1 VPS, 2 vCPU / 4 GB RAM (enough for every lab, including K3s)
- Provider / OS / domain: see docs/architecture.md