Skip to content

Learn GPU Infrastructure Without Owning a GPU

AI Factory Operations Lab

Hands-on GPU Infrastructure Engineering

Build, break, troubleshoot and operate production-style AI infrastructure on your laptop. Then validate the hardware-specific parts on real GPUs.

💻 $0 to start 🚫 No GPU required 🟥 Optional real-GPU capstone

Start Learning View Learning Path

You do not read this course, you run it: stand things up, break them on purpose, diagnose them the way you would on a real cluster, and capture the evidence. Most of it needs no GPU at all; one optional session uses a single cheap rented GPU and is clearly marked.


How you learn here: Build, Break, Diagnose, Prove

Every lesson runs the same loop, the one an operator actually lives in:

  • Build

    Stand up the system or capability: a fake GPU fleet, a scheduler, a serving stack.

  • Break

    Introduce a realistic failure, capacity limit, or queue condition, on purpose.

  • Diagnose

    Read the same signals and tools an operator uses in production to find the cause.

  • Prove

    Capture the evidence of what happened and what you fixed. A lesson is "done" only when its evidence exists, not when a command runs.


What it costs

Tier Lessons You pay You get
$0 simulation 1 through 5 (with 1B/1C/1D, 3B, 4A/4B) Nothing, a laptop runs it Scheduling, queueing, gang scheduling, GPU-sharing decisions, observability design, inference and capacity, lifecycle. Most of the course
$5-10 one-GPU capstone 6 A few hours on one entry-level GPU VM The real runtime path, enforced GPU sharing, real telemetry and inference benchmarks

Every lesson declares its mode and states exactly what it proves and what it does not:

  • 🟦 Simulation (no GPU). kind + KWOK fake nodes, the fake-gpu-operator, Slurm with fake GRES, synthetic DCGM and vLLM metrics. Proves control-plane behaviour: scheduling, queueing, sharing decisions, observability design. Nothing below the kubelet.
  • 🟥 Real GPU (one cheap NVIDIA GPU). Real driver, container toolkit, CUDA pod, DCGM telemetry, enforced GPU sharing, real inference numbers. Proves the runtime path, single node.

Knowing exactly where that line sits is itself one of the skills this course teaches.


The lessons

New here? Start Here first for orientation, then work the lessons in order. Each card leads with what you will be able to do.

  • 1 · Diagnose why a GPU pod stays Pending


    🟦 Simulation · No GPU · Beginner · ~30-45 min (est)

    Build a fake GPU fleet with kind + KWOK and the fake-gpu-operator, then walk the driver-to-pod path to find why work will not schedule.

    Open

  • 1B · Control GPU access between teams


    🟦 Simulation · No GPU · Intermediate · ~30 min (est)

    NVIDIA KAI Scheduler · queue quotas · borrowing · gang scheduling, enforced on a fake fleet.

    Open

  • 1C · Share one GPU between pods


    🟦 Simulation (+🟥 real half in 6) · Intermediate · ~30 min (est)

    HAMi fractional GPUs: prove the scheduling decision on fakes; the enforced memory slice is proven on a real GPU in the capstone.

    Open

  • 1D · Watch gang scheduling refuse a job


    🟦 Simulation · No GPU · Intermediate · ~30 min (est)

    Volcano · topology-driven fake fleets · Queue / PodGroup gang scheduling, all-or-nothing.

    Open

  • 2 · Schedule GPU jobs on an HPC cluster


    🟦 Simulation · No GPU · Intermediate · ~40 min (est)

    Slurm-in-Docker with fake GRES · GPU jobs · QoS caps · queue pressure · drain and resume.

    Open

  • 3 · See a GPU fleet and trip its alerts


    🟦 Simulation · No GPU · Beginner · ~20 min

    Prometheus · Grafana · synthetic DCGM · build dashboards and break them on purpose.

    Open

  • 3B · Alert on token-level inference SLOs


    🟦 Simulation · No GPU · Intermediate · ~25 min (est)

    Synthetic vLLM metrics · TTFT / goodput / queue depth / KV usage · SLO alerts that fire.

    Open

  • 4A · Find where a serving stack saturates


    🟦 Simulation (CPU) · No GPU · Intermediate · ~30 min (est)

    A load harness for TTFT · TPOT · p95/p99 · tokens/sec · goodput, and the saturation knee.

    Open

  • 4B · Size GPU memory to concurrent users


    🟦 Simulation · No GPU · Intermediate · ~25 min (est)

    The KV cache calculator · PagedAttention · GQA · prefix caching · the concurrency limit.

    Open

  • 5 · Run a node from provision to retire


    🟨 Concept + drill · No GPU · Intermediate · ~25 min (est)

    A provision to health-gate to patch to retire lifecycle drill, mapped to BCM-style ops.

    Open

  • 6 · The AI Factory Operator Capstone


    🟥 Real GPU · ~$5-10 · Advanced · ~1-2 hours

    One rented GPU. Prove what simulation cannot: the runtime path + real DCGM, enforced HAMi sharing, HAMi with the GPU Operator, and a real inference benchmark. Then tear it down.

    Open


Run the first loop

git clone https://github.com/ld-singh/ai-factory-ops-lab
cd ai-factory-ops-lab
make check          # verify docker, kind, kubectl, helm, jq
make phase1-up      # kind cluster + KWOK + fake GPU node pools
make phase1-demo    # schedulable + intentionally-Pending GPU workloads
make phase1-down    # tear it down

New to the course? Start Here explains the prerequisites, the modes, how evidence works, and exactly what to run first.


⭐ Finding this useful? Star it on GitHub so other engineers find the course.

Built by Lovedeep Singh, Cloud Infrastructure Architect (AWS, Azure, Kubernetes & DevSecOps), building secure, governed cloud platforms. See About for more.