Building the InfraForge AI Inference Lab
How we engineered a zero-to-edge inference cluster for 99.95% availability using k3s and vLLM — including cost-aware autoscaling and multi-tenant model routing.
Engineering insights distilled to essentials. Each piece includes runnable assets you can copy-paste.
How we engineered a zero-to-edge inference cluster for 99.95% availability using k3s and vLLM — including cost-aware autoscaling and multi-tenant model routing.
From the trenches: production Go binaries corrupted by Helm upgrade, and how we engineered image validation gates and per-namespace CI-driven deployments that outlived the team that shipped them.
GitHub shared tenancy at scale. When $216K/year uptime for 1M+ monthly API requests is on the line, reliability becomes a daily practice.
New essays weekly. Stay tuned.