Inference platform: a sanitized case study
How a GitOps-managed vLLM inference platform actually worked: one OpenAI-compatible router in front of per-GPU engine profiles, hot swaps with readiness gates, streaming passthrough, token-accurate telemetry, and an offline model cache.