Sai Ram
AI Infrastructure and Platform Architect · Chennai, India
I build and run the platforms behind large-model inference: Kubernetes, GPU fleets, serving engines, autoscaling, observability, and the cost models that keep them honest. My background is platform and DevOps engineering across cloud infrastructure, CI/CD, and production operations.
Recent work
Inference Cost Optimization is a reproducible, open case study for deciding when a managed model API, a private vLLM service, or a policy-routed hybrid is the lower-cost choice after quality, latency, security, and reliability constraints are applied. Eight published runs so far, measured with Prometheus, Grafana, and DCGM, including autoscaling under load with KEDA. Every number in it carries a date, a source, and a comparison boundary.
Consulting
I take on short, fixed-scope engagements: inference cost assessments, serving stack reviews, and platform builds. If your GPU bill is growing faster than your traffic, write to me and include a line about your current setup.