Senior Site Reliability Engineer — Token Factory (Inference Platform)
Design and maintain telemetry pipelines for metrics, logs, and traces at scale, while optimizing Kubernetes and Terraform configurations for GPU-intensive inference workloads. Focus on building self-healing, observable systems that ensure high reliability and performance across a distributed AI infrastructure. Collaborate with engineering teams to harden request routing, autoscaling, and incident response mechanisms for large-scale model deployment.