Senior Site Reliability Engineer (Noida, BLR, India)

Added
3 hours ago
Type
Full time
Salary
Salary not provided

Related skills

rust terraform python kubernetes gcp
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now โ†’

๐Ÿ“‹ Description

  • Senior SRE role blending backend engineering, infra ops, and FinOps
  • Own cost efficiency, Kubernetes right-sizing, and cost telemetry for backend teams
  • Run GPU throughput experiments on on-prem clusters with AI service owners
  • Build tooling, dashboards, and processes for teams to own cost and reliability budgets
  • Improve reliability instrumentation across new and offline flows
  • Contribute to security-focused platform changes in collaboration with DevOps

๐ŸŽฏ Requirements

  • 4-5 years hands-on systems experience
  • Backend depth: Python, Go/Rust; end-to-end service ownership
  • Kubernetes at scale: scheduler, resources, HPA/VPA, node pools, autoscaling
  • Cloud & on-prem infra: GCP, IaC (Terraform), CI/CD, hybrid on-prem GPU clusters
  • GPU workload understanding: throughput profiling, batching, cache behavior, inference tuning
  • Observability: metrics, traces, logs, SLOs; proper instrumentation

๐ŸŽ Benefits

  • Appropriate compensation and equity per level and market
  • Collaborative, growth-focused team in a world-class AI company
  • Exposure to cutting-edge GPU and on-prem hybrid infrastructure
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest โ€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs โ†’