Senior Site Reliability Engineer — Token Factory (Inference Platform)

Added
3 minutes ago
Location
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

devops terraform grafana prometheus python

📋 Description

  • Own the reliability, performance, and observability of the inference platform and its supporting
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large volumes of production
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience

🎯 Requirements

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and/or Bash.
  • Solid understanding of distributed systems and the ways production backends can fail under

🎁 Benefits

  • Competitive compensation.
  • Career growth and continuous learning opportunities.
  • Flexibility and significant ownership in your work.
  • Collaborative and innovative international working environment.
  • Opportunity to work on high-impact AI infrastructure and inference technologies.
  • Exposure to large-scale GPU infrastructure and complex distributed systems.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →