Senior Site Reliability Engineer — Token Factory (Inference Platform)

Added
1 day ago
Type
Full time
Salary
Salary not provided

Related skills

sre terraform bash grafana prometheus

📋 Description

  • Own reliability, performance, and observability of inference platform and supporting infrastructure.
  • Design, implement, and improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions for large volumes of production signals.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient
  • Tune Kubernetes autoscaling mechanisms to improve GPU resource efficiency.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns.

🎯 Requirements

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or related
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Advanced experience with Terraform and infrastructure-as-code.
  • Strong scripting and automation skills in Python and/or Bash.
  • Solid understanding of distributed systems and failure modes.

🎁 Benefits

  • Competitive compensation.
  • Career growth and continuous learning opportunities.
  • Flexibility and significant ownership in your work.
  • Collaborative and innovative international working environment.
  • Opportunity to work on high-impact AI infrastructure and inference technologies.
  • Exposure to large-scale GPU infrastructure and complex distributed systems.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →