Added
2 hours ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

sre azure devops grafana prometheus

📋 Description

  • Monitor, maintain, and support production and non-production environments, ensuring availability
  • Implement comprehensive observability through alerting, dashboards, health checks, synthetic
  • Lead incident response and troubleshooting, including root cause analysis, defect resolution
  • Analyze application, API, AI model, data pipeline, and infrastructure metrics to identify
  • Support capacity planning, resource sizing, autoscaling, and cost optimization across compute
  • Maintain backup and restore processes, disaster recovery procedures, high-availability

🎯 Requirements

  • Bachelor’s degree plus 15 years of relevant experience in SRE, DevOps, platform engineering
  • Ability to obtain an active DHS/EOD clearance as required.
  • Extensive knowledge of SRE principles, including monitoring, observability, incident response
  • Strong expertise with Azure cloud services covering compute, storage, networking, monitoring, and
  • Hands-on experience with observability and monitoring technologies such as Azure Monitor
  • Proven ability to troubleshoot complex issues across application, platform, and infrastructure

🎁 Benefits

  • Full-time position with hybrid work options and remote eligibility across the United States.
  • Proposed national salary range of $114,600–$252,100, with final compensation influenced by
  • Comprehensive healthcare and wellness benefits.
  • Financial and retirement benefits.
  • Family support programs.
  • Flexible time-off benefits designed to support work-life balance.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →