Platform Reliability Engineer

Added
9 days ago
Type
Full time
Salary
Salary not provided

Related skills

datadog aws grafana splunk new relic
JobCopilot logo
Meet JobCopilot: Your Personal Al Job Hunter
Automatically Apply to Your Dream Jobs While You Sleep
Try it now →

📋 Description

  • Build OpenTelemetry instrumentation standards for agents, tools and platform services, and make it
  • Connect AgentCore Observability and CloudWatch with the client’s existing observability tools
  • Set up end-to-end tracing across agent steps, model calls, tool calls and multi-agent handoffs so
  • Define SLOs for latency, availability and error rates, and build alerting that is useful and not
  • Track token usage and compute spend by team, agent and environment, with dashboards, budgets and
  • Write runbooks, support incident response and run post-incident reviews that lead to real fixes.

🎯 Requirements

  • 6+ years in SRE, platform reliability or observability engineering, with strong hands-on AWS
  • Hands-on experience with OpenTelemetry (SDKs, collectors, exporters) and distributed tracing.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog
  • Experience defining and running SLOs, error budgets and alerting strategies.
  • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging
  • Scripting and coding skills in Python, Go or TypeScript.
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest — finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs →