Jobs / Anna

Site Reliability Engineer

Anna · Huntsville, AL, United States
Huntsville, AL, United StatesExp: 3-5 yrs130,000-160,000 USD/yearlyOnsite
Remuneration
130,000-160,000 USD/yearly
Location
Huntsville, AL, United States
Visa sponsorship
Not specified

Job summary

Enhance operational maturity and incident response capabilities for a customer experience technology platform.

Qualifications

  • Strong Kubernetes operational experience
  • Experience with monitoring and alerting platforms
  • Familiarity with incident response processes
  • Ability to read and understand Terraform and Helm charts
  • Comfortable with GitOps workflows
  • Strong written communication

Responsibilities

  • Triage and respond to escalated platform issues
  • Develop and maintain operational runbooks and playbooks
  • Own the monitoring and alerting stack
  • Track and improve MTTR, availability, and change failure rate metrics
  • Define and propose SLOs for client-facing services
  • Operate EKS clusters day-to-day
  • Monitor cross-account drift as new tenant accounts come online
  • Document operational procedures and contribute to the team knowledge base
  • Mentor junior team members
  • Contribute to post-incident reviews

Skills

Argo CDDatadogDynamoDBEKSFluxGitHubGrafanaHelmKafkaKnativeKubernetesPrometheusTerraformTerraform Cloud

Relocation

No