Jobs / Anna
Site Reliability Engineer
Anna · Huntsville, AL, United States
Huntsville, AL, United StatesExp: 3-5 yrs130,000-160,000 USD/yearlyOnsite
Remuneration
130,000-160,000 USD/yearly
Location
Huntsville, AL, United States
Visa sponsorship
Not specified
Job summary
Enhance operational maturity and incident response capabilities for a customer experience technology platform.
Qualifications
- Strong Kubernetes operational experience
- Experience with monitoring and alerting platforms
- Familiarity with incident response processes
- Ability to read and understand Terraform and Helm charts
- Comfortable with GitOps workflows
- Strong written communication
Responsibilities
- Triage and respond to escalated platform issues
- Develop and maintain operational runbooks and playbooks
- Own the monitoring and alerting stack
- Track and improve MTTR, availability, and change failure rate metrics
- Define and propose SLOs for client-facing services
- Operate EKS clusters day-to-day
- Monitor cross-account drift as new tenant accounts come online
- Document operational procedures and contribute to the team knowledge base
- Mentor junior team members
- Contribute to post-incident reviews
Skills
Argo CDDatadogDynamoDBEKSFluxGitHubGrafanaHelmKafkaKnativeKubernetesPrometheusTerraformTerraform Cloud
Relocation
No