Remuneration
Not specified
Location
United States
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer role focuses on evolving and maintaining infrastructure for AI agents and distributed systems, utilizing Kubernetes, Go, and Terraform. Key responsibilities include operating Kubernetes controllers, managing Postgres databases, and ensuring observability at scale with Prometheus and OpenTelemetry, while contributing to architecture decisions and standards for deployment and security.
Qualifications
- production experience across AWS and Azure
- operating Kubernetes controllers in production
- infrastructure as code with real production work
- operating managed Postgres in production
- observability at scale with Prometheus
- securing Kubernetes clusters
Responsibilities
- help run and evolve infrastructure
- contribute to architecture decisions for deployment, observation, and security
- shape team standards
Skills
AWSAWS KMSAzureCloud SQLCortexFluxGCPGoKafkaKubernetesKustomizeLinkerdMimirOpenTelemetryPostgreSQLPrometheusPub/SubTerraformThanos
Relocation
No