Jobs / Akamai
Site Reliability Engineer (Guardicore AI Platform), Madrid
Akamai · Madrid, MD, Spain
Madrid, MD, SpainExp: 3+ yrsRemote
Remuneration
Not specified
Location
Madrid, MD, Spain
Visa sponsorship
Not specified
Job summary
Take technical ownership of the reliability, availability, performance, and operational readiness of the Guardicore Data and AI Platform.
Qualifications
- 3+ years in SRE, DevOps, or Platform Engineering
- Ability to design and implement monitoring strategies using Prometheus and Grafana
- Production experience with Kubernetes, Docker, Helm, and cloud services
- Exceptional troubleshooting across network, system, applications, and database layers
- Experience with GitOps, CI/CD, and Infrastructure as Code
- Proficiency in Python, Go, and Bash
- Leverage AI tools for operational tasks and propose automation initiatives
- Demonstrate technical leadership in defining tools and frameworks
Responsibilities
- Operate secure, highly available Kubernetes infrastructure
- Enhance platform reliability, observability, security, performance, and cost efficiency
- Guide engineers and developers on service performance
- Lead complex production investigations and drive improvements
- Leverage AI-driven automation for incident remediation
- Collaborate with DevOps, Software, Data, AI, and Security teams
- Participate in on-call rotations for service restoration
Skills
AkamaiAWSAzureBashDockerGCPGoGrafanaHelmKubernetesLinodeLinuxPrometheusPython
Relocation
No