Jobs / Apple

Site Reliability Engineer (SRE),

Apple · London, ENG, United Kingdom
London, ENG, United KingdomRemote
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

Apple is seeking a hardworking and passionate Site Reliability Engineer (SRE) to join their team, responsible for the availability and automation of critical systems and services. The role involves deploying, supporting, and monitoring services while enhancing software to improve availability, scalability, and security.

Qualifications

  • Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model, Subnetting and Load Balancing strategies.
  • Understanding of the Linux Operating System, including Kernel, Memory, Process, Threads, Static / Shared Libraries, IPC, Signals.
  • Experience with Infrastructure-as-Code and config-as-code tooling such as Pulumi, Terraform, or Pkl.
  • Experience with fleet/cluster lifecycle management, node provisioning, and hardware-adjacent reliability (e.g., GPU health, capacity management) at scale.
  • Experience building and operating CI/CD pipelines for cloud infrastructure.
  • In depth hands-on experience operating managed Kubernetes (GKE and/or EKS) in a public cloud, with experience scaling distributed systems.
  • Strong experience with deploying, supporting and supervising new and existing services, platforms and application stacks.
  • Experience with scale testing, disaster recovery, and capacity planning.
  • Passion for eliminating repetitive manual processes using automation to improve them through repeated iteration.
  • Confirmed ability to write programs using a high-level programming language like: Java, Swift, Python, or TypeScript.
  • Proclivity towards efficient programming emphasizing improvement via complexity analysis.
  • Experience with Nginx, Envoy, Prometheus, and/or Docker.

Responsibilities

  • Deploy, support and monitor new and existing services, platforms, and application stacks.
  • Use scale testing to measure, tune and optimize system performance.
  • Enhance, architect, author, and deliver software to improve the availability, scalability and security of Apple's internet services.
  • Build and run systems, infrastructure and applications through automation.
  • Participate in periodic on-call duties.

Skills

AWSDockerEKSEnvoyGCPGKEJavaKubernetesLinuxNGINXPrometheusPulumiPythonTerraformTypeScript

Relocation

No