Jobs / Heidi

Lead Site Reliability Engineer

Heidi · Sydney, NSW, Australia
Sydney, NSW, AustraliaExp: 7+ yrsRemote
Remuneration
Not specified
Location
Sydney, NSW, Australia
Visa sponsorship
Not specified

Job summary

The Lead Site Reliability Engineer will manage a small SRE team while actively participating in incident response, system reliability, and daily operations. Responsibilities include operating and enhancing Kubernetes clusters, managing cloud infrastructure, and improving observability through effective monitoring tools. The role demands extensive experience in production systems and strong skills in automation and alerting strategies.

Qualifications

  • 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles
  • A track record of hiring, coaching, and growing engineers
  • Deep experience supporting production systems
  • Strong experience operating cloud infrastructure at scale
  • Solid hands-on experience with Kubernetes and containerized workloads in production
  • Infrastructure as code experience
  • Hands-on experience with monitoring and alerting tools
  • Scripting or automation experience
  • Hands-on experience defining and owning SLOs, error budgets, and capacity planning

Responsibilities

  • Participate in on-call and incident response
  • Improve operational reliability
  • Own the production environment
  • Strengthen observability
  • Reduce operational toil
  • Support safe change
  • Contribute to operational practices
  • Collaborate closely with engineers
  • Lead and grow the SRE team
  • Shape the team's direction

Skills

AWSBashDatadogKubernetesPrometheusPythonTerraform

Relocation

No