Jobs / Heidi
Lead Site Reliability Engineer
Heidi · Sydney, NSW, Australia
Sydney, NSW, AustraliaExp: 7+ yrsRemote
Remuneration
Not specified
Location
Sydney, NSW, Australia
Visa sponsorship
Not specified
Job summary
The Lead Site Reliability Engineer will manage a small SRE team while actively participating in incident response, system reliability, and daily operations. Responsibilities include operating and enhancing Kubernetes clusters, managing cloud infrastructure, and improving observability through effective monitoring tools. The role demands extensive experience in production systems and strong skills in automation and alerting strategies.
Qualifications
- 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles
- A track record of hiring, coaching, and growing engineers
- Deep experience supporting production systems
- Strong experience operating cloud infrastructure at scale
- Solid hands-on experience with Kubernetes and containerized workloads in production
- Infrastructure as code experience
- Hands-on experience with monitoring and alerting tools
- Scripting or automation experience
- Hands-on experience defining and owning SLOs, error budgets, and capacity planning
Responsibilities
- Participate in on-call and incident response
- Improve operational reliability
- Own the production environment
- Strengthen observability
- Reduce operational toil
- Support safe change
- Contribute to operational practices
- Collaborate closely with engineers
- Lead and grow the SRE team
- Shape the team's direction
Skills
AWSBashDatadogKubernetesPrometheusPythonTerraform
Relocation
No