Jobs / Cerebras Systems
Staff Site Reliability Engineer – Automation and Platform
Cerebras Systems · Toronto, ON, Canada
Toronto, ON, CanadaFull timeExp: 8+ yrsHybrid
Remuneration
Not specified
Location
Toronto, ON, Canada
Visa sponsorship
Not specified
Job summary
Cerebras Systems is seeking a Staff Site Reliability Engineer to lead the engineering effort in building a high-performance SRE function for AI inference services.
Qualifications
- 8+ years in SRE, infrastructure engineering, or platform engineering.
- Deep expertise operating large scale heterogeneous clusters.
- Proven track record designing and delivering CI/CD or GitOps systems.
- Hands-on experience with observability systems.
- Ability to lead complex projects and communicate technical direction.
Responsibilities
- Define and implement a strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
- Architect self-service platforms and internal tooling for product teams and external customers.
- Define and evolve reliability practices for inference workloads, including SLOs and SLIs.
- Mentor mid-level SREs and support critical incident escalations.
- Measure and drive impact through metrics like toil reduction and deployment velocity.
Skills
Argo CDBazelLokiMimirPrometheusTempo
Relocation
No