Jobs / Cerebras Systems

Staff Site Reliability Engineer – Automation and Platform

Cerebras Systems · Toronto, ON, Canada
Toronto, ON, CanadaFull timeExp: 8+ yrsHybrid
Remuneration
Not specified
Location
Toronto, ON, Canada
Visa sponsorship
Not specified

Job summary

Cerebras Systems is seeking a Staff Site Reliability Engineer to lead the engineering effort in building a high-performance SRE function for AI inference services.

Qualifications

  • 8+ years in SRE, infrastructure engineering, or platform engineering.
  • Deep expertise operating large scale heterogeneous clusters.
  • Proven track record designing and delivering CI/CD or GitOps systems.
  • Hands-on experience with observability systems.
  • Ability to lead complex projects and communicate technical direction.

Responsibilities

  • Define and implement a strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
  • Architect self-service platforms and internal tooling for product teams and external customers.
  • Define and evolve reliability practices for inference workloads, including SLOs and SLIs.
  • Mentor mid-level SREs and support critical incident escalations.
  • Measure and drive impact through metrics like toil reduction and deployment velocity.

Skills

Argo CDBazelLokiMimirPrometheusTempo

Relocation

No