Jobs / ServiceNow

Staff Software Engineer – SRE & AIOps

ServiceNow · Vancouver, BC, Canada
Vancouver, BC, CanadaExp: 8+ yrs125,700-220,000 CAD/yearlyHybrid
Remuneration
125,700-220,000 CAD/yearly
Location
Vancouver, BC, Canada
Visa sponsorship
Not specified

Job summary

The Staff Software Engineer – SRE & AIOps role is centered on enhancing infrastructure automation and operational resilience within hybrid cloud and data center operations. Key responsibilities include designing enterprise-scale Kubernetes clusters, implementing auto-remediation systems, and developing SRE tooling to ensure high availability and reduce operational toil across global teams.

Benefits

Health plansFlexible spending accounts401(k) Plan with company matchESPPMatching donationsFlexible time away planFamily leave programs

Qualifications

  • Strong hands-on expertise operating production Kubernetes clusters at scale
  • Proven experience designing and implementing closed-loop automated remediation systems
  • Extensive hands-on experience with AWS, Azure, and GCP
  • Strong experience with Infrastructure-as-Code tools and GitOps platforms
  • Strong working knowledge of observability platforms
  • Solid understanding of distributed system challenges
  • Experience operating in follow-the-sun, 24/7 on-call models
  • Hands-on experience managing both on-premises infrastructure and public cloud environments
  • Demonstrated ability to apply machine learning and AI-driven insights
  • Proven ability to drive technical decisions across teams

Responsibilities

  • Design, deploy, and operate enterprise-scale Kubernetes clusters
  • Architect and implement closed-loop auto-remediation systems
  • Design and evolve the SRE tooling stack
  • Establish SLO frameworks, error budgets, and alerting policies
  • Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines
  • Architect hybrid cloud and data center operations
  • Drive adoption of containerization, microservices, and DevOps patterns
  • Design on-call rotation schedules, escalation policies, and incident command systems
  • Mentor and guide junior SRE engineers and infrastructure teams
  • Champion a culture of blameless incident analysis

Skills

AKSAWSAzureBashCloud SQLCloudFormationEKSGCPGKEGoKubernetesAWS LambdaLinuxPythonServiceNowTerraform

Certifications

Kubernetes certification (CKA, CKAD, or equivalent)

Degrees

Bachelor's degree in computer science, Computer Engineering, or related field

Relocation

No