Jobs / CaseWare International

Staff Site Reliability Engineer

CaseWare International · Toronto, ON, Canada
Toronto, ON, CanadaExp: 8+ yrs140,000-155,000 CAD/yearlyRemote
Remuneration
140,000-155,000 CAD/yearly
Location
Toronto, ON, Canada
Visa sponsorship
Not specified

Job summary

Caseware is seeking a Staff Site Reliability Engineer to enhance production resilience, security, and operational excellence. The role involves designing and evolving foundational systems and practices to support secure and scalable software delivery, while collaborating with various teams to optimize infrastructure and improve platform reliability.

Qualifications

  • 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.
  • Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3.
  • Advanced experience operating and scaling production Kubernetes environments.
  • Strong hands-on experience with Istio service mesh.
  • Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.
  • Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
  • Strong troubleshooting and incident management experience in distributed systems.
  • Excellent communication, collaboration, and technical leadership skills.

Responsibilities

  • Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.
  • Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems.
  • Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls.
  • Enhance platform observability through metrics, logs, tracing, and actionable alerting.
  • Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards.
  • Lead response efforts for high-severity incidents, ensuring timely resolution and effective communication.
  • Partner closely with engineering teams to strengthen platform standards and improve service resilience.
  • Mentor and guide engineers on cloud-native technologies and operational excellence practices.

Skills

AWSAWS CDKCloudFrontCloudWatchEKSGitHubGitHub ActionsIAMIstioKubernetesAWS LambdaNode.jsOpenTelemetryS3TypeScript

Relocation

No