Jobs / Valtech Group

Site Reliability Expert

Valtech Group · Canada
Canada100,000-150,000 CAD/yearlyRemote
Remuneration
100,000-150,000 CAD/yearly
Location
Canada
Visa sponsorship
Not specified

Job summary

The Site Reliability Expert is responsible for leading observability, reliability, and operational excellence initiatives in complex cloud-native environments. Key responsibilities include defining observability strategies, designing monitoring solutions using Dynatrace, and establishing SRE best practices such as SLIs and SLOs while collaborating with engineering and product teams to enhance system reliability.

Benefits

Comprehensive insurance planVirtual healthcare servicesEmployee and Family Assistance ProgramMental health support program500 Personal Spending AccountRetirement plan with RRSP matchingFlexible vacation policyPersonal Technology Reimbursement

Qualifications

  • Significant experience in Site Reliability Engineering (SRE) within large-scale production environments.
  • Deep understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and Symptom-based Alerting.
  • Proven expertise with enterprise observability platforms such as Dynatrace, Datadog, New Relic, and AppDynamics.
  • Strong experience with Application Performance Monitoring (APM), Real User Monitoring (RUM), Monitoring agents and instrumentation, Alerting strategies, Role-Based Access Control (RBAC), SLO management, and Tagging and governance models.
  • Strong knowledge of OpenTelemetry (OTEL) and distributed tracing.
  • Experience working within composable, microservices-based architectures.
  • Hands-on production experience with AWS and Kubernetes.
  • Experience with infrastructure and operational automation.
  • Practical knowledge of Terraform, Bash scripting, and Python scripting.
  • Experience with CI/CD tools such as GitLab CI or equivalent pipeline/workflow platforms.
  • Demonstrated ability to lead technical initiatives and workstreams.
  • Experience working within complex operational and agile environments.
  • Strong stakeholder management and collaboration skills.
  • Excellent communication skills in both French and English.
  • Strong documentation practices, organizational skills, and attention to detail.
  • High degree of autonomy and ownership.

Responsibilities

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
  • Partner with engineering and product teams to improve system reliability, performance, and operational maturity.
  • Develop standards for tagging, ownership, dashboard design, access management, and alerting governance.
  • Support teams that are not specialized in observability by providing guidance, coaching, and knowledge transfer.
  • Lead technical workstreams, prioritize initiatives, and ensure successful delivery within defined timelines and budgets.
  • Analyze distributed systems and troubleshoot complex production issues using monitoring and tracing data.
  • Promote documentation, operational rigor, and continuous improvement across engineering teams.
  • Collaborate effectively within a distributed, multilingual environment.

Skills

AppDynamicsAWSBashDatadogDynatraceGitLabGitLab CIJavaKubernetesNew RelicOpenTelemetryPythonTerraform

Languages

FrenchEnglish

Relocation

No