Jobs / Valtech Group

Site Reliability Expert

Valtech Group · Montréal, QC, Canada
Montréal, QC, Canada120,000-170,000 CAD/yearlyOnsite
Remuneration
120,000-170,000 CAD/yearly
Location
Montréal, QC, Canada
Visa sponsorship
Not specified

Job summary

Valtech is seeking an experienced Site Reliability Expert to lead observability, reliability, and operational excellence initiatives within complex cloud-native environments.

Qualifications

  • Significant experience in Site Reliability Engineering (SRE) in complex production environments.
  • Strong understanding of Service Level Indicators (SLI), Service Level Objectives (SLO), Error Budgets, and symptom-based alerting.
  • Proven expertise with observability platforms such as Dynatrace, Datadog, New Relic, AppDynamics.
  • Practical experience with Application Performance Monitoring (APM), Real User Monitoring (RUM), monitoring instrumentation, alert management, role-based access control (RBAC), SLO management, and tagging governance.
  • Good knowledge of OpenTelemetry (OTEL) and distributed tracing.
  • Experience in composable architectures based on microservices.
  • Practical experience in production environments with AWS and Kubernetes.
  • Experience in infrastructure and operations automation.
  • Proficiency in Terraform, Bash, Python.
  • Experience with GitLab CI or other CI/CD pipeline management tools.
  • Demonstrated ability to lead technical projects end-to-end.
  • Excellent communication skills in French and English, both oral and written.
  • Rigorous, autonomous, and highly organized with strong documentation skills.

Responsibilities

  • Define and implement observability strategy and governance standards.
  • Design, deploy, and maintain monitoring, alerting, and visualization solutions using Dynatrace or equivalent tools.
  • Establish and promote SRE best practices including SLI, SLO, Error Budgets, and symptom-based alerting.
  • Collaborate with development and product teams to enhance system reliability, performance, and resilience.
  • Set standards for tagging, governance, responsibility assignment, dashboards, and alert management.
  • Support and train teams lacking specific observability expertise.
  • Lead technical projects, manage priorities, and ensure timely delivery within defined budgets.
  • Analyze and resolve complex incidents in distributed microservices architectures.
  • Promote a culture of continuous improvement, documentation, and operational excellence.
  • Collaborate effectively with distributed teams in an international context.

Skills

AppDynamicsAWSBashDatadogDynatraceEnvoyGitLabGitLab CIJavaKubernetesNew RelicOpenTelemetryPythonTerraform

Languages

FrenchEnglish

Relocation

No