Jobs / Valtech Group
Site Reliability Expert
Valtech Group · Montréal, QC, Canada
Montréal, QC, Canada120,000-170,000 CAD/yearlyOnsite
Remuneration
120,000-170,000 CAD/yearly
Location
Montréal, QC, Canada
Visa sponsorship
Not specified
Job summary
Valtech is seeking an experienced Site Reliability Expert to lead observability, reliability, and operational excellence initiatives within complex cloud-native environments.
Qualifications
- Significant experience in Site Reliability Engineering (SRE) in complex production environments.
- Strong understanding of Service Level Indicators (SLI), Service Level Objectives (SLO), Error Budgets, and symptom-based alerting.
- Proven expertise with observability platforms such as Dynatrace, Datadog, New Relic, AppDynamics.
- Practical experience with Application Performance Monitoring (APM), Real User Monitoring (RUM), monitoring instrumentation, alert management, role-based access control (RBAC), SLO management, and tagging governance.
- Good knowledge of OpenTelemetry (OTEL) and distributed tracing.
- Experience in composable architectures based on microservices.
- Practical experience in production environments with AWS and Kubernetes.
- Experience in infrastructure and operations automation.
- Proficiency in Terraform, Bash, Python.
- Experience with GitLab CI or other CI/CD pipeline management tools.
- Demonstrated ability to lead technical projects end-to-end.
- Excellent communication skills in French and English, both oral and written.
- Rigorous, autonomous, and highly organized with strong documentation skills.
Responsibilities
- Define and implement observability strategy and governance standards.
- Design, deploy, and maintain monitoring, alerting, and visualization solutions using Dynatrace or equivalent tools.
- Establish and promote SRE best practices including SLI, SLO, Error Budgets, and symptom-based alerting.
- Collaborate with development and product teams to enhance system reliability, performance, and resilience.
- Set standards for tagging, governance, responsibility assignment, dashboards, and alert management.
- Support and train teams lacking specific observability expertise.
- Lead technical projects, manage priorities, and ensure timely delivery within defined budgets.
- Analyze and resolve complex incidents in distributed microservices architectures.
- Promote a culture of continuous improvement, documentation, and operational excellence.
- Collaborate effectively with distributed teams in an international context.
Skills
AppDynamicsAWSBashDatadogDynatraceEnvoyGitLabGitLab CIJavaKubernetesNew RelicOpenTelemetryPythonTerraform
Languages
FrenchEnglish
Relocation
No