Jobs / Realign
Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)
Realign · Toronto, ON, Canada
Toronto, ON, CanadaExp: 3+ yrsHybrid
Remuneration
Not specified
Location
Toronto, ON, Canada
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer role is essential for maintaining and optimizing the production reliability of Azure and Databricks platforms. Responsibilities include monitoring platform performance, responding to incidents, and managing infrastructure components, utilizing tools such as Azure Monitor, Grafana, and JIRA to ensure operational readiness and support.
Qualifications
- 3+ years supporting Azure production cloud infrastructure.
- 1+ years of Windows Server and Linux administration.
- Experience with Azure Storage services, including ADLS Gen2.
- Understanding of networking and security concepts in Azure.
- Hands-on experience with monitoring tools.
- Experience with incident management and operational runbooks.
Responsibilities
- Monitor and support production Azure and Databricks environments.
- Respond to production incidents and participate in on-call support.
- Support Databricks workspaces and maintain integrations with Azure services.
- Troubleshoot Azure infrastructure components.
- Manage alerts and platform health using monitoring tools.
- Perform root cause analysis and system maintenance.
- Maintain operational runbooks and track tasks in JIRA and ServiceNow.
Skills
AzureAzure Key VaultAzure MonitorDatabricksDatadogDynatraceGrafanaJiraLinuxNew RelicPrometheusServiceNowVaultWindowsWindows Server
Relocation
No