Remuneration
Not specified
Location
Canada
Visa sponsorship
Not specified
Job summary
Nexxa is seeking a Senior/Staff DevOps Engineer to build and operate infrastructure for AI and industrial systems, ensuring reliable production systems and managing CI/CD pipelines.
Qualifications
- 6+ years in DevOps, Site Reliability Engineering, or related roles
- Hands-on experience with cloud platforms at production scale
- Kubernetes in production, including GPU workload scheduling
- Infrastructure-as-code tooling experience
- CI/CD systems experience
- Track record designing and operating observability stacks
- Experience supporting ML/AI infrastructure
- Ability to scope and lead infrastructure projects
- Experience operating infrastructure bridging cloud and edge environments
Responsibilities
- Own and evolve core infrastructure end-to-end
- Design and operate CI/CD pipelines for fast, safe iteration
- Build and maintain infrastructure-as-code for reproducible environments
- Architect and manage Kubernetes-based platforms for workloads
- Support infrastructure for data warehouses and lakehouse architectures
- Define and drive observability practices across distributed systems
- Establish and enforce reliability practices
- Design for security and compliance across cloud infrastructure
- Collaborate with leadership to define infrastructure roadmap
- Mentor engineers on infrastructure best practices
Skills
Argo CDAWSAzureBashBigQueryCircleCIDatabricksDatadogGCPGitHubGitHub ActionsGitLabGitLab CIGoGrafanaJenkinsKubernetesMakeOpenTelemetryPrometheusPulumiPythonRedshiftSnowflakeTerraform
Relocation
No