Jobs / Newfold Digital
Senior DevOps Engineer, AI Platform
Newfold Digital · Canada
CanadaExp: 7+ yrsRemote
Remuneration
Not specified
Location
Canada
Visa sponsorship
Not specified
Job summary
The Senior DevOps Engineer is responsible for building and operating the infrastructure for AI platforms, web applications, and backend services. This role involves designing and managing Kubernetes environments, supporting AI workloads, and implementing CI/CD pipelines using tools like Jenkins and Bitbucket, primarily within Microsoft Azure and Oracle Cloud Infrastructure.
Qualifications
- 7 or more years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related role.
- Strong hands on experience operating production Kubernetes environments and deep knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
- Strong Microsoft Azure experience, including AKS, networking, identity, storage, and monitoring.
- Strong cloud networking knowledge across virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
- Strong experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
- Proven experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
- Hands on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies.
- Experience implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools.
Responsibilities
- Translate application and platform technical designs into production ready cloud infrastructure with minimal supervision.
- Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
- Support AI workloads including LiteLLM based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
- Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service to service communication.
- Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event driven workloads.
- Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
- Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
- Implement end to end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
- Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
- Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.
Skills
AKSArgo CDAzureBashBitbucketCloudflareC#DockerEnvoyGoGrafanaHelmJavaJavaScriptJenkinsKubernetesLinuxOpenTelemetryOracle CloudPostgreSQLPrometheusPythonRabbitMQRedisRESTSentryCloud MonitoringTerraformTypeScript
Languages
PythonC#JavaGoJavaScriptTypeScript
Relocation
No