Jobs / Future Secure AI

Site Reliability Engineer

Future Secure AI · Toronto, ON, Canada
Toronto, ON, CanadaExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Toronto, ON, Canada
Visa sponsorship
Not specified

Job summary

Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate platforms for AI Co‑Workers.

Qualifications

  • 5+ years of professional experience in Site Reliability Engineering or DevOps Engineering
  • Kubernetes experience on EKS, AKS, GKE, or self-managed
  • Terraform experience for infrastructure provisioning
  • Helm experience for application deployment
  • Experience with at least two programming or scripting languages
  • Experience with reliability engineering and incident response
  • DevOps or DevSecOps experience including CI/CD and automation

Responsibilities

  • Design, build, and operate reliable production infrastructure supporting AI Co‑Workers
  • Own Kubernetes-based platforms for AI workloads
  • Build and maintain infrastructure as code using Terraform
  • Implement and maintain Helm-based deployment workflows
  • Define, measure, and improve system reliability using SLIs, SLOs, and SLAs
  • Participate in on-call rotation and incident response
  • Reduce operational toil through automation
  • Build and improve observability across monitoring and logging
  • Partner with engineers to ensure system resilience and security
  • Operate across software lifecycle phases

Skills

AKSArgo CDAWSAzureBashEKSGCPGKEGoHelmJavaKubernetesPowerShellPythonRubyTerraform

Certifications

CKACKAD

Degrees

Bachelors Degree in Computer ScienceInformation SystemsRelated field

Relocation

No