Jobs / Future Secure AI
Site Reliability Engineer
Future Secure AI · Toronto, ON, Canada
Toronto, ON, CanadaExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Toronto, ON, Canada
Visa sponsorship
Not specified
Job summary
Future Secure AI is seeking a Site Reliability Engineer to design, build, and operate platforms for AI Co‑Workers.
Qualifications
- 5+ years of professional experience in Site Reliability Engineering or DevOps Engineering
- Kubernetes experience on EKS, AKS, GKE, or self-managed
- Terraform experience for infrastructure provisioning
- Helm experience for application deployment
- Experience with at least two programming or scripting languages
- Experience with reliability engineering and incident response
- DevOps or DevSecOps experience including CI/CD and automation
Responsibilities
- Design, build, and operate reliable production infrastructure supporting AI Co‑Workers
- Own Kubernetes-based platforms for AI workloads
- Build and maintain infrastructure as code using Terraform
- Implement and maintain Helm-based deployment workflows
- Define, measure, and improve system reliability using SLIs, SLOs, and SLAs
- Participate in on-call rotation and incident response
- Reduce operational toil through automation
- Build and improve observability across monitoring and logging
- Partner with engineers to ensure system resilience and security
- Operate across software lifecycle phases
Skills
AKSArgo CDAWSAzureBashEKSGCPGKEGoHelmJavaKubernetesPowerShellPythonRubyTerraform
Certifications
CKACKAD
Degrees
Bachelors Degree in Computer ScienceInformation SystemsRelated field
Relocation
No