Jobs / Haven

Site Reliability Engineer

Haven · Hemel Hempstead, ENG, United Kingdom
Hemel Hempstead, ENG, United KingdomRemote
Remuneration
Not specified
Location
Hemel Hempstead, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

Haven is seeking a hands-on Site Reliability Engineer to enhance the performance, scalability, and reliability of their digital platforms.

Qualifications

  • Deep knowledge of AWS and cloud engineering best practices
  • Experience with CI/CD and configuration management tools
  • Hands-on experience with Infrastructure as Code, ideally Terraform
  • Experience with Docker and Kubernetes
  • Working knowledge of NodeJS and TypeScript
  • Understanding of observability, incident management, and security practices
  • Grounding in cloud and network security principles
  • Experience with legacy and modern database systems
  • Active use of AI-assisted coding tools
  • Understanding of AI agent architectures
  • Consultative style with ability to coach on cloud architecture

Responsibilities

  • Support infrastructure engineering from design to implementation
  • Design and improve CI/CD processes and pipelines using Git Actions
  • Troubleshoot build and deployment issues
  • Maintain internally developed engineering tools
  • Own monitoring, tracing, and observability
  • Drive database reliability across relational and non-relational systems
  • Ensure security controls are in place
  • Identify performance and scalability improvements
  • Define incident management approach
  • Ensure disaster recovery documentation and training are ready
  • Shape release process and deployment strategies
  • Define guardrails and cloud policies with Platform Team
  • Share cloud best practices

Skills

AWSAzureDockerGCPGitKubernetesNode.jsOpenSearchPostgreSQLTerraformTypeScriptGitHub ActionsOracle Cloud

Certifications

AWS or other cloud accreditations

Relocation

No