Jobs / Allwyn UK

Site Reliability Engineer

Allwyn UK · Watford, ENG, United Kingdom
Watford, ENG, United KingdomRemote
Remuneration
Not specified
Location
Watford, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

The Site Reliability Engineer at Allwyn UK will support the reliability and performance of digital services by operating production systems, building automation, and improving observability.

Qualifications

  • Experience in cloud environments (AWS preferred)
  • Working knowledge of containers
  • Infrastructure as Code (Terraform)
  • Basic programming/scripting capability
  • Ability to diagnose issues in distributed systems
  • Familiarity with monitoring and logging tools
  • Understanding of Linux systems and networking fundamentals
  • Understanding of monitoring and alerting concepts
  • Reliability principles
  • Incident response processes
  • Willingness to be part of an on-call rotation
  • Exposure to SLOs / SLIs / error budgets
  • CI/CD pipelines
  • Performance and load testing
  • Familiarity with observability platforms
  • Experience in customer-facing, high-availability systems

Responsibilities

  • Maintain reliable production services across digital platforms
  • Improve monitoring, alerting, and observability coverage
  • Reduce operational toil through automation
  • Support incident response and continuous improvement
  • Contribute to performance and scaling of services
  • Participate in on-call rotation
  • Respond to incidents
  • Support out-of-hours diagnosis
  • Monitor system health across platforms
  • Troubleshoot issues across application, infrastructure, and network layers
  • Support incident triage and resolution
  • Participate in post-incident reviews
  • Maintain and improve operational documentation
  • Implement and maintain monitoring tools
  • Improve logging quality and alerting accuracy
  • Develop scripts to reduce manual tasks
  • Contribute to infrastructure management
  • Support deployment processes and CI/CD improvements
  • Assist in performance tuning and capacity planning
  • Work closely with engineers to improve service reliability

Skills

AWSBashCloudWatchECSEKSGrafanaKubernetesLinuxPythonSplunkTerraform

Relocation

No