Jobs / bet365
Site Reliability Engineer
bet365 · Manchester, ENG, United Kingdom
Manchester, ENG, United KingdomHybrid
Remuneration
Not specified
Location
Manchester, ENG, United Kingdom
Visa sponsorship
Not specified
Job summary
Enhance stability and performance of systems supporting global products through software engineering, automation, and incident response.
Qualifications
- Software engineering background with Python, Golang, JavaScript or similar language.
- Knowledge of modern development practices and delivery lifecycles.
- Understanding of SRE principles and incident management.
- Experience with observability tools like OpenTelemetry and Grafana.
- Proficiency in shell scripting for automation.
- Experience with Infrastructure as Code, including Terraform and Ansible.
- Knowledge of Cloudflare or comparable edge platforms.
- Ability to troubleshoot distributed systems.
- Experience in a large-scale, 24/7 enterprise environment.
- Practical experience using LLM platforms to improve productivity.
Responsibilities
- Develop and maintain resilient tools and automation for effective system management.
- Use orchestration and scripting to improve operational consistency.
- Contribute to code and instrumentation that enhance service reliability.
- Build dashboards using telemetry from observability platforms.
- Manage Cloudflare edge services using Infrastructure as Code.
- Diagnose incidents and coordinate effective remediation.
- Participate in live incident response and root-cause analysis.
- Administer monitoring and alerting toolsets.
- Drive initiatives for reliability and continuous improvement.
- Mentor colleagues and share knowledge.
Skills
AnsibleBashCloudflareGoGrafanaJavaScriptNew RelicOpenTelemetryPagerDutyPythonSplunkTerraform
Relocation
No