Jobs / FDM***

Site Reliability Engineer- London

FDM*** · London, ENG, United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
London, ENG, United KingdomHybrid
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Sponsors visa

Job summary

FDM*** is seeking a Site Reliability Engineer to enhance monitoring and operational visibility for a client in the Finance sector. The role involves deploying and managing OpenSearch, along with developing observability solutions and collaborating with engineering teams.

Benefits

Career coachingMentoringAccess to upskillingAnnual leaveWork-place pension

Qualifications

  • Hands-on experience deploying and managing OpenSearch in production.
  • Understanding of OpenSearch architecture and performance tuning.
  • Experience with log aggregation and observability use cases using OpenSearch.
  • Knowledge of OpenSearch security and operational best practices.
  • Experience with Grafana for dashboard development and alerting.
  • Experience with enterprise monitoring platforms like Geneos.
  • Understanding of observability principles including metrics and logs.
  • Scripting skills in Python, Shell/Bash, or PowerShell.
  • Experience in SRE, Platform Engineering, DevOps, or Infrastructure Engineering.
  • Strong troubleshooting and problem-solving skills.
  • Experience with highly available and business-critical systems.

Responsibilities

  • Design, deploy, configure, and manage OpenSearch clusters and observability tooling.
  • Develop and maintain monitoring, logging, and alerting solutions for critical applications.
  • Build and enhance observability dashboards using Grafana.
  • Support and optimize Geneos monitoring implementations.
  • Create and maintain automation scripts for operational processes.
  • Collaborate with teams to improve system resilience and performance.
  • Implement SRE best practices, including monitoring standards and incident response.
  • Perform troubleshooting and root cause analysis of issues.
  • Support capacity planning and platform optimization.
  • Contribute to documentation and knowledge sharing.

Skills

AWSAzureBashDockerElasticsearchGCPGrafanaKubernetesOpenSearchPowerShellPython

Relocation

No