Jobs / PulsePoint

Site Reliability Engineer, Big Data (Remote, International)

PulsePoint · United Kingdom · Remote
United KingdomExp: 5+ yrsHybrid
Remuneration
Not specified
Location
United Kingdom · Remote
British Summer Time (UTC+1)
Visa sponsorship
Not specified

Job summary

The Site Reliability Engineer at PulsePoint will build and maintain streaming and storage systems, focusing on Kafka and Ceph within a hybrid infrastructure.

Qualifications

  • 5+ years operating distributed systems at scale
  • Deep expertise in Kafka, Ceph, or similar infrastructure
  • Ability to design for scale and reliability
  • Experience mentoring engineers and making decisions

Responsibilities

  • Build streaming and storage systems
  • Own the lifecycle from architecture to incident response
  • Optimize Kafka architecture and topic design
  • Manage Ceph operations and capacity planning
  • Automate operations and improve incident response
  • Support SQL Server backup and recovery
  • Develop data team tooling with observability

Skills

AnsibleArgo CDCephGrafanaHadoopKafkaKubernetesPagerDutyPrometheusPuppetTerraform

Work schedule

9am-6pm ET U.S. hours

Relocation

No