Jobs / CrowdStrike

Engineer II, Site Reliability (Remote, GBR)

CrowdStrike · United Kingdom · Remote
United KingdomExp: 5+ yrsRemote
Remuneration
Not specified
Location
United Kingdom · Remote
British Summer Time (UTC+1)
Visa sponsorship
Not specified

Job summary

CrowdStrike is seeking an Engineer II for their TechOps SRE team, focusing on the Commercial Cloud. The role involves developing automation and tooling for large-scale distributed systems, ensuring operational excellence, and collaborating with a global team of engineers.

Qualifications

  • Five years of experience in a large scale production environment
  • Two years of experience in software engineering
  • Two years of experience in C++, Java, Python, or Go
  • Experience with storage technologies such as SAN, NAS, NFS, Object Storage, FreeNAS, iSCSI
  • Experience with infrastructure technologies including Linux, Windows, VMware, Docker, Kubernetes
  • Experience writing technical documentation
  • Configuration management experience with tools like Puppet, Chef, Ansible
  • Understanding of application design and operational trade-offs
  • Strong analytical skills and sense of urgency
  • Ability to work in a diverse, team-focused environment
  • Ability to communicate and present reliability conventions
  • Experience utilizing AI technologies to enhance decision-making and improve efficiency

Responsibilities

  • Expertise in Linux engineering and administration for bare metal servers and virtual machines
  • Responsible for operational aspects of platform including availability, latency, throughput, monitoring, issue response, and capacity planning
  • Collaborate with a global team of engineers
  • Participate in on-call rotation
  • Troubleshoot server hardware issues
  • Ensure platform operates flawlessly 24/7
  • Champion new technologies and raise technical IQ of the team
  • Gain broad exposure to architecture and process flow
  • Drive improvements
  • Engage in small and larger development projects
  • Experience with modern monitoring and telemetry stacks
  • Gather and analyze metrics for performance tuning and fault finding
  • Lead incident analysis and drive resolution

Skills

AnsibleChefC++DockerGoGrafanaJavaKubernetesLinuxPrometheusPuppetPythonVMwareWindows

Degrees

Bachelor's degree in Computer Science or equivalent experience

Relocation

No