Jobs / T. Rowe Price

Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price · London, ENG, United Kingdom
London, ENG, United KingdomExp: 10+ yrsHybrid
Remuneration
Not specified
Location
London, ENG, United Kingdom
Visa sponsorship
Not specified

Job summary

The Principal Site Reliability Engineer, Infrastructure Observability at T. Rowe Price will lead a team focused on enhancing the observability, sustainability, and scalability of cloud and on-prem solutions. The role requires a strong operations and engineering background, with expertise in cloud environments and DevOps practices.

Qualifications

  • Bachelor's degree or equivalent and 10+ years of experience in cloud infrastructure
  • 5+ years building and supporting solutions in Amazon AWS
  • 5+ years in a DevOps and/or SRE function
  • Experience with chaos model implementation at scale
  • Strategic and program-level implementation experience
  • Experience implementing new technology and tools
  • System administration and scripting experience
  • Experience leveraging automation for incident prevention
  • Fluent in multiple programming languages
  • Proficient in database development
  • Proficient in defining and tracking Service Level Objectives
  • Experience managing Error Budgets
  • Proficient in explaining incident situations and recovery plans
  • Knowledge of dashboard standardization for observability
  • Experience with observability and cloud management tools
  • Ability to work independently in complex situations
  • Sound decision-making with limited resources
  • Balance strategic and pragmatic problem-solving
  • Adjust communication style for different audiences
  • Articulate operational principles and practices

Responsibilities

  • Possesses extensive knowledge in own area of expertise
  • Design technology solutions to prevent service disruptions
  • Prevent service disruptions through recommendations and automations
  • Foster a culture of deep learning through blameless post-mortems
  • Transform operations teams to adopt SRE methodologies
  • Analyze incidents impacting technology availability
  • Drive initiatives to reduce technology failures
  • Create cohesive views of the technology portfolio
  • Demonstrate awareness of complexities in tech and asset management
  • Lead initiatives of varying complexity
  • Contribute to target state architecture and design

Skills

AnsibleAWSAzure.NETGrafanaJavaMySQLNew RelicNode.jsPostgreSQLPrometheusPythonSplunkTerraformVagrantVaultGo

Relocation

No