Jobs / Boeing
Lead Site Reliability Engineer
Boeing · Berkeley, MO, United States
Berkeley, MO, United StatesExp: 8+ yrs198,050-267,950 USD/yearlyOnsite
Remuneration
198,050-267,950 USD/yearly
Location
Berkeley, MO, United States
Visa sponsorship
No visa sponsorship
Job summary
The Lead Site Reliability Engineer is responsible for the reliability strategy, architecture, and operational maturity of mission-critical developer platforms used by engineering teams. Key responsibilities include defining technical strategy for GitLab, CI/CD runners, and PostgreSQL, leading complex technical investigations, and ensuring the security and reliability of developer tooling through automation and Infrastructure as Code practices.
Benefits
Health insuranceFlexible spending accountsHealth savings accountsRetirement savings plansLife and disability insurance programsTuition assistance programFertility, adoption, and surrogacy benefitsGift match for nonprofit organizations
Qualifications
- Bachelors Degree
- Ability to obtain a US Security Clearance for which the US Government requires US Citizenship
- Ability to obtain access to Special Access Programs (SAP)
- 8+ years of experience with CI/CD tools such as Jenkins or Bamboo
- 8+ years of enterprise architecture experience, including cloud architecture, security, data privacy, integration, and deployment
- 5+ years of experience in root cause analysis and corrective action
- 5+ years of technical leadership and team leadership
Responsibilities
- Define and lead the Site Reliability Engineering technical strategy for GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube, and related developer tooling infrastructure
- Establish platform reliability architecture, operational standards, SLIs, SLOs, SLAs, KPIs, error budgets, observability patterns, capacity models, backup strategies, and disaster recovery approaches
- Serve as the senior technical authority for complex reliability, performance, scalability, integration, database, automation, and security-related platform decisions
- Lead architecture and design reviews for developer tooling infrastructure, CI/CD runner topology, PostgreSQL operations, cloud-based and on-premises infrastructure, monitoring, alerting, access controls, and platform integrations
- Drive automation, Infrastructure as Code, Ansible, configuration management, and repeatable operational patterns that reduce toil and improve reliability
- Guide major upgrades, migrations, lifecycle planning, patch strategies, recovery planning, and technical roadmaps for supported platforms
- Lead the most complex incidents and technical investigations, including root cause analysis, corrective action planning, and systemic reliability improvements
- Mentor and technically guide SREs in operational excellence, troubleshooting, automation, secure administration, and architectural thinking
- Partner with program leadership, cybersecurity, infrastructure, software engineering, database, networking, suppliers, customers, and other stakeholders
- Identify platform risks, technical debt, capacity constraints, single points of failure, compliance concerns, and operational gaps, then drive remediation plans
- Lead efforts to operationally field higher-quality end-to-end system software more frequently
- Participate in after-hours support and escalation for urgent or mission-impacting issues as required
Skills
AnsibleArtifactoryAWSAzureBambooConfluenceDockerGitLabGitLab CIJenkinsJiraKubernetesPostgreSQLSonarQube
Certifications
Security+ certification
Degrees
Bachelor's Degree
Travel
10%
Security clearance
Active U.S. Secret Security Clearance (U.S. Citizenship Required)
Relocation
No