Jobs / Evercommerce

EverCommerce - Senior Site Reliability Engineer

Evercommerce · United States · Remote
United StatesExp: 8+ yrs110,000-130,000 USD/yearlyRemote
Remuneration
110,000-130,000 USD/yearly
Location
United States · Remote
Eastern Daylight Time (UTC-4)
Visa sponsorship
No visa sponsorship

Job summary

The Senior Site Reliability Engineer is responsible for ensuring the reliability, performance, and operational stability of business-critical SaaS infrastructure across cloud and traditional hosting environments.

Qualifications

  • 8 years of systems, infrastructure, or cloud engineering experience
  • Experience administering Windows Server and Linux systems
  • Experience supporting infrastructure in AWS environments
  • Experience with enterprise virtualization platforms
  • Experience with infrastructure automation using scripting languages
  • Experience leading technical investigations for production issues
  • Understanding of networking fundamentals and operational resiliency
  • Familiarity with enterprise monitoring platforms
  • Familiarity with operational support for production database infrastructure
  • Strong communication and technical documentation skills
  • Ability to manage complex production infrastructure independently

Responsibilities

  • Maintain uptime, reliability, and performance for production SaaS environments
  • Support Windows and Linux infrastructure, virtualization platforms, storage systems, and networking components
  • Lead infrastructure modernization and operational automation initiatives
  • Troubleshoot complex infrastructure and production incidents
  • Design and implement improvements to monitoring and operational visibility
  • Collaborate with security teams for compliance initiatives
  • Lead vulnerability remediation and infrastructure maintenance activities
  • Develop and maintain technical documentation and operational procedures
  • Support reliable infrastructure for production database systems
  • Lead medium-sized infrastructure projects from planning to implementation
  • Participate in disaster recovery testing and operational improvement initiatives
  • Evaluate emerging technologies for infrastructure operations

Skills

AnsibleAWSBashChefLinuxPowerShellPuppetPythonTerraformWindowsWindows Server

Relocation

No