Jobs / For***

Site Reliability Engineer

For*** · United States · Remote
Visa sponsorship details are locked. Unlock company name and apply link with .
United StatesExp: 3+ yrs85,400-192,900 USD/yearlyRemote
Remuneration
85,400-192,900 USD/yearly
Location
United States · Remote
Eastern Daylight Time (UTC-4)
Visa sponsorship
Sponsors visa

Job summary

For*** is seeking a Site Reliability Engineer to enhance their global monitoring and observability platform. The role involves blending software and systems engineering to ensure the uptime and scalability of critical cloud services, while collaborating with development teams to improve system reliability.

Qualifications

  • 3+ years of experience as an SRE, Software Engineer, DevOps Engineer or similar role
  • Solid programming skills in Golang and scripting languages
  • Proficient with monitoring and observability tools
  • Experience with cloud services, especially Kubernetes and GCP
  • Experience with relational and document databases
  • Ability to debug, optimize code, and automate tasks
  • Strong problem-solving skills under pressure
  • Excellent verbal and written communication skills

Responsibilities

  • Write, configure, and deploy code in Go and Javascript to improve service reliability
  • Optimize performance and cost within Google Cloud Platform (GCP)
  • Provide feedback and review for code or production changes
  • Drive repair and optimization of complex systems
  • Lead debugging and troubleshooting of service architecture
  • Participate in on-call rotation
  • Write documentation including design and runbooks
  • Implement and manage SRE monitoring application backends
  • Collaborate with development teams to enhance system reliability
  • Develop automated solutions for operational aspects
  • Troubleshoot and resolve issues in development and production environments
  • Participate in postmortem analysis and create preventative measures
  • Implement and maintain security best practices
  • Participate in capacity planning and forecasting
  • Identify and address performance bottlenecks
  • Develop and test disaster recovery plans
  • Contribute to internal knowledge bases

Skills

DynatraceGCPGoJavaScriptKubernetesMakeOpenTelemetryPostgreSQLTerraform

Relocation

No