Jobs / App***

Software Engineer, Platform Reliability Engineering, AiDP

App*** · Sunnyvale, CA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Sunnyvale, CA, United StatesExp: 7+ yrs184,700-277,600 USD/yearlyHybrid
Remuneration
184,700-277,600 USD/yearly
Location
Sunnyvale, CA, United States
Visa sponsorship
Sponsors visa

Job summary

The Software Engineer in Platform Reliability Engineering is responsible for designing, operating, and optimizing large-scale distributed systems that support generative AI, machine learning, and big data platforms. This role involves maintaining scalable multi-tenant systems, responding to production incidents, and collaborating with engineering teams to enhance system reliability and performance using technologies such as Kubernetes, Spark, and Flink.

Benefits

Comprehensive medical and dental coverageRetirement benefitsDiscounted products and free servicesReimbursement for certain educational expenses

Qualifications

  • 7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale
  • Proficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments
  • Familiarity with open source codebases; ability to read, understand, and explain complex system implementations
  • Strong understanding of system architecture and proven ability to collaborate effectively across engineering teams
  • Hands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure)
  • Strong foundational knowledge of Linux, databases, and security principles
  • Excellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership
  • Demonstrated track record of designing and operating systems at scale
  • Proficiency in at least one systems programming language (Python, Go, Java, or similar)
  • Strong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles
  • Hands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies)

Responsibilities

  • Design, build, and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scale
  • Own the full lifecycle of infrastructure and platform projects from architectural design and implementation through deployment, monitoring, and optimization
  • Operate and optimize high-throughput, mission-critical services to ensure reliability, performance, and cost-efficiency
  • Participate in on-call rotations to respond to production incidents; diagnose root causes, implement rapid fixes, and drive post-incident improvements
  • Lead cross-functional collaboration with engineering teams to define requirements, validate designs, and deliver customer-impacting features and improvements
  • Proactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resilience
  • Establish observability practices and continuously refine operational excellence standards across the platform

Skills

GoJavaKubernetesLinuxPythonSparkTrino

Degrees

Bachelor's degree in Computer ScienceComputer Engineering

Relocation

No