Jobs / Sourceo
Site Reliability Engineer
Sourceo · Launceston, TAS, Australia
Launceston, TAS, AustraliaExp: 5+ yrsOnsite
Remuneration
Not specified
Location
Launceston, TAS, Australia
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer role focuses on supporting the daily operations and maintenance of AI-accelerated high-performance computing (HPC) infrastructure. Key responsibilities include deploying and maintaining GPU servers, troubleshooting hardware and software issues, and collaborating with engineering teams on tailored customer environments using technologies such as Kubernetes and Slurm.
Responsibilities
- Support deployment, configuration, and maintenance of high-end GPU servers, storage servers, and networking equipment in secure environments.
- Perform hardware diagnostics, systems functionality, and firmware updates.
- Collaborate with engineering teams for tailored customer environment deployment.
- Serve as first line of engineering support for onsite operational issues.
- Troubleshoot incidents, escalate critical issues, and provide feedback for improvements.
- Participate in on-call rotation for 24/7 availability.
- Provide technical support to GOC Support Specialist team.
- Document incident details, resolutions, and lessons learned.
- Maintain clear and up-to-date documentation for knowledge sharing.
- Communicate effectively with internal teams and stakeholders.
- Participate in team meetings and knowledge-sharing sessions.
Skills
BashGrafanaKubernetesLinuxPrometheusPython
Degrees
Bachelor’s degree in computer engineeringComputer scienceOr a related technical field
Relocation
No