Jobs / Focused
Site Reliability Engineer II - AI & Infrastructure (f/m/d)
Focused · Berlin, BE, Deutschland
Berlin, BE, DeutschlandOnsite
Remuneration
Not specified
Location
Berlin, BE, Deutschland
Visa sponsorship
Not specified
Job summary
The Site Reliability Engineer II role is centered on ensuring the reliability, security, and observability of internal applications while managing deployments and incident responses. Responsibilities include driving automation improvements, implementing code fixes, and supporting AI workflows in cloud environments like Azure, utilizing technologies such as PostgreSQL and CI/CD tools.
Qualifications
- Take ownership of technical work from investigation to deployment.
- Calmly respond to live incidents and diagnose interconnected system issues.
- Focus on root-cause resolution rather than temporary fixes.
- Communicate incidents and recovery plans clearly to stakeholders.
- Proactively identify reliability risks and operational improvements.
- Make independent decisions on tactical fixes and configuration changes.
- Document changes for effective system support by other engineers.
- Seek help early when risks are unclear and keep management informed.
Responsibilities
- Own the reliability and operation of internal applications and deployment platforms.
- Lead investigation and resolution of incidents and deployment failures.
- Drive automation improvements across deployment and monitoring processes.
- Implement code and configuration fixes to enhance service reliability.
- Develop proposals for cloud and database migrations and scaling approaches.
- Maintain operational documentation including runbooks and recovery procedures.
- Collaborate with stakeholders to execute infrastructure changes safely.
- Provide diagnostic support for AI workflows and tools.
Skills
AzureAzure DevOpsGitHubGitHub ActionsGitLabGitLab CIIAMLinuxOpenTofuPostgreSQLTerraform
Relocation
No