Jobs / ARANGO

Site Reliability Engineer

ARANGO · En remoto, Spain
En remoto, SpainRemote
Remuneration
Not specified
Location
En remoto, Spain
Visa sponsorship
Not specified

Job summary

Arango is seeking a Site Reliability Engineer (SRE) to enhance the reliability, scalability, and performance of their cloud-native infrastructure supporting distributed database systems. The role involves designing and maintaining cloud infrastructure, automating tasks, and collaborating with development teams to ensure high availability and performance of cloud-based systems.

Qualifications

  • Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
  • Proficiency with Kubernetes in managing large-scale, distributed systems.
  • Experience with cloud providers such as AWS and Google Cloud (GCP).
  • Solid understanding of networking, security practices, and troubleshooting methods.
  • Understanding of Linux internals (processes, environment variables etc.).
  • Familiarity with containerization technologies (e.g., Docker).
  • Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).
  • Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
  • Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.
  • Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence.
  • Strong ability to self-organize and to work independently as part of a remote team.
  • Knowledge of version control systems, particularly Git.
  • Familiarity with programming languages such as Golang or Python.

Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
  • Develop strategies for disaster recovery, high availability, and fault tolerance.
  • Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
  • Participate in on-call rotations to support critical production systems and respond to incidents.
  • Collaborate with cross-functional teams to improve overall system reliability and scalability.
  • Collaborate with the Customer Success team to resolve customer issues.

Skills

AWSBashCircleCIDockerGCPGitGoGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform

Relocation

No