Site Reliability Engineer

Arango, Inc.
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Artificial Intelligence Amazon Web Services Bash Shell Big Data Cloud Computing Cloud Engineering Computer Programming Continuous Integration Data Infrastructure DevOps Disaster Recovery
+29 more
Distributed Data Store Distributed Systems Fault Tolerance Monitoring of Systems Python (Programming Language) Linux Kernel Reliability Engineering Prometheus Circleci Data Logging Scripting Google Cloud Cloud Platform System System Availability Grafana Reliability of Systems Git Data Layers Containerization Kubernetes Infrastructure Automation Frameworks Production Code Terraform Software Version Control Docker Elk Stack Jenkins Golang Programming Languages

Job description

As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams

Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you’re passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!, * Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.

  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
  • Develop strategies for disaster recovery, high availability, and fault tolerance.
  • Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
  • Participate in on-call rotations to support critical production systems and respond to incidents.
  • Collaborate with cross-functional teams to improve overall system reliability and scalability.
  • Collaborate with the Customer Success team to resolve customer issues., At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.

Requirements

  • Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
  • Proficiency with Kubernetes in managing large-scale, distributed systems.
  • Experience with cloud providers such as AWS and Google Cloud (GCP).
  • Solid understanding of networking, security practices, and troubleshooting methods
  • Understanding of Linux internals (processes, environment variables etc.)
  • Familiarity with containerization technologies (e.g., Docker).
  • Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).
  • Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
  • Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.
  • Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence.

Strong ability to self-organize and to work independently as part of a remote team Knowledge of version control systems, particularly Git

  • Familiarity with programming languages such as Golang or Python

Nice-to-Have:

Experience managing distributed databases or large-scale data storage systems. Knowledge of security best practices in cloud environments.

  • Experience with scripting languages like Python or Bash.

Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. Experience working with GitOps

  • Strong programming skills in Golang, with experience in developing automation tools, scripts, or services.

Location: EU Timezone, preferably within the EU itself (Remote)

About the company

Arango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.

The Arango Contextual Data Platform connects fragmented enterprise data with LLMs, copilots, and AI agents through a simplified architecture delivered out of the box. By combining graph, vector, document, key-value, and search capabilities in a single platform, Arango eliminates the complex stacks many organizations build to operationalize enterprise AI.

Trusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air Force, NIH, Siemens, Transient.AI, Matpriskollen, and Articul8, Arango helps enterprises move from AI pilots to reliable production systems faster while lowering infrastructure complexity and total cost of ownership. Arango is a proud member of the NVIDIA Inception Program and the AWS ISV Accelerate Program. Learn more at arango.ai, LinkedIn, and G2.

Values at ArangoDB:

  • Innovation: We continually push boundaries in database technology to meet the evolving needs of developers and organizations.
  • Customer-Centric Focus: We prioritize the success of our users by delivering reliable, scalable, and performant database solutions.
  • Collaboration and Growth: We value a supportive work environment where employees can grow, learn, and share knowledge.

Joining ArangoDB means becoming part of a forward-thinking company that is transforming how data is handled across diverse applications, while working in an inclusive, growth-oriented team.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all