Site Reliability Engineer

deepset
Barcelona, Spain
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Software as a Service Continuous Integration Software Debugging Github Machine Learning Reliability Engineering Site Reliability Engineering Practices Prometheus Service-Oriented Architecture Private Cloud Environment
+4 more
Datadog Large Language Models Kubernetes Terraform

Job description

You won’t just “keep things running” - you’ll help define how our platform is built, deployed, and scaled across cloud and customer environments.

  • Build and operate real-world infrastructure. Design, configure, and evolve infrastructure that runs both in our cloud and inside customer environments (SaaS, private cloud, on-prem).
  • Make self-hosted production-ready. Help us deliver a production-grade, self-hosted platform that can be deployed on any Kubernetes setup in weeks - not months.
  • Drive automation & platform maturity. Improve CI/CD pipelines, GitHub workflows, and GitOps setups so teams can ship faster with confidence.
  • Reduce complexity and cost. Continuously simplify systems and optimize infrastructure spend without compromising performance or reliability.
  • Shape how we build. Champion best practices in reliability, scalability, and security across the organization, not as rules, but as working systems.

Requirements

Do you have experience in Terraform?, Do you have a Master’s degree?, * 2-5 years of experience working with large-scale production infrastructure

  • Experience with distributed or service-oriented architectures
  • Hands-on expertise with:
  • AWS
  • Kubernetes
  • CI/CD and GitOps (e.g. ArgoCD)
  • Working knowledge of Infrastructure as Code (Terraform preferred)
  • Solid troubleshooting skills - you can debug across systems, not just within one layer
  • A pragmatic mindset: you balance speed, simplicity, and reliability
  • Ownership and accountability - you take responsibility for systems end-to-end
  • Ability to work independently while staying aligned with the team’s goals

Nice to have

  • Familiarity with observability stacks (e.g. Datadog, Prometheus)
  • Experience optimizing cloud costs at scale
  • Interest or experience in Machine Learning / LLM systems
  • Experience improving developer experience and platform tooling using AI agents
  • Contributions to SRE practices like postmortems, SLIs/SLOs, and reliability engineering culture

Benefits & conditions

  • Remote-first setup with flexible hours & tech of your choice
  • 30 days vacation + extra days for family sick leave
  • Competitive salary & stock options for every team member
  • Monthly sports & mental health support allowance with Oliva
  • Annual learning & development budget
  • Monthly team socials & in-person meetups
  • Dog-friendly Berlin HQ

About the company

Founded in 2018, deepset builds open and enterprise-grade tools that help teams build AI with purpose. From Haystack, our open-source framework, to the Haystack Enterprise Platform, we give developers and organizations the building blocks to solve complex, high impact challenges with AI with full control, transparency, and sovereignty. Backed by GV and Balderton, we’re growing the world’s production AI community and customer base solving challenges too critical to get wrong.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters ¡ WWC 2023

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 ¡ LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis ¡ LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley ¡ WWC 2021

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle ¡ Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 ¡ Coffee With Developers

Videos

See all

Related articles

See all