Senior Site Reliability Engineer (SRE / Backend)
nilo
Berlin, Germany
yesterday
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
SeniorJob location
Berlin, Germany
Tech stack
API
Amazon Web Services (AWS)
Data Migration
Software Debugging
DevOps
Amazon DynamoDB
Fault Tolerance
Identity and Access Management
Python
Key Management
PostgreSQL
Query Optimization
Reliability Engineering
Amazon Web Services (AWS)
Software Vulnerability Management
Web Applications
Working Model 2D
Datadog
Grafana
Indexer
Backend
Servicebus
Event Driven Architecture
Sentry
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Terraform
New Relic (SaaS)
Data Pipelines
Serverless Computing
Job description
We're looking for a Senior Site Reliability Engineer to take ownership of the infrastructure behind our platform. We run web applications and two mobile apps entirely on AWS, built around serverless and event-driven services and managed with Terraform., * Own our AWS infrastructure end to end - Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS
- Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow
- Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure
- Design our event-driven services for resilience: retries, dead-letter queues, idempotency, graceful degradation
- Own our Datadog and Sentry setup - define SLOs, build dashboards, and keep alerting actionable instead of noisy
- Lead incident response and run blameless postmortems that actually change how we build
- Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation
- Protect sensitive health data and support our GDPR and compliance requirements
- Monitor and optimize AWS spend without compromising reliability
- Contribute to backend development - APIs, event consumers, data pipelines, and Postgres and DynamoDB performance
- Participate in architectural discussions and mentor engineers on operational excellence, * Real ownership: a small team, short feedback loops, and no layers of approval between you and production
- Work that matters: the reliability you build directly affects people reaching for mental health support
- Free access to the nilo app (incl. family support)
- A dedicated learning budget for your personal and professional development
- Work abroad for up to 90 days per year (within the EU)
- Hybrid working model: 2 days/week from the office, 3 days from home
- Urban Sports Club membership at a discounted price
- Equity options: you benefit from any increase in nilo's valuation that you've helped to create
- Regular team and company events
- Bring your dog to work: we have 4 office dogs
Requirements
Must-have
- 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
- Deep AWS experience across serverless and containers - Lambda, ECS Fargate, and debugging both under pressure
- Strong Terraform skills, including module design and managing state across multiple environments
- Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
- Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
- Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
- Comfortable writing production backend code in [Python / Node.js / Go]
- Genuine on-call and incident response experience - you've led an incident and written the postmortem
- Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
- Pragmatic about complexity - you reach for the simplest thing that meets the reliability bar
- Effective communicator who can explain a technical tradeoff without jargon
Nice-to-have
- AWS DMS or other data migration and replication tooling
- Compliance experience (GDPR, SOC 2, ISO 27001)
- Experience in a B2B SaaS environment or healthcare-related product