Site Reliability Engineer
Manychat
Barcelona, Spain
13 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
PHP (Programming Language)
Agile Methodology
Ubuntu (Operating System)
Cloud Computing Security
Continuous Integration
Software Debugging
Linux
Github
Identity and Access Management
Python (Programming Language)
PostgreSQL
Nginx
+7 more
Ansible
Prometheus
Reverse Proxy
Grafana
AI Platforms
Kubernetes
Terraform
Job description
- Maintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch)
- Operate and evolve our EKS clusters powering Python-based AI services
- Migrate existing services to Kubernetes using Terraform and Helm
- Codify infrastructure with Terraform and manage host-level automation via Ansible
- Build and improve CI/CD pipelines with GitHub Actions
- Own observability efforts: Prometheus, Grafana, alerting, and on-call readiness
- Support OS-level patching, certs, WAF rules, and general infra hygiene
- Partner with engineers to guide best practices and drive platform reliability
- Create clean, maintainable infrastructure documentation and playbooks
- Occasionally support rare off-hours incidents (don’t worry, really rare)
Requirements
- 5+ years of experience managing Linux in production (Ubuntu, Amazon Linux)
- Strong experience with Kubernetes (ideally EKS), Helm, and Terraform
- Comfort with running and debugging Python workloads in containers
- Solid understanding of networking, IAM, and cloud security best practices
- Hands-on Nginx experience (Ingress and reverse proxy setups)
- Excellent communication skills; you can explain complex infra to devs clearly, * Strong Ansible skills beyond the basics
- PostgreSQL or Amazon RDS tuning and operations experience
- Deep understanding of observability tools (Prometheus, Grafana, Loki, etc.)
- Familiarity with PHP production environments
- Experience with TDD, CI/CD best practices, and agile development
- Any previous SRE-like exposure such as building resilience, automation, or incident tooling
Benefits & conditions
- Hybrid onboarding to start work remotely and relocation support for you and your family.
- Comprehensive health insurance for both you and your family.
- Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
- Flexible benefits package to tailor perks that matters most for you.
- Hybrid work and generous leave options to prioritize your work-life balance.
- In-office perks, including free meals and snacks.
- Company-funded sport activities, annual offsites and team-building events.
About the company
We help creators get more out of every conversation with Instagram-focused automations and support for other channels like Messenger, WhatsApp, and TikTok. The result? Better engagement, more sales, and real, sustainable growth. With a diverse team of 350+ people spread across three continents, we’re building the leading Chat Marketing platform that is used - and loved - by more than 1.5 million customers worldwide.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.jobleads.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
about 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
EM
Eli McGarvie
The Best X (Twitter) Accounts for Developers
almost 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
about 2 years ago