Senior Sre Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
Senior SRE EngineerPara presentar una candidatura, simplemente lea la siguiente descripción del puesto y asegúrese de adjuntar los documentos pertinentes.Join DempoOwn reliability at scale with us!We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience.This role sits at the Senior level on our Cloud/Platform/SRE career path - reliability engineering with a heavy focus on metrics and production systems.You’ll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go.You’ll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load.ResponsibilitiesReliability & Incident ManagementLead incidents as commander: set and revise severity, and know when to mitigate first and diagnose laterOwn the incident record and timeline standard, including the link between deployments and incidentsCommunicate with stakeholders while an incident is activeConduct blameless postmortems and drive toil identification and elimination as measured workObservabilityImplement the three pillars of observability (logs, metrics, traces) end to endDesign metrics and query strategy - recording rules, dashboard design that separates on-call needs from analyst needsDefine SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that’s easiest to measureDesign alerting systems - routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts)Platform & Production SystemsDesign workload health signals - liveness, readiness, and startup probes - and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining)Build runbook automation and self-healing systems to reduce operational toilContribute to CI/CD framework improvements and cost optimization initiativesTechnical LeadershipMake architectural decisions for reliability-critical systemsMentor cloud/platform engineersInfluence technical direction on infrastructure and platform decisionsRequirements5+ years of experience in SRE, Cloud, or Platform engineering rolesProfound knowledge of SRE principles and incident response practicesProfound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracingProfound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modesSolid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxiesHands-on experience with Infrastructure as Code (Terraform or equivalent)Profound knowledge of process and OS-level architecture trade-offs as they apply to containers (immutable infrastructure, image strategy)Scripting proficiency (Bash or Python) for tooling and automationStrong Git and collaborative workflow experienceNice to HaveExperience with chaos engineering or failure injection programsExposure to multi-region or multi-cloud design trade-offsFamiliarity with service mesh implementationsPrior mentoring or technical leadership experienceCertifications such as Site Reliability Engineering (SRE) Foundation or Observability FoundationContributions to open-source observability or Kubernetes toolingWhat We OfferPermanent contract.Flexible working hours (core hours: 09:***:30).Remote-first culture, with the option to work from our Granada office.30 working days of annual leave, plus December 24th and December 31st as additional company days off that do not count against your holiday allowance.Private health insurance.Your choice of MacBook or Lenovo.Continuous professional development:Unlimited access to learning platforms.Learning time during working hours.Budget for certifications and specialised training.Employee referral programme.Opportunity-based bonuses.Stable, long-term projects.Real opportunities for professional growth.The opportunity to play a key role in the growth of a modern Software Engineering company.Selection ProcessWith pleasure, we receive your CV and we give you a callNow it’s when we put faces to names, we’d love to get to chat with you!Let’s deepen a bit more with a technical interview, a chance to meet your People PartnerWe get back to you with offer and feedSalary Range· €45,000 - €55,000Dempo is where technology, teamwork, and your professional growth come together.xsgfvudHay opciones de teletrabajo/trabajo desde casa disponibles para este puesto.
Requirements
5+ years of experience in SRE, Cloud, or Platform engineering roles Profound knowledge of SRE principles and incident response practices Profound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracing Profound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modes Solid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxies Hands-on experience with Infrastructure as Code (Terraform or equivalent) Profound knowledge of process and OS-level architecture trade-offs as they apply to containers (immutable infrastructure, image strategy) Scripting proficiency (Bash or Python) for tooling and automation Strong Git and collaborative workflow experience Nice to Have Experience with chaos engineering or failure injection programs Exposure to multi-region or multi-cloud design trade-offs Familiarity with service mesh implementations Prior mentoring or technical leadership experience Certifications such as Site Reliability Engineering (SRE) Foundation or Observability Foundation Contributions to open-source observability or Kubernetes tooling
Benefits & conditions
Permanent contract. Flexible working hours (core hours: 09:***:30). Remote-first culture, with the option to work from our Granada office. 30 working days of annual leave, plus December 24th and December 31st as additional company days off that do not count against your holiday allowance. Private health insurance. Your choice of MacBook or Lenovo. Continuous professional development: Unlimited access to learning platforms. Learning time during working hours. Budget for certifications and specialised training. Employee referral programme. Opportunity-based bonuses. Stable, long-term projects. Real opportunities for professional growth. The opportunity to play a key role in the growth of a modern Software Engineering company. Selection Process With pleasure, we receive your CV and we give you a call Now it’s when we put faces to names, we’d love to get to chat with you! Let’s deepen a bit more with a technical interview, a chance to meet your People Partner We get back to you with offer and feed Salary Range · €45,000 - €55,000 Dempo is where technology, teamwork, and your professional growth come together. xsgfvud Hay opciones de teletrabajo/trabajo desde casa disponibles para este puesto.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Dev Digest 120 - Apple and peers
The Best X (Twitter) Accounts for Developers
The Best Software Developer Blogs to Read