Senior Site Reliability Engineer

Runware
Málaga, Spain
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work

Tech stack

PHP (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Databases Software Debugging DevOps Distributed Systems Python (Programming Language) MySQL Operational Databases Software Architecture Queueing Systems
+11 more
RabbitMQ Redis Reliability Engineering Load Balancing Kubernetes Low Latency Deployment Automation Bare Metal Vertica Dynatrace Serverless Computing

Job description

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.What you’ll doOwn and improve the reliability, availability and performance of critical production services across the Runware platformDefine and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standardsInvestigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotationLead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvementsReduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilienceWork closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform growsRequirementsHave strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar roleHave a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructureHave experience designing and operating observability systems using metrics, logs and distributed tracingUnderstand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toilHave experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHPTake strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediationBonusExperience operating high-throughput or low-latency APIs and distributed systemsExperience with bare-metal infrastructure, GPU environments or AI and ML workloadsExperience with RabbitMQ or other distributed messaging and queueing systemsExperience operating MySQL, Redis, ClickHouse or similar production data systemsExperience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environmentsExperience building automated scaling, capacity management or self-healing systemsBenefitsWe’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.Generous paid time off- vacation, sick days, public holidaysMeaningful stock options- share in the upside you createRemote-first setup- work from home anywhere we can employ youFlexible hours- own your schedule outside core collaboration blocksFamily leave- paid maternity, paternity, and caregiver timeCompany retreats- twice-yearly gatherings in inspiring locations

Requirements

Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure Have experience designing and operating observability systems using metrics, logs and distributed tracing Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation Bonus Experience operating high-throughput or low-latency APIs and distributed systems Experience with bare-metal infrastructure, GPU environments or AI and ML workloads Experience with RabbitMQ or other distributed messaging and queueing systems Experience operating MySQL, Redis, ClickHouse or similar production data systems Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments Experience building automated scaling, capacity management or self-healing systems

Benefits & conditions

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life. Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap. Generous paid time off

  • vacation, sick days, public holidays Meaningful stock options

  • share in the upside you create Remote-first setup

  • work from home anywhere we can employ you Flexible hours

  • own your schedule outside core collaboration blocks Family leave

  • paid maternity, paternity, and caregiver time Company retreats

  • twice-yearly gatherings in inspiring locations

About the company

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · WWC Europe 2026

Videos

See all

Related articles

See all