> Markdown version of [/jobs/ext/1951149-principal-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1951149-principal-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - **Company:** Commercetools - **Location:** Granada, Spain - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Cloud Computing, Operational Data Store, Reliability Engineering - **Published:** August 6, 2026 - **Apply:** https://www.buscojobs.com.es/principal-site-reliability-engineer-en-granada-ID-365648385 ## About the Role Experience: 7+ years driving incident management/operational excellence; 5+ years leading org-wide resiliency and reliability initiatives. Peak Load Execution: Proven track record of managing and scaling systems for high-stakes, multi-team operational events (e.g., Black Friday, major launches). Data Literacy: Strong ability to analyze metrics to diagnose technical issues and measure process improvements. Leadership & Influence: Ability to evaluate both technical bugs and the organizational/human dynamics behind them. Project Management: Demonstrated success managing large-scale initiatives that span multiple engineering teams in an Agile environment. Communication: Fluent English with exceptional written and verbal communication skills; experience running technical training or onboarding is a plus. Soft Skills: High self-awareness, a strong customer focus, and a passion for mentoring others and learning new technologies. ## Description About commercetoolsReal innovation starts with a strong foundation, and at commercetools, that comes from the perfect balance of our product and our people.Behind every leap forward is a collective of builders, explorers, doers, makers, and problem-solvers.The kind of people who not only pioneered a more flexible approach to commerce architecture but also shaped the culture of experimentation that approach unlocked. Together they are the engine of commerce innovation today.At commercetools, we power the next era of autonomous commerce for our customers. Whether it's AI-driven solutions that help enterprises make smarter business decisions, bridging digital and physical shopping experiences, or enabling entirely new ways for industries to connect with their customers, we help the world's most ambitious companies experiment, scale, and grow without limits.Here the best idea wins, not the loudest voice. You will have the tools, trust, and space to not only build the future of commerce, but to build your own.Your ImpactAs the Principal Engineer championing Resiliency , you'll be the driving force behind how commercetools prepares for, responds to, and learns from operational incidents at scale. Our customers rely on us for mission-critical commerce infrastructure, including during their highest-stakes moments of the year, like Black Friday. You'll own the discipline of resiliency end-to-end: mature incident management, strong operational visibility, data-driven process improvement, and organization-wide readiness for peak-traffic events.Standardize Incident Management: Build and champion intuitive, end-to-end processes for incident detection, response, communication, and postmortems across the company.Enhance System Visibility: Develop clear, real-time metrics, dashboards, and signals to track system health and incident trends.Drive Data-Backed Improvements: Use operational data to find process gaps, partnering with product engineering teams to fix them.Own Peak-Event Readiness: Scale and lead the organization-wide readiness program for massive traffic spikes like Black Friday .Lead Cross-Team Initiatives: Identify resilience gaps, collaborate with infrastructure/product teams on solutions, and turn ideas into concrete technical outcomes.Cross-Functional Collaboration: Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers (Cloud, Security, API, Architecture, Performance).Foster Knowledge Sharing: Drive organizational communication, documentation, and training around resiliency and operational excellence.This role is to based 3 days a week in our hubs with a hybrid modality - Berlin, Valencia or London.What Sets You ApartYou're a creative problem-solver who is wired to find solutions. You confidently dive into complex challenges and have a talent for making them simple for others. Your curiosity drives you to constantly grow and contribute to an environment of trust and teamwork. Great ideas come from many paths, and your unique perspective matters more than checking every box. What matters most is the mindset you bring to the work.You BringExperience: 7+ years driving incident management/operational excellence; 5+ years leading org-wide resiliency and reliability initiatives.Peak Load Execution: Proven track record of managing and scaling systems for high-stakes, multi-team operational events (e.g., Black Friday, major launches).Data Literacy: Strong ability to analyze metrics to diagnose technical issues and measure process improvements.Leadership & Influence: Ability to evaluate both technical bugs and the organizational/human dynamics behind them.Project Management: Demonstrated success managing large-scale initiatives that span multiple engineering teams in an Agile environment.Communication: Fluent English with exceptional written and verbal communication skills; experience running technical training or onboarding is a plus.Soft Skills: High self-awareness, a strong customer focus, and a passion for mentoring others and learning new technologies.Our BenefitsComprehensive health benefits for you and your dependents, including access to OpenUp for personalized mental health supportLearning and development opportunities including an annual learning budget, access to self-paced learning platforms and language training, personalized coaching, mentorship, and leadership programsFamily Leave Plus gives you additional fully paid weeks of parental leave on top of government-provided leave, so you can spend more time with your new additionOur equity participation program allows you to share in our successFor more information on our benefits, visit this page.Come as you are. Build with us. ...#J-*****-Ljbffr ## Related Videos - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [ShapeShift: Reinventing Agile for a B2B SaaS Scale-Up](https://www.wearedevelopers.com/videos/1655-shapeshift-reinventing-agile-for-a-b2b-saas-scale-up) - [Dos and don'ts with react hooks. An opinionated approach](https://www.wearedevelopers.com/videos/957-dos-and-don-ts-with-react-hooks-an-opinionated-approach) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)