> Markdown version of [/jobs/ext/2275128-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2275128-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Eversports - **Location:** Wien, Austria - **Experience:** Expert - **Salary:** €70,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Bash Shell, Continuous Integration, Relational Databases, DevOps, Github, Python (Programming Language), PostgreSQL, MySQL, Node.Js, Reliability Engineering, Prometheus, Runbook, Amazon Simple Notification Service (SNS), TypeScript, AWS Cdk, GitHub Copilot, Grafana, Infrastructure as Code (IaC), Servicebus, Amazon Relational Database Service, Deployment Automation, AWS Fargate, Functional Programming, Amazon Simple Queue Service (SQS), Programming Languages - **Published:** August 28, 2026 - **Apply:** https://www.adzuna.at/details/5840360309 ## About the Role * 8+ years in SRE, DevOps, or Platform Engineering in mid/large-scale production systems, including experience inside an environment where mature SRE/reliability practices were already established, ideally having built up or significantly contributed to said environment. * Proven track record of operating and reasoning about large, complex, interconnected production systems end-to-end (not only deep expertise in a single component). * Strong experience operating and running production infrastructure on AWS. * Deep, practical experience with reliability frameworks: SLIs, SLOs, error budgets, and alerting strategy. * Experience with relational databases (MySQL or Postgres) from a reliability/operations perspective. * Solid observability background (OpenTelemetry, Grafana/Prometheus, or similar). * Experience defining and running incident management and on-call processes. * Experience with Infrastructure as Code (IaC), preferably AWS CDK. * Comfortable with at least 2 programming languages e.g. TypeScript/Node.js, Bash, Python * Proficiency with CI/CD and deployment automation (GitHub Actions or similar). * Curiosity and openness towards AI tools in an infra/ops workflow (e.g. Claude, GitHub Copilot, or similar). * Fluent written and verbal communication skills in English, * Hands on experience with our stack: ECS Fargate, Lambda, SQS, SNS, EventBridge, RDS MySQL. * Experience coaching or enabling other teams to adopt reliability practices. * Verbal communication skills in German. ## Description * Own Production Reliability: Take ownership of monitoring, alerting, and observability across our AWS environment so issues are detected and resolved before they reach users. * Build a System-Wide View: Develop and maintain a holistic understanding of how our production systems fit together end-to-end, using it to find the highest-leverage reliability improvements and to prioritize where to act first. * Lead Incident Response: Investigate incidents, run root-cause analysis, and drive long-term fixes that prevent recurrence. Establish and evolve incident-handling and on-call practices, and participate in the on-call rotation. * Enable Team Ownership: Empower the value-aligned teams to own the health of their own applications - their key metrics, alarms, and runbooks - rather than centralizing reliability in the Platform Team. * Tame Reactive Load: Provide a structured intake for infrastructure-related requests from teams (triage, planning, proposal, delivery), reducing today's unpredictable operational load. * Automate Away Toil: Drive automation that improves resilience and reduces manual, repetitive operational work. * Strengthen Recoverability & Compliance: Improve backup, recovery, and disaster-recovery practices, and support security and compliance requirements (NIS2, GDPR). * Partner with Platform Engineering: Work closely with the Platform Engineers, feeding operational reality into their forward-looking platform initiatives (including CI/CD). ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Coding for Good: Achieving social change with an app](https://www.wearedevelopers.com/videos/1645-coding-for-good-achieving-social-change-with-an-app) ## Related Articles - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs)