> Markdown version of [/jobs/ext/1938145-site-reliability-engineering-sre](https://www.wearedevelopers.com/jobs/ext/1938145-site-reliability-engineering-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineering (Sre) - **Company:** Worldfirst - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Applicant Tracking Systems, Computer Networks, Computer Literacy, Data Centers, Data Integrity, Disaster Recovery, Payment Systems, Python (Programming Language), Online Analytical Processing, Reliability Engineering, Data Logging, Google Cloud, System Availability, Apache Flink, Oracle Cloud Infrastructure, Programming Languages - **Published:** August 5, 2026 - **Apply:** https://www.buscojobs.com.es/site-reliability-engineering-sre-en-madrid-ID-365547451 ## About the Role Solid knowledge of Computer Science, and familiar with the principles of Operating System (Unix/Linux), Computer Storage, Computer Networking and other related principles. Proficient in at least one programming language, such as Java/Python/Shell with experience in developing operations and maintenance tools. The strong ability to resolve system problems, good communication skills and a sense of ownership. Experiences in operating Google Cloud Platform (GCP) / Oracle Cloud Infrastructure(OCI), OLAP platform (like DPDI, Flink, AntSpark), OcenBase (OB), Ant Trust-Native Service (ATS) is a plus. #J-*****-Ljbffr ## Description Description Position at Ant GroupKey ResponsibilitiesEnsuring Payment System Stability and High Availability: Lead technical initiatives to strengthen reliability of our payment systems, designing and implementing monitoring tools, logging frameworks, dashboards, diagnostic utilities, and disaster recovery plans.Conduct routine drills, develop contingency strategies, and participate in on-call rotations to ensure rapid response and resolution of production issues across regions.Incident Handling and Emergency Response: Conduct routine drills, develop contingency strategies, and participate in on-call rotations to ensure rapid response and resolution of production issues.Analyze and Optimize Production Issues: Investigate and analyze real-world production cases, such as performance bottlenecks or system inefficiencies, to derive actionable insights and establish technical best practices.Contribute to the evolution of a highly available and resilient payment architecture.Design and Implement Infrastructure Solutions: Architect and set up new Internet Data Centers (IDCs) to meet scalability and performance requirements.Develop and execute comprehensive data protection plans that adhere to industry standards and compliance requirements, ensuring data integrity and security.Technical RequirementsSolid knowledge of Computer Science, and familiar with the principles of Operating System (Unix/Linux), Computer Storage, Computer Networking and other related principles.Proficient in at least one programming language, such as Java/Python/Shell with experience in developing operations and maintenance tools.The strong ability to resolve system problems, good communication skills and a sense of ownership.Experiences in operating Google Cloud Platform (GCP) / Oracle Cloud Infrastructure(OCI), OLAP platform (like DPDI, Flink, AntSpark), OcenBase (OB), Ant Trust-Native Service (ATS) is a plus.#J-*****-Ljbffr ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Build Delightful Mobile Experiences with Kotlin, Realm, and Atlas Device Sync](https://www.wearedevelopers.com/videos/694-build-delightful-mobile-experiences-with-kotlin-realm-and-atlas-device-sync) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)