Site Reliability Engineering (Sre)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
DescriptionKey Responsibilities* Ensuring Payment System Stability and High Availability: Lead technical initiatives to strengthen the reliability of our payment systems. This includes designingand implementing monitoring tools, logging frameworks, dashboards, diagnostic utilities, and disaster recovery plans. Conduct routine drills, develop contingencystrategies, and participate in on-call rotations to ensure rapid response and resolution of production issues across regions.* Incident Handling and Emergency Response: Conduct routine drills, develop contingency strategies, and participate in on-call rotations to ensure rapid responseand resolution of production issues.* Analyze and Optimize Production Issues: Investigate and analyze real-world production cases, such as performance bottlenecks or system inefficiencies, to deriveactionable insights and establish technical best practices. Contribute to the evolution of a highly available and resilient payment architecture.* Design and Implement Infrastructure Solutions: Architect and set up new Internet Data Centers (IDCs) to meet scalability and performance requirements. Developand execute comprehensive data protection plans that adhere to industry standards and compliance requirements, ensuring data integrity and security.Technical Requirements* Solid knowledge of Computer Science, and familiar with the principles of Operating System (Unix/Linux), Computer Storage, Computer Networking and otherrelated principles.* Proficient in at least one programming language, such as Java/Python/Shell with experience in developing operations and maintenance tools.* The strong ability to resolve system problems, good communication skills and a sense of ownership.* Experiences in operating Google Cloud Platform (GCP) / Oracle Cloud Infrastructure(OCI), OLAP platform (like DPDI, Flink, AntSpark), OcenBase (OB), Ant Trust-Native Service (ATS)is a plus.
Requirements
-
Solid knowledge of Computer Science, and familiar with the principles of Operating System (Unix/Linux), Computer Storage, Computer Networking and other related principles.
- Proficient in at least one programming language, such as Java/Python/Shell with experience in developing operations and maintenance tools.
- The strong ability to resolve system problems, good communication skills and a sense of ownership.
- Experiences in operating Google Cloud Platform (GCP) / Oracle Cloud Infrastructure(OCI), OLAP platform (like DPDI, Flink, AntSpark), OcenBase (OB), Ant Trust-Native Service (ATS) is a plus.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top Big Data Technologies That You Need to Know
Where To Find Software Engineering Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Top-Paying Tech Jobs (with Salaries)