Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
We are seeking a Site Reliability Engineer (SRE) to join the GDI&A SRE team, focused on observability, monitoring, and technical consulting across Google Cloud Platform-based data platforms. This role is responsible for ensuring the availability, reliability, performance, and scalability of cloud infrastructure and services through automation, proactive monitoring, and continuous improvement initiatives., * Collaborate with infrastructure and engineering teams to design and implement automation solutions that eliminate manual operational activities.
- Monitor and manage production environments, proactively identifying, troubleshooting, and resolving system issues.
- Develop and maintain tooling for system access monitoring, session recording, logging, and reliability management across distributed environments.
- Partner with engineering teams to improve on-call processes, incident response, root cause analysis, and post-incident reviews.
- Perform capacity planning and resource optimization to support evolving platform demands and traffic patterns.
- Configure and maintain monitoring, alerting, and health-check systems to ensure platform stability and availability.
- Drive continuous improvements in system performance, reliability, security, and operational efficiency through data-driven analysis.
- Create and maintain technical documentation, architecture diagrams, operational runbooks, and knowledge-sharing materials.
- Support BigQuery workloads, Google Cloud Platform infrastructure services, and CI/CD pipeline reliability initiatives.
Requirements
The ideal candidate will have hands-on experience with Google Cloud Platform (Google Cloud Platform), BigQuery, observability tools, CI/CD practices, and incident management., * Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 4+ years of overall IT experience.
- 3+ years of software development, platform engineering, cloud engineering, or SRE experience.
- Hands-on experience with Google Cloud Platform (Google Cloud Platform).
- Experience supporting cloud-based production environments and distributed systems.
- Proficiency with monitoring and observability platforms such as Dynatrace, Datadog, New Relic, or similar tools.
- Experience with BigQuery and cloud data platform operations.
- Familiarity with IT Service Management (ITSM) tools, including ServiceNow for incident, problem, and change management.
- Experience with at least one programming language or automation framework.
Preferred Qualifications
- Experience with Google Cloud Platform Cloud Run and related cloud-native services.
- Python development and automation experience.
- Strong troubleshooting and problem-solving skills.
- Experience defining, measuring, and reporting Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Familiarity with AI-powered tools and platforms, including Copilot, LLMs, agents, and automation solutions.
About the company
Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.
Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
7 Cloud Computing Trends Coming in 2025 for Developers
Data Engineer Salary UK
Software Engineer Salary London