> Markdown version of [/jobs/ext/2719367-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2719367-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Everforth Apex - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** BigQuery, Cloud Computing, Cloud Engineering, Continuous Integration, Distributed Systems, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Data Driven Tests, Software Engineering, Software Testing Automation Framework, Datadog, Data Logging, Google Cloud, Application Enhancement Tool, Cloud Platform System, Large Language Models, Grafana, Information Technology, Performance Monitor, Data Management, New Relic (SaaS), Dynatrace, Serverless Computing, Servicenow, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://www.dice.com/job-detail/ec1822b8-d57d-4a5a-b5f3-f723af02300b ## About the Role The ideal candidate will have hands-on experience with Google Cloud Platform (Google Cloud Platform), BigQuery, observability tools, CI/CD practices, and incident management., * Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field. * 4+ years of overall IT experience. * 3+ years of software development, platform engineering, cloud engineering, or SRE experience. * Hands-on experience with Google Cloud Platform (Google Cloud Platform). * Experience supporting cloud-based production environments and distributed systems. * Proficiency with monitoring and observability platforms such as Dynatrace, Datadog, New Relic, or similar tools. * Experience with BigQuery and cloud data platform operations. * Familiarity with IT Service Management (ITSM) tools, including ServiceNow for incident, problem, and change management. * Experience with at least one programming language or automation framework. Preferred Qualifications * Experience with Google Cloud Platform Cloud Run and related cloud-native services. * Python development and automation experience. * Strong troubleshooting and problem-solving skills. * Experience defining, measuring, and reporting Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs). * Familiarity with AI-powered tools and platforms, including Copilot, LLMs, agents, and automation solutions. ## Description We are seeking a Site Reliability Engineer (SRE) to join the GDI&A SRE team, focused on observability, monitoring, and technical consulting across Google Cloud Platform-based data platforms. This role is responsible for ensuring the availability, reliability, performance, and scalability of cloud infrastructure and services through automation, proactive monitoring, and continuous improvement initiatives., * Collaborate with infrastructure and engineering teams to design and implement automation solutions that eliminate manual operational activities. * Monitor and manage production environments, proactively identifying, troubleshooting, and resolving system issues. * Develop and maintain tooling for system access monitoring, session recording, logging, and reliability management across distributed environments. * Partner with engineering teams to improve on-call processes, incident response, root cause analysis, and post-incident reviews. * Perform capacity planning and resource optimization to support evolving platform demands and traffic patterns. * Configure and maintain monitoring, alerting, and health-check systems to ensure platform stability and availability. * Drive continuous improvements in system performance, reliability, security, and operational efficiency through data-driven analysis. * Create and maintain technical documentation, architecture diagrams, operational runbooks, and knowledge-sharing materials. * Support BigQuery workloads, Google Cloud Platform infrastructure services, and CI/CD pipeline reliability initiatives. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)