> Markdown version of [/jobs/ext/2445857-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2445857-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Everforth Apex - **Location:** Irvine, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Application Configuration Access Protocols, Business Logic, User Authentication, Automation of Tests, Microsoft Azure, Cloud Computing, Computer Networks, Databases, Computer Literacy, Extract Transform Load (ETL), Data Structures, Database Development, Software Debugging, DevOps, Graph Database, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, Liquibase, MongoDB, Neo4j, Role-Based Access Control, Reliability Engineering, Ansible, Prometheus, Data Streaming, TypeScript, Transport Layer Security, Google Cloud, Istio, Large Language Models, Grafana, Multi-Cloud, Generative AI, Change Data Capture, Containerization, AI Platforms, Debezium, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Apache Flink, Apache Kafka, Bitbucket, Stream Processing, Data Pipelines, Dynatrace, Serverless Computing - **Published:** August 20, 2026 - **Apply:** https://www.dice.com/job-detail/57305765-a7e8-41ef-947d-2609b98e6e53 ## About the Role BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & education & experience in lieu of degree, 6+ years of relevant expertise is required. Strong, hands-on Kubernetes experience - deploying, scaling, and operating containerized workloads across multiple clusters and environments, including Helm-based packaging. Solid DevOps and GitOps practice, including CI/CD pipelines, infrastructure-as-code, and automated, GitOps-driven deployment (for example, with ArgoCD). Strong AWS skills, including object storage (S3), managed identity (IAM), secrets management, managed AI services, and managed Kubernetes (EKS). Working proficiency in React and TypeScript, sufficient to build and modify the platform's internal dashboard and graph-visualization UI. Strong Python engineering with a data-engineering focus, including building ETL services (incremental extraction, schema handling) and writing columnar data to object storage. Preferred: BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & experience in lieu of degree, 6+ years of relevant expertise is required. Experience with graph databases, particularly Neo4j and Cypher, is a strong plus. The platform's curated business layer is built on Neo4j, so this experience is highly valued. Experience with Kafka real-time event streaming components and workloads. Knowledge/Skills/Abilities: Ability to multi-task effectively without compromising the quality of the work. Excellent interpersonal, oral and written communication skills. Detail oriented, organized, process focused problem solver, proactive, ambitious, customer service focused. Ability to draw conclusions and make independent decisions with limited information. Ability to respond to common inquiries from customers, staff, regulatory agencies, vendors and other members of the business community. Self-motivated, reliable individual capable of working independently as well as part of the team. Motivated to drive improvement in a challenging environment. Strong background in data structures, algorithms and debugging. Demonstrated technical leadership, and successful participation in projects involving multiple engineers. Ability to learn quickly, understand complex systems and to work closely with others across multiple teams. Ability to handle uncertainty, time pressure and large technical challenges. Ability to deliver high-quality work on time. Strong attention to details, highly organized, computer literate ## Description As a Site Reliability Engineer, you will join a team responsible for the reliability, scalability, and operational excellence of a product platform. This role involves working with engineering and product teams to design, build, and operate resilient platform services. A key focus will be building and owning a new custom observability platform to provide a clear, end-to-end view of critical pipelines. This position requires some coding and a willingness to learn., Observability Platform Engineering o Build and own a new observability platform that gives both engineers and business stakeholders a single, end-to-end view of the health of the company's critical multi-system pipelines. o Bring together real-time telemetry for engineers and a curated, longer-term view for the business, so the organization can see not just why a request was slow but whether a pipeline as a whole is healthy and where it is failing. o Map how work flows across systems so that failures can be traced to their source, downstream impact is understood, and every issue has a clear, accountable owner. o Deliver the dashboards and views that let both technical and non-technical users understand pipeline health and drill into the underlying detail when needed. Cloud & Infrastructure Monitoring o Build capability for monitoring the health, performance, and availability of the company's multi-cloud infrastructure (AWS, Google Cloud Platform, and Azure) across enterprise-scale healthcare operations. o Track the state of Kubernetes clusters and cloud-native services, surfacing scaling, resiliency, and reliability issues before they cause outages. o Enhance observability for traffic and service-to-service communication across multi-cluster environments (for example, through the Istio service mesh) to catch latency and routing problems early. o Provide visibility into environment consistency, disaster-recovery readiness, and rollout health so that availability and business-continuity goals are met for mission-critical applications. Data Pipeline & Streaming Monitoring o Develop monitoring capabilities for real-time stream processing (for example, Apache Flink) for throughput, latency, backpressure, and job failures. o Track the health of event-streaming infrastructure such as Apache Kafka (Amazon MSK) and Kafka Connect, including consumer lag and data-ingestion issues. o Enhance observability for change-data-capture flows (for example, Debezium and MongoDB) and schema management for breakages and compatibility problems. o Assist with data orchestration team in enhancing observability for data-pipeline orchestration (for example, Dagster) so that failed or delayed jobs are detected and attributed quickly. AI & ML Monitoring o Track the performance, accuracy, and reliability of AI applications in production, including generative-AI and LLM/RAG workloads. o Provide visibility into the AI-serving stack, such as model endpoints and vector databases, for latency, scaling, and indexing health. o Develop AI guardrails and safety controls so that issues are detected and addressed quickly. Security, Compliance & Observability o Own the core monitoring and observability tooling (Prometheus, Grafana, OpenTelemetry) that provides distributed tracing and system-performance visibility. o Maintain security controls such as automated credential rotation (Conjur Cloud) and default-deny network policies. o Manage authentication and access for database clusters, including SASL/SSL (SCRAM-SHA-512) and RBAC permissions. o Help keep infrastructure compliant with HIPAA, SOC 2, and ISO 27001 through automated policy enforcement and encrypted communications. o Build "golden signal" dashboards (latency, traffic, errors, saturation) that populate automatically for every new microservice. o Reduce operational toil through automation and self-healing remediation, and help track and optimize platform costs across all layers. DevOps & Automation o Maintain CI/CD pipelines (BitBucket Pipelines, ArgoCD) and GitOps-based automated testing and deployment. o Support database deployment pipelines (Liquibase). o Automate infrastructure and application configuration (Ansible). o Enable automated health checks and failover for zero-downtime updates. o Partner with and mentor engineering teams on reliability, observability, and DevOps best practices. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)