Site Reliability Engineer

Everforth Apex
Irvine, CA, United States
16 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Application Configuration Access Protocols Business Logic User Authentication Automation of Tests Microsoft Azure Cloud Computing Computer Networks Databases
+42 more
Computer Literacy Extract Transform Load (ETL) Data Structures Database Development Software Debugging DevOps Graph Database Monitoring of Systems Identity and Access Management Python (Programming Language) Key Management Liquibase MongoDB Neo4j Role-Based Access Control Reliability Engineering Ansible Prometheus Data Streaming TypeScript Transport Layer Security Google Cloud Istio Large Language Models Grafana Multi-Cloud Generative AI Change Data Capture Containerization AI Platforms Debezium Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Apache Flink Apache Kafka Bitbucket Stream Processing Data Pipelines Dynatrace Serverless Computing

Job description

As a Site Reliability Engineer, you will join a team responsible for the reliability, scalability, and operational excellence of a product platform. This role involves working with engineering and product teams to design, build, and operate resilient platform services. A key focus will be building and owning a new custom observability platform to provide a clear, end-to-end view of critical pipelines. This position requires some coding and a willingness to learn., Observability Platform Engineering

o Build and own a new observability platform that gives both engineers and business stakeholders a single, end-to-end view of the health of the company’s critical multi-system pipelines.

o Bring together real-time telemetry for engineers and a curated, longer-term view for the business, so the organization can see not just why a request was slow but whether a pipeline as a whole is healthy and where it is failing.

o Map how work flows across systems so that failures can be traced to their source, downstream impact is understood, and every issue has a clear, accountable owner.

o Deliver the dashboards and views that let both technical and non-technical users understand pipeline health and drill into the underlying detail when needed.

Cloud & Infrastructure Monitoring

o Build capability for monitoring the health, performance, and availability of the company’s multi-cloud infrastructure (AWS, Google Cloud Platform, and Azure) across enterprise-scale healthcare operations.

o Track the state of Kubernetes clusters and cloud-native services, surfacing scaling, resiliency, and reliability issues before they cause outages.

o Enhance observability for traffic and service-to-service communication across multi-cluster environments (for example, through the Istio service mesh) to catch latency and routing problems early.

o Provide visibility into environment consistency, disaster-recovery readiness, and rollout health so that availability and business-continuity goals are met for mission-critical applications.

Data Pipeline & Streaming Monitoring

o Develop monitoring capabilities for real-time stream processing (for example, Apache Flink) for throughput, latency, backpressure, and job failures.

o Track the health of event-streaming infrastructure such as Apache Kafka (Amazon MSK) and Kafka Connect, including consumer lag and data-ingestion issues.

o Enhance observability for change-data-capture flows (for example, Debezium and MongoDB) and schema management for breakages and compatibility problems.

o Assist with data orchestration team in enhancing observability for data-pipeline orchestration (for example, Dagster) so that failed or delayed jobs are detected and attributed quickly.

AI & ML Monitoring

o Track the performance, accuracy, and reliability of AI applications in production, including generative-AI and LLM/RAG workloads.

o Provide visibility into the AI-serving stack, such as model endpoints and vector databases, for latency, scaling, and indexing health.

o Develop AI guardrails and safety controls so that issues are detected and addressed quickly.

Security, Compliance & Observability

o Own the core monitoring and observability tooling (Prometheus, Grafana, OpenTelemetry) that provides distributed tracing and system-performance visibility.

o Maintain security controls such as automated credential rotation (Conjur Cloud) and default-deny network policies.

o Manage authentication and access for database clusters, including SASL/SSL (SCRAM-SHA-512) and RBAC permissions.

o Help keep infrastructure compliant with HIPAA, SOC 2, and ISO 27001 through automated policy enforcement and encrypted communications.

o Build “golden signal” dashboards (latency, traffic, errors, saturation) that populate automatically for every new microservice.

o Reduce operational toil through automation and self-healing remediation, and help track and optimize platform costs across all layers.

DevOps & Automation

o Maintain CI/CD pipelines (BitBucket Pipelines, ArgoCD) and GitOps-based automated testing and deployment.

o Support database deployment pipelines (Liquibase).

o Automate infrastructure and application configuration (Ansible).

o Enable automated health checks and failover for zero-downtime updates.

o Partner with and mentor engineering teams on reliability, observability, and DevOps best practices.

Requirements

BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & education & experience in lieu of degree, 6+ years of relevant expertise is required.

Strong, hands-on Kubernetes experience - deploying, scaling, and operating containerized workloads across multiple clusters and environments, including Helm-based packaging.

Solid DevOps and GitOps practice, including CI/CD pipelines, infrastructure-as-code, and automated, GitOps-driven deployment (for example, with ArgoCD).

Strong AWS skills, including object storage (S3), managed identity (IAM), secrets management, managed AI services, and managed Kubernetes (EKS).

Working proficiency in React and TypeScript, sufficient to build and modify the platform’s internal dashboard and graph-visualization UI.

Strong Python engineering with a data-engineering focus, including building ETL services (incremental extraction, schema handling) and writing columnar data to object storage.

Preferred:

BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & experience in lieu of degree, 6+ years of relevant expertise is required.

Experience with graph databases, particularly Neo4j and Cypher, is a strong plus. The platform’s curated business layer is built on Neo4j, so this experience is highly valued.

Experience with Kafka real-time event streaming components and workloads.

Knowledge/Skills/Abilities:

Ability to multi-task effectively without compromising the quality of the work.

Excellent interpersonal, oral and written communication skills.

Detail oriented, organized, process focused problem solver, proactive, ambitious, customer service focused.

Ability to draw conclusions and make independent decisions with limited information.

Ability to respond to common inquiries from customers, staff, regulatory agencies, vendors and other members of the business community.

Self-motivated, reliable individual capable of working independently as well as part of the team.

Motivated to drive improvement in a challenging environment.

Strong background in data structures, algorithms and debugging.

Demonstrated technical leadership, and successful participation in projects involving multiple engineers.

Ability to learn quickly, understand complex systems and to work closely with others across multiple teams.

Ability to handle uncertainty, time pressure and large technical challenges.

Ability to deliver high-quality work on time. Strong attention to details, highly organized, computer literate

About the company

Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:30 min

Introduction to Neo4j and remote developer relations work

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

Videos

See all

Related articles

See all