Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+42 more
Job description
As a Site Reliability Engineer, you will join a team responsible for the reliability, scalability, and operational excellence of a product platform. This role involves working with engineering and product teams to design, build, and operate resilient platform services. A key focus will be building and owning a new custom observability platform to provide a clear, end-to-end view of critical pipelines. This position requires some coding and a willingness to learn., Observability Platform Engineering
o Build and own a new observability platform that gives both engineers and business stakeholders a single, end-to-end view of the health of the company’s critical multi-system pipelines.
o Bring together real-time telemetry for engineers and a curated, longer-term view for the business, so the organization can see not just why a request was slow but whether a pipeline as a whole is healthy and where it is failing.
o Map how work flows across systems so that failures can be traced to their source, downstream impact is understood, and every issue has a clear, accountable owner.
o Deliver the dashboards and views that let both technical and non-technical users understand pipeline health and drill into the underlying detail when needed.
Cloud & Infrastructure Monitoring
o Build capability for monitoring the health, performance, and availability of the company’s multi-cloud infrastructure (AWS, Google Cloud Platform, and Azure) across enterprise-scale healthcare operations.
o Track the state of Kubernetes clusters and cloud-native services, surfacing scaling, resiliency, and reliability issues before they cause outages.
o Enhance observability for traffic and service-to-service communication across multi-cluster environments (for example, through the Istio service mesh) to catch latency and routing problems early.
o Provide visibility into environment consistency, disaster-recovery readiness, and rollout health so that availability and business-continuity goals are met for mission-critical applications.
Data Pipeline & Streaming Monitoring
o Develop monitoring capabilities for real-time stream processing (for example, Apache Flink) for throughput, latency, backpressure, and job failures.
o Track the health of event-streaming infrastructure such as Apache Kafka (Amazon MSK) and Kafka Connect, including consumer lag and data-ingestion issues.
o Enhance observability for change-data-capture flows (for example, Debezium and MongoDB) and schema management for breakages and compatibility problems.
o Assist with data orchestration team in enhancing observability for data-pipeline orchestration (for example, Dagster) so that failed or delayed jobs are detected and attributed quickly.
AI & ML Monitoring
o Track the performance, accuracy, and reliability of AI applications in production, including generative-AI and LLM/RAG workloads.
o Provide visibility into the AI-serving stack, such as model endpoints and vector databases, for latency, scaling, and indexing health.
o Develop AI guardrails and safety controls so that issues are detected and addressed quickly.
Security, Compliance & Observability
o Own the core monitoring and observability tooling (Prometheus, Grafana, OpenTelemetry) that provides distributed tracing and system-performance visibility.
o Maintain security controls such as automated credential rotation (Conjur Cloud) and default-deny network policies.
o Manage authentication and access for database clusters, including SASL/SSL (SCRAM-SHA-512) and RBAC permissions.
o Help keep infrastructure compliant with HIPAA, SOC 2, and ISO 27001 through automated policy enforcement and encrypted communications.
o Build “golden signal” dashboards (latency, traffic, errors, saturation) that populate automatically for every new microservice.
o Reduce operational toil through automation and self-healing remediation, and help track and optimize platform costs across all layers.
DevOps & Automation
o Maintain CI/CD pipelines (BitBucket Pipelines, ArgoCD) and GitOps-based automated testing and deployment.
o Support database deployment pipelines (Liquibase).
o Automate infrastructure and application configuration (Ansible).
o Enable automated health checks and failover for zero-downtime updates.
o Partner with and mentor engineering teams on reliability, observability, and DevOps best practices.
Requirements
BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & education & experience in lieu of degree, 6+ years of relevant expertise is required.
Strong, hands-on Kubernetes experience - deploying, scaling, and operating containerized workloads across multiple clusters and environments, including Helm-based packaging.
Solid DevOps and GitOps practice, including CI/CD pipelines, infrastructure-as-code, and automated, GitOps-driven deployment (for example, with ArgoCD).
Strong AWS skills, including object storage (S3), managed identity (IAM), secrets management, managed AI services, and managed Kubernetes (EKS).
Working proficiency in React and TypeScript, sufficient to build and modify the platform’s internal dashboard and graph-visualization UI.
Strong Python engineering with a data-engineering focus, including building ETL services (incremental extraction, schema handling) and writing columnar data to object storage.
Preferred:
BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & experience in lieu of degree, 6+ years of relevant expertise is required.
Experience with graph databases, particularly Neo4j and Cypher, is a strong plus. The platform’s curated business layer is built on Neo4j, so this experience is highly valued.
Experience with Kafka real-time event streaming components and workloads.
Knowledge/Skills/Abilities:
Ability to multi-task effectively without compromising the quality of the work.
Excellent interpersonal, oral and written communication skills.
Detail oriented, organized, process focused problem solver, proactive, ambitious, customer service focused.
Ability to draw conclusions and make independent decisions with limited information.
Ability to respond to common inquiries from customers, staff, regulatory agencies, vendors and other members of the business community.
Self-motivated, reliable individual capable of working independently as well as part of the team.
Motivated to drive improvement in a challenging environment.
Strong background in data structures, algorithms and debugging.
Demonstrated technical leadership, and successful participation in projects involving multiple engineers.
Ability to learn quickly, understand complex systems and to work closely with others across multiple teams.
Ability to handle uncertainty, time pressure and large technical challenges.
Ability to deliver high-quality work on time. Strong attention to details, highly organized, computer literate
About the company
Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.
Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
Is Software Engineering Over-Saturated?
Dev Digest 121 - AI goes offline