> Markdown version of [/jobs/ext/2238057-site-reliability-engineer-application-operations](https://www.wearedevelopers.com/jobs/ext/2238057-site-reliability-engineer-application-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site-Reliability Engineer, Application Operations - **Company:** Caris Life Sciences - **Location:** Irving, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Applications, Application Portfolio Management, Information Systems, Continuous Integration, Software Design Patterns, DevOps, EHealth, HP Systems Insight Manager, Laboratory Information Management Systems, Reliability Engineering, Software Engineering, Scripting, IT General Controls (ITGC), Delivery Pipeline, Kubernetes, Information Technology, Performance Monitor, Static Application Security Testing, Dynamic Application Security Testing - **Published:** August 26, 2026 - **Apply:** https://dejobs.org/x/x/AF4957DCF87044B8903DF472C826A1AE/job/ ## About the Role * Bachelor's degree in Computer Science, Information Systems, Software Engineering, or a closely related technical field, or equivalent practical experience. * 5+ years of professional experience in SRE, production operations, DevOps, or platform engineering in a production-support capacity. * 2+ years of direct, hands-on experience participating in an on-call rotation as a primary production responder for Tier 1 or business-critical systems. * Experience executing production runbooks, maintenance scripts, or change procedures in a production environment with documented privileged-action controls. * Experience implementing or maintaining application monitoring, alerting thresholds, or dashboards in a production observability platform. * Experience participating in post-incident review or blameless retrospective processes, including timeline reconstruction and corrective-action tracking. * Hands-on use of AI coding assistants for automation, scripting, or operational tooling. Preferred Qualifications * Familiarity with SLO frameworks, error budgets, and associated alerting design patterns. * Experience reducing operational toil through scripting or automation in an SRE or production-operations setting. * Experience working in a SOX-controlled IT environment or a CLIA/CAP-regulated laboratory setting, including change-control ticketing, access-review processes, or audit-evidence collection. * Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms. * Domain experience in clinical diagnostics, laboratory information systems, or digital health software. * Familiarity with SAST/DAST tooling and secure CI/CD pipelines, including pipeline-embedded security scanning. * Familiarity with deployment pipelines, container orchestration, and release-automation tooling from an operate and release-support perspective. * Certifications in cloud platforms or information-security disciplines relevant to production operations. Physical Demands * Ability to sit, stand, and work at a computer for extended periods. ## Description Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it healthy is this role's mission. The Site-Reliability Engineer, Application Operations, works on the dedicated Application Site-Reliability Engineering (App-SRE) team. The work is site-reliability engineering across the clinical application portfolio, with real production telemetry and a direct hand in shaping the run-operate discipline. This is a production-operations craft role, distinct from application feature development, operating under governed privileged access and segregation-of-duties discipline. The engineer serves as a primary responder in the on-call rotation; executes runbooks and approved maintenance scripts with audit-grade discipline; implements application instrumentation, SLO monitors, and error-budget tracking; supports releases, deployment health verification, and rollback execution; and contributes to post-incident reviews. The team runs automation-first: recurring manual work is engineering backlog, and the engineer progressively automates away the toil they encounter rather than absorbing it. Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning runbook authoring, alert-quality analysis, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability, spanning application-performance monitoring and production telemetry, and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance. Reporting to the Director, Application Site-Reliability Engineering, this practitioner-level individual contributor operates production clinical applications subject to SOX financial controls and FDA regulatory requirements, working safely and accurately under established, documented procedures., * Participate in a scheduled on-call rotation as a primary responder for production incidents across the SOX- and FDA-regulated clinical application portfolio; execute triage, escalation, and initial remediation steps in accordance with documented runbooks and incident-response playbooks. * Execute approved runbooks and maintenance scripts in the production environment; document all privileged actions in compliance with SOX ITGCs and FDA audit-trail requirements. * Implement and maintain application-layer instrumentation, including dashboards, alert thresholds, and SLO monitors, in coordination with the centrally operated observability platform. * Support production releases and deployments: coordinate deployment health verification, monitor application behavior after deployment, and execute rollback procedures when required. * Participate in post-incident reviews: contribute timeline reconstructions, identify contributing factors, and track remediation action items through to closure. * Maintain audit-ready production-access logs, privileged-action records, and role-change documentation to support SOX ITGC and FDA regulatory audits. * Automate recurring operational work: convert repeated manual interventions, diagnostics, and maintenance procedures into scripted, reviewed, pipeline-executed automation under the team's change controls. * Track the manual-intervention rate for assigned services and drive it down over time by retiring runbook steps into automation. * Author and update runbook entries and known-issue documentation as operational knowledge is gained; contribute to continuous improvement of the operate discipline. * Collaborate with application engineering teams to gather context during incidents and to validate fixes deployed to the production environment. * Monitor application SLO attainment and error-budget consumption; escalate proactively when budgets are at risk. * Contribute to the ongoing development of on-call tooling, alert quality, and operational dashboards within the application-layer observability framework. * Work AI-first: use AI coding assistants and agentic workflows as daily practice in monitoring, triage, runbook work, and automation scripting, with review as the quality gate., * All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee. Other * This role includes participation in a scheduled on-call rotation with required after-hours response to production incidents and critical service events, including evenings and weekends. Periodic travel may be required to support business needs and team on-sites. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)