> Markdown version of [/jobs/ext/138113-lead-medtech-technology-service-reliability](https://www.wearedevelopers.com/jobs/ext/138113-lead-medtech-technology-service-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead, MedTech Technology Service Reliability... - **Company:** Johnson & Johnson - **Location:** Raritan, NJ, United States - **Experience:** Expert - **Salary:** $94,000.0 - $151,800.0 - **Contract:** Permanent contract - **Skills:** Microsoft Access, Agile Methodology, Artificial Intelligence, Amazon Web Services, Automation of Tests, Microsoft Azure, Continuous Integration, DevOps, Fault Tolerance, Human-Computer Interaction, Key Management, Networking Basics, Systems Development Life Cycle, Reliability Engineering, Runbook, Software Engineering, Software Vulnerability Management, IT General Controls (ITGC), Mttr, Git, Cloudformation, Information Technology, Terraform, Meditech, Jenkins - **Published:** May 21, 2026 - **Apply:** https://www.juju.com/job/00000000g1s0qg ## About the Role + Bachelor's degree in Computer Science, Engineering, or related discipline, or equivalent experience. + 5+ years of experience in SRE, DevOps, platform engineering, or software engineering with substantial production operations responsibilities. + Hands-on experience with observability and incident management practices, including monitoring/alerting design, on-call operations, and root-cause analysis. + Experience with infrastructure-as-code and CI/CD (e.g., Terraform/CloudFormation, Git, Azure DevOps/Jenkins or similar) and automated testing/release practices. + Experience operating services in cloud-hosted or hybrid enterprise environments (AWS and/or on-prem), including networking fundamentals, secure configuration, and environment management. + Strong communication skills with the ability to explain technical issues, incident impact, reliability risks, and tradeoffs to both technical and non-technical stakeholders. + Working knowledge of Agile delivery practices and ability to collaborate across cross-functional teams (Product, Engineering, QA/Validation, Security, Infrastructure) to deliver reliable, well-managed releases. + Experience working in MedTech, Life Sciences, or other regulated environments, including familiarity with validated systems, documentation expectations, and controlled change processes. + Demonstrates **AI Fluency** -the ability to **use and evaluate AI technologies responsibly** (with a primary focus on **generative AI in the workplace** )-to improve productivity and decision quality while maintaining human accountability, managing risk, and complying with applicable governance, privacy, security, and policy requirements, Product Lifecycle Management (PLM), Reliability Engineering Preferred Skills: Agile Product Development, Analytical Reasoning, Coaching, Collaborating, Competitive Landscape Analysis, Critical Thinking, Customer Alignment, Demand Forecasting, Human-Computer Interaction (HCI), Organizing, Product Development, Product Improvements, Product Strategies, Requirements Analysis, Research and Development, Software Development Life Cycle (SDLC), Software Development Management, Stakeholder Management, Technical Credibility, Technical Writing, Technologically Savvy ## Description We are searching for the best talent for a** **Lead, MedTech Technology Service Reliability Engineer, R&D** **to be located in Raritan, NJ, The Service Reliability Engineer (SRE) designs, builds, and operates reliability practices and technical capabilities that ensure critical engineering and enterprise services are available, performant, secure, and resilient. This is a hands-on, non-manager role focused on improving service reliability through observability, incident response, automation, and engineering excellence. This role partners closely with Product Owners, development teams, infrastructure/platform engineering, Quality/Validation, Security, and Enterprise Architecture to define reliability targets, implement operational controls, and maintain documentation appropriate for regulated environments. The SRE helps standardize operational patterns across environments (dev/test/prod), including monitoring baselines, access controls, runbooks, change management, and deployment readiness. Key outcomes include establishing and measuring Service Level Indicators/Objectives (SLIs/SLOs), improving alert quality and troubleshooting speed, reducing incident frequency and Mean Time to Recovery (MTTR), and enabling safe, repeatable releases through automation and operational readiness. The SRE identifies reliability risks and technical gaps, recommends scalable and resilient designs, implements reusable operational tooling, and participates in Agile ceremonies and on-call support aligned to the team's ways of working. Major Duties & Responsibilities + Define, implement, and continuously improve reliability standards for production services, including SLIs/SLOs, error budgets, and operational readiness criteria. + Build and maintain observability capabilities (metrics, logs, traces, dashboards) and establish actionable alerts that reflect customer impact. + Participate in on-call rotations, lead incident triage and restoration, and drive root-cause analysis with corrective and preventive actions. + Engineer reliability improvements through automation (self-healing, auto-remediation, runbook automation) and eliminate toil through scripting and tooling. + Partner with engineering teams to design and validate resilient architectures (timeouts/retries, circuit breaking, queuing, graceful degradation) and to improve deployment safety. + Perform capacity planning and performance analysis; proactively identify bottlenecks and reliability risks, and validate scaling strategies. + Establish and maintain operational runbooks, playbooks, and escalation paths; conduct game days and resilience testing (e.g., failover/chaos exercises) as appropriate. + Improve change management by defining deployment/rollback standards, validating monitoring coverage, and supporting release readiness reviews across dev/test/prod. + Create and maintain operational documentation (service catalogs, SLIs/SLOs, runbooks, monitoring standards) and ensure knowledge transfer across teams. + Support validation and audit readiness by following SDLC/IT controls, producing required evidence (e.g., monitoring/test results), and supporting controlled releases in regulated environments. + Develop reliability reporting (availability, latency, error rates, MTTR, incident trends) and present insights and recommendations to stakeholders. + Apply security-by-design principles (identity/access, secrets management, vulnerability management, data protection) and ensure operational practices meet company standards. + Collaborate with internal teams and vendors as needed to implement reliability improvements, manage platform upgrades, and continuously improve maintainability and supportability. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [User 1st! Technology 2nd! Stop building AI nobody uses - start delivering real business outcomes](https://www.wearedevelopers.com/videos/100337-user-1st-technology-2nd-stop-building-ai-nobody-uses-start-delivering-real-business-outcomes) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer)