Senior Staff Reliability Engineer, Software Engineering

MSD
Morristown, NJ, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$142,400.0 - $224,100.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing Configuration Management System Configuration Information Engineering Data Visualization Network Planning and Design Open Source Technology Performance Tuning Release Management Reliability Engineering Software Deployment
+7 more
Software Engineering Software Project Management Systems Integration Web Platforms Delivery Pipeline Information Technology Build Tools

Job description

Join our company as we transform and innovate. We are at the forefront of delivering reliable, scalable, and resilient digital solutions that support critical scientific and business outcomes across our global organization. Our Digital Platforms & Services organization provides the technical foundation powering our company’s applications. We are seeking a highly experienced engineer who brings deep expertise in Site Reliability Engineering (SRE), Observability, and Resilience to help define and mature our reliability engineering practices. As a Senior Principal Reliability Engineer, you will lead the evolution of how reliability is engineered, measured, and improved across IT systems. You will play a critical role in enabling engineering teams to build systems that are reliable by design, while shaping enterprise practices that scale across the organization. This is a highly visible and impactful role with the potential to significantly improve the reliability, resilience, and operational effectiveness of the IT products that power our company’s mission. Responsibilities Build relationships across the broader IT organization to increase adoption and maturity of SRE, Observability, and Resilience practices Define and evolve the strategic vision for enterprise reliability engineering and ensure alignment across product, platform, and ITSM teams Establish and enforce standards for Service Level Objectives, observability frameworks, and resilience engineering practices Collaborate with engineering teams to ensure reliability is embedded into architecture, design, and delivery processes Drive adoption of Service Level Objectives using Nobl9 as the system of record for reliability governance Lead evaluation and introduction of new technologies that improve reliability outcomes while integrating with existing platforms Apply AI capabilities to enhance reliability practices, including incident triage, diagnostics, and automation, in a governed and controlled manner Collaborate within efforts to standardize observability across logs, metrics, traces, and events to improve system visibility and decision-making Consult and promote resilience patterns including fault isolation, failover strategies, and recovery mechanisms Guide improvements surrounding incident lifecycle effectiveness, including detection, response, root cause analysis, and continuous improvement Lead and mentor a community of reliability practitioners to grow organizational capability and maturity Represent reliability engineering practice in architecture reviews, governance forums, and key IT initiatives Drive continuous improvement of reliability practices through research, innovation, and feedback from engineering teams, San Francisco Residents Only: We will consider qualified applicants with arrest and conviction records for employment in compliance with the San Francisco Fair Chance Ordinance Los Angeles Residents Only: We will consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with the requirements of applicable state and local laws, including the City of Los Angeles’ Fair Chance Initiative for Hiring Ordinance Search Firm Representatives Please Read Carefully Merck & Co., Inc., Rahway, NJ, USA, also known as Merck Sharp & Dohme LLC, Rahway, NJ, USA, does not accept unsolicited assistance from search firms for employment opportunities. All CVs / resumes submitted by search firms to any employee at our company without a valid written search agreement in place for this position will be deemed the sole property of our company. No fee will be paid in the event a candidate is hired by our company as a result of an agency referral where no pre-existing agreement is in place. Where agency agreements are in place, introductions are position specific. Please, no phone calls or emails. Employee Status: Regular Relocation: No relocation VISA Sponsorship: Yes Travel Requirements: 10% Flexible Work Arrangements: Hybrid Shift: 1st - Day Valid Driving License: No Hazardous Material(s): N/A Job Posting End Date: 06/14/2026 *A job posting is effective until 11:59:59PM on the day BEFORE the listed job posting end date. Please ensure you apply to a job posting no later than the day BEFORE the job posting end date.

Requirements

Bachelors degree in IT, Engineering, Computer Science, or related field Minimum 7 years experience in site reliability engineering Expertise in capacity management, system integration, software development, release management, network design, configuration management (CM), software development life cycle (SDLC), system administration, change controls, and solution architecture Proficiency in designing, managing, developing, and maintaining technological products, particularly in the animal health domain Strong expertise in hardware, mechanics, artificial intelligence, and software development Experience in program management, including product definition, development, testing, maintenance, and tier 4 support Ability to conduct technological and product research and drive innovation Skilled in developing and managing CI/CD pipelines for product development cycles Knowledge of performance optimization and server software management Experience with application deployment to both cloud and on-premises production environments Understanding of product security, company development policies, and open source usage Strong leadership skills including strategic planning, entrepreneurship, innovation, and business savviness Proven track record in coaching and development, talent growth, and execution excellence Strong commitment to inclusion, with the ability to influence and motivate others Excellent emotional intelligence, decision-making skills, and a strong sense of ownership and accountability Networking and partnerships should be a key strength Required Skills: Data Engineering, Data Visualization, Design Applications, Software Configurations, Software Development, Software Development Life Cycle (SDLC), Solution Architecture, System Designs, System Integration, Testing Preferred Skills: Current Employees apply

Benefits & conditions

As an Equal Employment Opportunity Employer, we provide equal opportunities to all employees and applicants for employment and prohibit discrimination on the basis of race, color, age, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability status, or other applicable legally protected characteristics. As a federal contractor, we comply with all affirmative action requirements for protected veterans and individuals with disabilities. For more information about personal rights under the U.S. Equal Opportunity Employment laws, visit: We are proud to be a company that embraces the value of bringing together, talented, and committed people with diverse experiences, perspectives, skills and backgrounds. The fastest way to breakthrough innovation is when people with diverse ideas, broad experiences, backgrounds, and skills come together in an inclusive environment. We encourage our colleagues to respectfully challenge one another’s thinking and approach problems collectively. The salary range for this role is $142,400.00 - $224,100.00 This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee’s position within the salary range will be based on several factors including, but not limited to relevant education, qualifications, certifications, experience, skills, geographic location, government requirements, and business or organizational needs. The successful candidate will be eligible for annual bonus and long-term incentive, if applicable.

About the company

Our company is committed to inclusion, ensuring that candidates can engage in a hiring process that exhibits their true capabilities. Please if you need an accommodation during the application or hiring process.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri ¡ WWC 2023

48 sec

Exploring alternative build tools and experimental web components

Sasha Shynkevich ¡ LIVE

2:31 min

Web platform advancements and shared technological ecosystems

Nico Martin ¡ LIVE

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo ¡ WWC 2024

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley ¡ WWC 2021

2:09 min

Configuring IDEs and build tools for Java 17

Daniel Strmečki · LIVE

Videos

See all

Related articles

See all