> Markdown version of [/jobs/ext/2311007-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2311007-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** Morgan Stanley - **Location:** Edison, NJ, United States - **Experience:** Expert - **Salary:** $175,000.0 - **Contract:** Permanent contract - **Skills:** Adobe Creative Cloud, Artificial Intelligence, Amazon Web Services, CA Workload Automation Ae, Microsoft Azure, Cloud Computing, Software Documentation, Computer Programming, Databases, Computer Engineering, IBM DB2, Software Debugging, DevOps, Disaster Recovery, Perl (Programming Language), Monitoring of Systems, Web Browsers, Python (Programming Language), Knowledge Management, Unix Shell, Networking Basics, Network Administration, Oracle (Applications), Performance Tuning, Systems Development Life Cycle, Reliability Engineering, Prometheus, Shell Script, Software Deployment, Software Engineering, Strategies of Testing, Web Analytics, Scripting, System Availability, Grafana, Sybase, Mttr, Information Technology, Performance Monitor, Kibana, Splunk - **Published:** August 30, 2026 - **Apply:** https://www.careerbuilder.com/job-details/lead-site-reliability-engineer-edison-nj--eccc8f75-9ec3-47f2-81b6-77197a5775bd ## About the Role * 10+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end-to-end troubleshooting, networking fundamentals and appropriate attention to detail * BS/MS or equivalent, preferably in quantitative discipline (Computer Science, Computer Engineering). * 5+ years' experience in leading a small to medium team of alike skillset. * 5+ years of experience in driving SRE principles and Chaos Engineering. * Experienced, technically hands-on professional that understands both code and infrastructure * Strong experience in scripting language (Shell scripting, Python, Perl, etc.) and cloud driven development * Strong database skills with DB2, Sybase or Oracle * Hands-on experience with Autosys or other batch scheduling software * Experience in AWS/GCP/Azure Cloud technologies * Working knowledge on any of the DevOps & observability tools (Grafana, Prometheus, Splunk, Kibana) * Solid analytical skills, problem determination, and resolution recovery processes * Ability to interface and cultivate excellent working relationships with technology teams, business analysts, and vendors * Experience in web analytics tools (preferably Adobe Experience Cloud tools) is Plus * Should be a fast learner of technologies in a quick paced environment. * Have strong organizational skills and the ability to manage multiple tasks and high-pressure situations for outage handling, management, or resolution * Is driven to learn about new technologies, techniques and what it takes to be an integral member of this team * Hands-on experience administering large-scale, high-availability systems and the tools to monitor performance and availability * Excellent communication and writing skills specific to technical discussions across the management layers * Experience with incident "on call" and ability to respond to emergencies on a 24/7 basis * Hands-on with AI and implementation of AI tools for operational efficiency * Strong ownership mentality with a focus on customer satisfaction * Be able to manage an outage incident, coordinating user communications, and other teams to help resolve an incident. * Experience working with Financial Services area will be a plus, Adobe Product Family, Analysis Skills, Artificial Intelligence (AI), Automation, Budgeting, Business Analysis, Business impact analysis (BIA), Call Monitoring, Capacity Management, Change Management, Cloud Computing, Communication Skills, Computer Engineering, Computer Science, Corrective Action, Customer Experience, Customer Relations, Customer Satisfaction, Customer Support/Service, Debugging Skills, Detail Oriented, DevOps, Disaster Recovery, Diversity, Documentation Standards, Emergency Response, Employee Benefits, Event Management, Finance, Financial Services, High Availability, Identify Issues, Investment Management, Knowledge Management, Leadership, Maintain Compliance, Metrics, Multitasking, Network Administration/Management, On Call, Operational Audit, Operational Support, Operations Management, Operations Processes, Organizational Skills, People Management, Performance Analysis, Performance Tuning/Optimization, Perl Programming Language, Power Outages, Problem Solving Skills, Process Improvement, Production Control, Production Support, Production Systems, Python Programming/Scripting Language, Quality Management, Reliability Engineering, Risk, Risk Analysis, Risk Management, Risk Management Framework (RMF), Scripting (Scripting Languages), Software Development, Software Development Lifecycle (SDLC), Splunk, Systems Administration/Management, Team Player, Technical Leadership, Telephone Skills, Test Plan/Schedule, Test Strategy, Time Management, Unix Shell Programming, Vendor/Supplier Evaluation, Wealth Management, Web Analytics, Web Browsers, Writing Skills ## Description The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk. This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues. What you'll do in the role: Drive Reliability Engineering Practices * Champion SRE principles, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational risk management frameworks. * Define and track reliability metrics that measure platform health, customer experience, and operational effectiveness. * Lead initiatives focused on reducing Mean Time to Detect (MTTD), Mean Time to Identify (MTTI), and Mean Time to Restore (MTTR). Own Production Reliability * Provide leadership for the proactive detection, triage, and resolution of production issues impacting business-critical applications and services. * Serve as the primary owner for escalated production incidents, driving resolution efforts across application, infrastructure, vendor, and external partner teams until service is restored and client impact is mitigated. * Establish a culture of operational excellence focused on stability, resiliency, and continuous service improvement. * Ensure clear, concise, and timely communication during outages, providing accurate business impact assessments and recovery updates to senior leadership. Production Governance & Change Management * Serve as a key gatekeeper for the production environment, ensuring adherence to change management policies, release controls, operational readiness standards, and risk management practices. * Assess the operational impact of technology changes and ensure appropriate testing, rollback strategies, monitoring, and support models are in place prior to production deployment. * Partner with development teams throughout the software lifecycle to ensure reliability, observability, and operational supportability are built into new applications and services. Automation & Operational Efficiency * Identify opportunities to eliminate manual effort and operational toil through automation, self-healing capabilities, and AI-driven operational workflows. * Lead the development and adoption of automation solutions that improve reliability, reduce risk, and increase operational efficiency across the organization. * Promote a culture of engineering-led operations and continuous process optimization. Operational Readiness & Knowledge Management * Establish and maintain a comprehensive knowledge management framework, ensuring runbooks, troubleshooting guides, standards, and operational procedures are accurate, current, and accessible. * Drive operational readiness programs that improve first-level diagnosis and reduce dependency on development teams for routine issue resolution. Create End-to-End Know your system diagrams. * Ensure support teams maintain high-quality documentation and standardized troubleshooting practices to accelerate incident resolution. Technical Leadership * Act as a senior technical leader and trusted advisor for reliability, resiliency, observability, and production support strategies. * Provide guidance on architecture reviews, platform scalability, capacity planning, disaster recovery, and resiliency testing initiatives. * Partner with engineering teams to identify systemic risks and implement long-term solutions that improve platform stability and customer experience. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Add Location-based Searching to Site with ElasticSearch](https://www.wearedevelopers.com/videos/77-add-location-based-searching-to-site-with-elasticsearch) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)