> Markdown version of [/jobs/ext/1891985-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1891985-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** Walt Disney Studios - **Location:** Bay Lake, FL, United States - **Experience:** Expert - **Salary:** $148,300.0 - $198,800.0 - **Contract:** Franchise - **Skills:** Java (Programming Language), Agile Methodology, Artificial Intelligence, Amazon Web Services, Systems Engineering, Microsoft Azure, C++ (Programming Language), Cyber Security, Information Systems, Continuous Integration, Decision Support Systems, DevOps, Distributed Systems, Perl (Programming Language), Python (Programming Language), Reliability Engineering, Ansible, Ruby, Rust (Programming Language), Google Cloud, Cloud Platform System, Delivery Pipeline, Gitlab, Cloudformation, Procedural Programming, Build Management, Infrastructure Automation Frameworks, Information Technology, Terraform, Golang - **Published:** August 1, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27898659/Lead-Site-Reliability-Engineer-Florida-Bay-Lake-1015 ## About the Role * Minimum 7 years of related work experience * Proficient in agile environments * Applied expertise in observability principles and tools, including defining and implementing SLIs, SLOs, and SLAs * Hands-on experience with CI/CD tools like Gitlab, AWS CodeBuild , Azure DevOps * Proficient in configuration management tools: Terraform, CloudFormation, Ansible, Chef * Experience in procedural programming languages ( Python, Perl, Ruby, Java, Go, Rust, C/C++ ) * Skilled in Cloud environments ( AWS, Azure, Google Cloud ) * Proven ability to design and build reliable, scalable enterprise systems * Capable of leading reliability efforts and identifying root causes in large-scale distributed systems * Proficient in UNIX/Linux administration, troubleshooting, and security * Demonstrated experience leading technical projects and ensuring smooth delivery * Collaborative work with Security Operations teams for secure solutions * Strong troubleshooting skills across systems, network, and code * Proven experience mentoring, guiding, or training other engineers, with strong written and verbal communication and the ability to influence without direct authority * Proactive demeanor toward continuous learning and mastering emerging tools and methodologies, * Bachelor's degree in Computer Science , Information Systems, Software, Electrical or Electronics Engineering, or comparable field of study, and/or equivalent work experience required ## Description "We Power the Magic!" That's our motto at Disney Experiences (DX). Our team creates world-class immersive digital experiences for the Company's premier vacation brands including Disney's Parks & Resorts worldwide, Disney Cruise Line, Aulani, a Disney Resort & Spa, and Disney Vacation Club. We are responsible for the end-to-end digital and physical Guest experience for all technology & digital-led initiatives across the Attractions & Entertainment, Food & Beverage, Resorts & Transportation and Merchandise lines of business as well as other initiatives including MyDisneyExperience and Hey, Disney! The US Parks Site Reliability Organization is accountable for the reliability and resilience of a portfolio of critical applications and services. We partner with product, engineering, and SRE teams across the organization to define and evolve reliability standards rooted in SRE and DevOps principles. Rather than simply operating systems, we apply engineering approaches-observability, automation, and data - driven decision making-to proactively improve service health. By designing reliability into platforms and reducing operational toil, we enable teams to focus on innovation and delivering world - class guest experiences. This role sits in the US Parks & Resorts organization within Disney Experiences Technology and works closely with other site reliability engineers, application delivery teams, and systems engineers from across the company. The Lead Site Reliability Engineer will report to the Manager, System Engineering . What You Will Do * Serve as the SRE subject matter expert and technical lead for assigned products and platforms, owning the reliability strategy and embedding SRE and DevOps best practices * Drive adoption and tracking of service level management-defining and operationalizing SLIs, SLOs, and SLAs-for the systems and applications in your assigned portfolio * Lead the design, build, and support of products and platforms; consult on and build development pipelines, automate infrastructure and operations, and create telemetry for monitoring * Engineer high reliability and reinforce best practices to secure company data across systems, network, performance, capacity, and operational excellence * Mentor and guide other site reliability and systems engineers, providing coaching, feedback, and technical direction to elevate team performance and hold self and others accountable to commitments * Lead Major Incident response for owned services-minimizing Mean Time to Resolve and delivering comprehensive retrospectives that result in measurable improvements to prevent future failures * Partner with engineering, product, and program management to align priorities, manage dependencies, contribute to estimation and planning, and negotiate solutions to complex reliability challenges * Champion a DevOps culture and a shift-left, reliability-by-design mindset among peers and developers * Stay current with emerging technologies and apply AI/automation to reduce toil and improve service health ## Related Videos - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)