> Markdown version of [/jobs/ext/1899428-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1899428-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** LEANDATA, INC. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Systems Engineering, C Sharp (Programming Language), Cloud Computing, Cloud Computing Security, Computer Programming, Continuous Delivery, Continuous Integration, DevOps, Disaster Recovery, Data Intelligence, Python (Programming Language), Unix Shell, Microsoft Dynamics, Network Architecture, Node.Js, Redis, Reliability Engineering, PM2, Salesforce.Com, Shell Script, Amazon Simple Notification Service (SNS), Apex Code, Autoscaling, Amazon ElastiCache, System Availability, Software Security, AWS Lambda, AngularJS, Information Technology, Npm(Software), Functional Programming, Api Gateway, Amazon Simple Queue Service (SQS), Terraform, New Relic (SaaS) - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-site-reliability-engineer-santa-clara-ca--49d19190-19f9-4de1-a739-85dbe1ec6d6f ## About the Role * Experienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record of managing complex AWS environments. * Proven Incident Commander: You demonstrate calm, decisive leadership during high-pressure outages. You have extensive experience running blameless postmortems and, crucially, driving the remediation work needed to prevent recurrence. * Observability Pro: You have deep experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs. * Automation Advocate: You believe that manual intervention is a bug. You have deep experience with Terraform and a "Code-First" approach to infrastructure. * Strategic Problem Solver: You can look at a complex, "needs-based" architecture and formulate a clear, prioritized roadmap to move it toward industry best practices. * Collaborative Leader: You enjoy working with feature engineers to help them build "reliability-by-design" into their services. * Education: A Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent professional experience)., AWS Lambda, Amazon Elastic Compute Cloud (EC2), Amazon Web Services (AWS), Applications Security, Architectural Services, Automation, Autoscaling, Best Practices, Business Continuity Planning (BCP), Capacity Management, Change Management, Cloud Computing, Computer Science, Continuous Deployment/Delivery, Continuous Integration, DevOps, Disaster Recovery, High Availability, Incident Response, Microsoft Dynamics, Network Architecture/Engineering, Philosophy, Problem Solving Skills, Redis, Reliability Engineering, Relocation Services, Reporting Dashboards, Salesforce.com, Strategic Planning, Systems Engineering, Team Player, Unix Shell Programming ## Description We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a builder - someone who wants to move beyond maintenance and into the realm of architectural transformation. You will have the autonomy to evaluate our existing AWS footprint and lead the charge in modernizing our environment. Your mission is to take a high-velocity system and implement the best practices, guardrails, and automated architectures that will support our next 10x of scale. You will be the primary authority on reliability, performance, and infrastructure security. Please note: This is a hybrid role based in our Santa Clara, CA office, with an in-office schedule of two days per week - Monday and Wednesday., * Architectural Modernization: Lead the design and implementation of a scalable, "Cloud-First" AWS architecture. You will drive the transition toward fully automated, state-of-the-art Infrastructure as Code (Terraform). * High Availability & Resilience: Design and implement robust Disaster Recovery (DR) and Business Continuity plans, moving our services toward a zero-downtime deployment model. * Performance & Capacity Engineering: Own the strategy for capacity planning and autoscaling. You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns with precision and cost-efficiency. * Advanced Observability: Define our monitoring and alerting philosophy using New Relic for deep APM and system insights. Partner this with IncidentIO to ensure we catch and resolve issues before they impact customers. * Streamlined CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to "production" safely and predictably. * Cloud Security: Harden our network architecture and application security posture, including WAF management and secure service-to-service communication. The Tech Stack * Cloud Infrastructure: AWS (EC2, Lambda, SQS, SNS, ALB, API Gateway, S3, WAF). * Observability & Incident Response:New Relic (APM/Infrastructure), IncidentIO. * Automation & Tools: Terraform, Redis/Elasticache, Shell Scripting, NPM/PM2. * Application Ecosystem: NodeJS, Python, C#, Angular, Apex. * Integration: Salesforce Managed Packages, MSFT Dynamics365. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)