> Markdown version of [/jobs/ext/225185-site-reliability-engineer-lead-applications-domains](https://www.wearedevelopers.com/jobs/ext/225185-site-reliability-engineer-lead-applications-domains). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, Lead. - Applications/Domains - **Company:** Toyota Motor Sales, U.S.A., Inc. - **Location:** Plano, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Macintosh Application Environment, Build Automation, Microsoft Azure, DevOps, Github, Monitoring of Systems, Python (Programming Language), Key Management, Reliability Engineering, Prometheus, Software Deployment, Software Engineering, Data Logging, Scripting, Google Cloud, Grafana, Reliability of Systems, Git Flow, Kubernetes, Cloudwatch, Terraform, Dynatrace, Serverless Computing, Elk Stack - **Published:** May 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f1428a371dcd3cf6 ## About the Role Do you have experience in Terraform?, * 6+ years of experience with DevOps tools like GitHub, Harness & Dynatrace. * 7+ years of experience building self-healing systems and automated remediation workflows. * 5+ years of experience in Site Reliability Engineering, DevOps, or related field. * Demonstrated experience in problem-solving, key SRE/DevOps concepts & tools with a proven track record of achieving high system reliability and performance. * Strong experience with Terraform for AWS IaC. * Proficient in scripting and automation with Python and familiar with monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack). * Deep knowledge of container orchestration (Kubernetes/EKS). * Deep understanding of cloud platforms (e.g., AWS, GCP, Azure) and container orchestration technologies (e.g., Kubernetes). * Effective communication skills, with the ability to convey complex technical concepts to diverse audiences. Added bonus if you have * AWS certifications (DevOps Engineer, Solutions Architect, etc.). * Familiarity with GitOps, secrets management, and infrastructure monitoring best practices. * Experience building self-healing systems and automated remediation workflows. ## Description Toyota Financial Services is building out a new Site Reliability Engineering (SRE) team for application domains, and we are seeking an SRE Lead to ensure reliability, performance and availability of the applications within each domain. As an SRE Lead - Applications/Domains, you will be working with development engineers, product owners, SRE Infrastructure, production engineers and Technology Operations Center personnel with a primary focus on improving observability, automation, overall system health, reliability and uptime. What you'll be doing * Design, code, and maintain automation to streamline operations, reduce manual tasks, and improve system efficiency to enable a robust application environment. * Working with observability engineers to enable actionable insights into applications and infrastructure health and performance. Foster a collaborative team-culture and support professional development. * Ensure scalable & repeatable code deployments with CI/CD pipelines using GitHub & Harness, repeatable deployments with infrastructure as code (IaC) using Terraform. * Build automation and operational runbooks primarily using Python scripting. * Manage container orchestration platforms and related cloud-native services. * Drive reliability improvements through Service Level Objectives (SLOs), error budgets, and Service Level Agreements (SLAs) aligned with business goals. * Design & implement observability improvements using Dynatrace & CloudWatch. * Lead major incident responses and coordinate with stakeholders for resolution and drive problem management to prevent recurrence. * Conduct blameless post-incident reviews and drive continuous improvement. * Collaborate cross-functionally to embed SRE principles into application design and operation meeting reliability goals. * Participate in architectural reviews, providing input on reliability and scalability. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)