> Markdown version of [/jobs/ext/2534927-director-of-cloud-sre](https://www.wearedevelopers.com/jobs/ext/2534927-director-of-cloud-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director of Cloud SRE - **Company:** Ford Motor Company - **Location:** Springfield, IL, United States (Remote available) - **Experience:** Experienced - **Salary:** $141,700.0 - $268,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Data Centers, Programming Tools, Reliability Engineering, Datadog, Information Technology, Enterprise Integration, Virtual Agents, Terraform, Splunk, New Relic (SaaS), Software Version Control, Dynatrace - **Published:** August 23, 2026 - **Apply:** https://www.businessworkforce.com/job.asp?id=3363340322&tx=DT10596UYT&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience. * 10+ years of experience in Site Reliability Engineering, platform engineering, or infrastructure engineering * 4+ years in a people leadership role managing engineering leaders and/or engineers. * Demonstrated experience building and operating observability platforms at scale, with hands-on depth in OpenTelemetry and at least one enterprise observability platform (Dynatrace, Datadog, New Relic, Splunk, or similar). * Proven track record designing and delivering internally-built developer tooling, including integration with CI/CD pipelines, source control platforms, and developer workflows. * Strong working knowledge of public cloud architecture (GCP strongly preferred; AWS/Azure acceptable) and demonstrated ability to extend reliability practices into hybrid or on-premise environments. * Deep understanding of core SRE principles: SLIs/SLOs, error budgets, incident management and postmortem practice, toil reduction, capacity planning, and reliability-by-design. * Experience operating in environments with heterogeneous infrastructure - cloud, data center, and OT/manufacturing or industrial environments a strong plus. * Demonstrated ability to build a long-term technical strategy and translate it into an executable roadmap, balancing tactical remediation against multi-year platform investment. * Strong executive communication skills - able to represent technical strategy to senior leadership and align cross-functional stakeholders around a unified direction. * Experience with infrastructure-as-code (Terraform or equivalent) and modern software delivery practices (agile/PI planning experience a plus). Preferred: * Experience building or scaling an SRE function within a large, matrixed enterprise. * Familiarity with emerging AI/agentic observability standards (OTel GenAI semantic conventions) and their application to platform tooling. * Experience establishing SRE maturity models or capability frameworks used to guide staged team adoption. This role requires up to 10% travel. ## Description This is a builder's role as much as a leader's role. You will guide a team that develops internally-owned tooling built intentionally to remain vendor-agnostic so the organization is never architecturally locked to a single observability provider and drives Agentic AI deeper into our SRE ecosystem. While our primary application runtime is GCP, this leader must be equally comfortable partnering across the SRE organization to extend reliability and observability standards into data center, manufacturing, distribution, and global campus environments - meeting engineering and operations teams where they are, not just where the platform lives. The ideal candidate blends technical depth with organizational fluency: someone who can sit in an architecture review and a roadmap planning session with equal credibility, who has personally built and operated production systems, and who can partner effectively across a broader SRE leadership team to advance a long-term, unified observability strategy. * Partner with fellow SRE leaders to define and drive a multi-year, holistic strategy for unified observability and SRE platform offerings, spanning cloud-native (GCP) and on-premise (data center, manufacturing, distribution, campus) environments. * Lead, develop, and grow a team of engineering managers/leads and individual contributor engineers, building organizational depth in SRE practice and platform engineering. * Drive the roadmap for internally-built observability tooling, ensuring architecture remains vendor-agnostic and portable across telemetry backends (OpenTelemetry-first design, current integration with Dynatrace) with a focus on Agentic AI platforms to simplify correlation data. * Help federate core SRE principles - SLIs/SLOs, error budgets, incident management, toil reduction, capacity and reliability engineering - across application and platform teams enterprise-wide, working alongside peer SRE leaders rather than centralizing reliability as a bottleneck. * Partner with developer experience and platform engineering teams to embed observability and reliability tooling directly into CI/CD pipelines and source repositories, shifting reliability left in the development lifecycle. * Contribute to an SRE maturity model, providing application teams a clear, staged path to deepen their own reliability practice with SRE org support and self-service tooling. * Build cross-domain relationships with manufacturing, plant, and OT engineering leadership, in partnership with other SRE leaders, to extend reliability and observability discipline into environments with materially different constraints (legacy protocols, air-gapped or constrained networks, safety-critical operations). * Represent SRE platform direction to senior technology leadership, including architecture governance bodies, and act as an escalation point for major reliability and observability initiatives within your team's scope. * Contribute to vendor relationship and technology decisions related to observability tooling, balancing build-vs-buy tradeoffs against long-term platform and cost strategy. * Ensure the team maintains hands-on technical currency - reviewing designs, contributing to architecture decisions, and staying credible as a technical leader. ## Related Videos - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)