> Markdown version of [/jobs/ext/459329-sr-director-platform-engineering-sre](https://www.wearedevelopers.com/jobs/ext/459329-sr-director-platform-engineering-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Director, Platform Engineering & SRE - **Company:** Brand Ministry LLC - **Location:** Alpharetta, GA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Software as a Service, Cloud Computing, Cloud Engineering, Continuous Integration, Monitoring of Systems, IT Management, PCI Data Security Standards, Reliability Engineering, Prometheus, Datadog, Google Cloud, Grafana, Mttr, Multi-Cloud, Information Technology - **Published:** June 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=589aef5d1efe4d29 ## About the Role Do you have experience in Technology management?, * Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience * 10+ years of overall experience in software, infrastructure, or platform engineering * 6+ years in engineering leadership or management roles * Demonstrated track record building or scaling a Site Reliability Engineering or Platform Engineering function and improving availability/reliability outcomes * Deep, hands-on cloud experience at SaaS scale (Azure and/or AWS), including infrastructure-as-code and CI/CD * Strong background across: Site Reliability Engineering, Observability & Monitoring, Cloud & Infrastructure Engineering, Incident & Performance Management, Capacity Planning, and Production Operations, * Experience operating multi-cloud and multi-tenant SaaS environments * Hands-on implementation of SLO/error-budget frameworks and modern observability tooling (e.g., Datadog, Grafana, Prometheus, OpenTelemetry) * Experience standing up reliability in a distributed or embedded (product-team) model * Exposure to SOC 2 and PCI DSS 4.0.1 control evidence at the infrastructure layer * Background in a private-equity-backed or high-growth SaaS environment * Demonstrated business acumen and sound decision-making in complex, multi-product environments Benefit offerings designed to promote a life of balance! ## Description As Sr. Director, Platform Engineering & SRE, you will build and lead the function responsible for the reliability, performance, and operational excellence of the Ministry Brands platform. You will own site reliability engineering, observability, production operations, and cloud engineering across our multi-cloud SaaS portfolio - establishing the practices, tooling, and standards that keep our products available and performant for the organizations we serve. This is a hands-on leadership role at the center of our most important technical priority: platform stability. You will define and drive measurable improvements in availability and incident response, stand up a modern SRE discipline, and partner closely with R&D, Product, and Security leaders to embed reliability into how we build and operate software. You will be accountable to executive leadership for platform availability and performance., Reliability & SRE Practice * Establish and own service-level objectives (SLOs), service-level indicators (SLIs), and error-budget policy across the product platform * Lead incident command, on-call rotation, escalation, and a blameless postmortem culture; drive measurable reduction in MTTR and change-failure rate * Set reliability standards and partner with embedded reliability engineers in R&D product teams to apply them at the point of system design * Drive availability toward enterprise targets and own the reliability roadmap and its reporting to executive stakeholders Operations & Observability * Build and operate the observability platform - metrics, logs, traces, and alerting - and define the golden signals and dashboards used across products * Lead capacity planning, performance engineering, and operational-readiness reviews for new and existing services * Own production operations practices, runbooks, and escalation workflows that improve transparency, stability, and stakeholder communication * Deliver metrics-based reporting on platform availability and performance Cloud & Platform Engineering * Lead cloud engineering across our multi-cloud footprint (Azure, AWS, GCP), balancing reliability, performance, security posture, and cost * Own infrastructure-as-code, CI/CD platform standards, and the internal developer platform that product teams build on * Drive consolidation and standardization of fragmented infrastructure and pipeline tooling * Partner with Security to implement and evidence platform-layer controls in support of SOC 2 and PCI DSS objectives Leadership & Stakeholder Management * Define team culture and objectives aligned to Enterprise IT & Security strategic goals; build, coach, and develop the Platform Engineering & SRE team * Build and maintain strong partnerships with R&D, Product, Security, and IT leaders * Develop and manage the platform engineering budget, balancing run-the-business needs with strategic investment, and author clear business cases for technology investments * Manage key cloud and tooling vendor relationships in partnership with IT and Procurement * Present updates, metrics, and recommendations to both technical and business stakeholders ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)