> Markdown version of [/jobs/ext/2643574-sre-and-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2643574-sre-and-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE and Platform Engineer - **Company:** Yoh Services LLC - **Location:** Addison, TX, United States - **Experience:** Expert - **Salary:** $116,480.0 - $166,400.0 - **Contract:** Temporary to permanent - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Application Firewall, Application Performance Management, Cloud Computing, Cyber Security, Continuous Integration, Cursor (Graphical User Interface Elements), Disaster Recovery, Domain Name System (DNS), Monitoring of Systems, Reliability Engineering, Site Reliability Engineering Practices, Web Application Security, Web Applications, Web Platforms, Automatic Programming, SSL Certificate Management, Data Logging, Transport Layer Security, Enterprise Software Applications, Load Balancing, Cloud Platform System, Microsoft Power Automate, GitHub Copilot, Delivery Pipeline, Software Security, Rate Limiting, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Web Technologies, Terraform, Software Version Control, Dynatrace - **Published:** August 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=79f9cde8789824e7 ## About the Role * 8+ years supporting enterprise applications, cloud platforms, infrastructure services, or web technologies. * 3+ years serving as a Technical Lead, Senior Engineer, Architect, or equivalent. * Strong experience with AWS, Terraform, Kubernetes/EKS, DNS, CDN, AWS WAF, SSL/TLS, observability platforms, CI/CD, and SRE practices. * Experience troubleshooting large-scale customer-facing web applications. * Experience managing global domain portfolios. * Experience with enterprise observability platforms and global support models. * Experience developing enterprise Infrastructure-as-Code frameworks and reusable Terraform modules. * Experience leveraging AI-assisted development tools and coding agents such as GitHub Copilot, Microsoft Copilot, Claude Code, Cursor, Amazon Q Developer, or similar technologies to accelerate software delivery, infrastructure automation, troubleshooting, and operational efficiency. * AWS, Terraform, Kubernetes, SRE, networking, security, or cloud certifications. ## Description Seeking a highly experienced Technical Lead - SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of enterprise platforms and modern web applications. This role is ideal for a senior engineer, technical lead, or architect who enjoys solving complex technical challenges, mentoring engineers, and influencing technical direction while remaining hands-on. The position offers a clear growth path into a future Technical Manager - SRE and Platform Engineering role as organizational needs and leadership responsibilities expand. The Technical Lead will serve as a senior technical leader for Site Reliability Engineering (SRE), monitoring and observability, domain portfolio management, cloud platform operations, infrastructure automation, web application security, and application availability. The role will work closely with engineering teams in Dallas, Europe, and Asia Pacific to establish consistent operational standards, improve platform reliability, and drive technology modernization across global technology landscape. The successful candidate will play a critical role in ensuring the reliability, performance, security, and operational health of Mary Kay's externally facing digital platforms through ownership of key platform services including observability, domain services, DNS, certificate management, CDN technologies, web application firewalls, cloud platform infrastructure, and Infrastructure-as-Code (IaC) solutions. KEY RESPONSIBILITIES Technical Leadership * Provide technical leadership across Platform Engineering and Site Reliability Engineering functions. * Establish engineering standards, operational best practices, and reliability objectives. * Lead technical decision-making for cloud infrastructure, observability platforms, domain services, infrastructure automation, application security, and operational tooling. * Mentor engineers and provide technical coaching across multiple disciplines. * Drive technical roadmaps and continuous improvement initiatives. * Evaluate, promote, and help operationalize emerging engineering capabilities, including AI-assisted development tools, coding agents, Infrastructure-as-Code automation, and other technologies that improve engineering productivity, quality, and speed of delivery. * Site Reliability Engineering (SRE) * Lead enterprise reliability initiatives focused on availability, scalability, resiliency, performance, and operational excellence. * Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis. * Drive automation initiatives that reduce operational overhead and improve service reliability. * Serve as the technical owner for the enterprise monitoring and observability platform. * Support APM, infrastructure monitoring, synthetic monitoring, Real User Monitoring (RUM), centralized logging, and distributed tracing. * Define dashboards, alerting standards, operational metrics, and reporting. * Lead governance of the company's global domain portfolio. * Manage registrations, renewals, DNS services, certificate lifecycle management, and related vendor relationships. * Ensure domain-related services remain secure, compliant, and highly available. * Provide technical leadership for AWS infrastructure, Kubernetes/EKS, CDN, DNS, SSL/TLS, load balancing, Web Application Firewalls (WAF), and edge security services. * Design and support Infrastructure-as-Code solutions using Terraform. * Establish standards for cloud provisioning, automation, and environment consistency. * Design, maintain, and optimize Terraform modules and deployment pipelines. * Promote automated provisioning, version control, testing, and infrastructure governance. * Drive reduction of manual deployment activities through automation. * Administer and optimize AWS WAF and comparable WAF technologies. * Manage WAF rules, rate limiting, bot protection, IP reputation controls, and application-layer threat mitigation. * Partner with Information Security to improve web application protection capabilities. * Serve as a senior escalation point for complex production issues. * Lead troubleshooting across client-side and server-side technologies. * Diagnose issues involving browser behavior, APIs, DNS, CDN, WAF, load balancing, networking, cloud infrastructure, and application performance. * Drive reliability, resiliency, and end-user experience improvements. * Work closely with engineering teams across North America, Europe, and Asia Pacific. * Participate in technical reviews, architecture discussions, operational planning, and knowledge sharing. * Help evolve a follow-the-sun operating model. * Participate in scheduled on-call rotations supporting critical platforms and services. * Provide leadership during major incidents and after-hours escalations. * Support maintenance, upgrades, deployments, and disaster recovery activities. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)