SRE and Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+23 more
Job description
Seeking a highly experienced Technical Lead - SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of enterprise platforms and modern web applications.
This role is ideal for a senior engineer, technical lead, or architect who enjoys solving complex technical challenges, mentoring engineers, and influencing technical direction while remaining hands-on. The position offers a clear growth path into a future Technical Manager - SRE and Platform Engineering role as organizational needs and leadership responsibilities expand.
The Technical Lead will serve as a senior technical leader for Site Reliability Engineering (SRE), monitoring and observability, domain portfolio management, cloud platform operations, infrastructure automation, web application security, and application availability. The role will work closely with engineering teams in Dallas, Europe, and Asia Pacific to establish consistent operational standards, improve platform reliability, and drive technology modernization across global technology landscape.
The successful candidate will play a critical role in ensuring the reliability, performance, security, and operational health of Mary Kay’s externally facing digital platforms through ownership of key platform services including observability, domain services, DNS, certificate management, CDN technologies, web application firewalls, cloud platform infrastructure, and Infrastructure-as-Code (IaC) solutions.
KEY RESPONSIBILITIES Technical Leadership
- Provide technical leadership across Platform Engineering and Site Reliability Engineering functions.
- Establish engineering standards, operational best practices, and reliability objectives.
- Lead technical decision-making for cloud infrastructure, observability platforms, domain services, infrastructure automation, application security, and operational tooling.
- Mentor engineers and provide technical coaching across multiple disciplines.
- Drive technical roadmaps and continuous improvement initiatives.
- Evaluate, promote, and help operationalize emerging engineering capabilities, including AI-assisted development tools, coding agents, Infrastructure-as-Code automation, and other technologies that improve engineering productivity, quality, and speed of delivery.
- Site Reliability Engineering (SRE)
- Lead enterprise reliability initiatives focused on availability, scalability, resiliency, performance, and operational excellence.
- Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis.
- Drive automation initiatives that reduce operational overhead and improve service reliability.
- Serve as the technical owner for the enterprise monitoring and observability platform.
- Support APM, infrastructure monitoring, synthetic monitoring, Real User Monitoring (RUM), centralized logging, and distributed tracing.
- Define dashboards, alerting standards, operational metrics, and reporting.
- Lead governance of the company’s global domain portfolio.
- Manage registrations, renewals, DNS services, certificate lifecycle management, and related vendor relationships.
- Ensure domain-related services remain secure, compliant, and highly available.
- Provide technical leadership for AWS infrastructure, Kubernetes/EKS, CDN, DNS, SSL/TLS, load balancing, Web Application Firewalls (WAF), and edge security services.
- Design and support Infrastructure-as-Code solutions using Terraform.
- Establish standards for cloud provisioning, automation, and environment consistency.
- Design, maintain, and optimize Terraform modules and deployment pipelines.
- Promote automated provisioning, version control, testing, and infrastructure governance.
- Drive reduction of manual deployment activities through automation.
- Administer and optimize AWS WAF and comparable WAF technologies.
- Manage WAF rules, rate limiting, bot protection, IP reputation controls, and application-layer threat mitigation.
- Partner with Information Security to improve web application protection capabilities.
- Serve as a senior escalation point for complex production issues.
- Lead troubleshooting across client-side and server-side technologies.
- Diagnose issues involving browser behavior, APIs, DNS, CDN, WAF, load balancing, networking, cloud infrastructure, and application performance.
- Drive reliability, resiliency, and end-user experience improvements.
- Work closely with engineering teams across North America, Europe, and Asia Pacific.
- Participate in technical reviews, architecture discussions, operational planning, and knowledge sharing.
- Help evolve a follow-the-sun operating model.
- Participate in scheduled on-call rotations supporting critical platforms and services.
- Provide leadership during major incidents and after-hours escalations.
- Support maintenance, upgrades, deployments, and disaster recovery activities.
Requirements
- 8+ years supporting enterprise applications, cloud platforms, infrastructure services, or web technologies.
- 3+ years serving as a Technical Lead, Senior Engineer, Architect, or equivalent.
- Strong experience with AWS, Terraform, Kubernetes/EKS, DNS, CDN, AWS WAF, SSL/TLS, observability platforms, CI/CD, and SRE practices.
- Experience troubleshooting large-scale customer-facing web applications.
- Experience managing global domain portfolios.
- Experience with enterprise observability platforms and global support models.
- Experience developing enterprise Infrastructure-as-Code frameworks and reusable Terraform modules.
- Experience leveraging AI-assisted development tools and coding agents such as GitHub Copilot, Microsoft Copilot, Claude Code, Cursor, Amazon Q Developer, or similar technologies to accelerate software delivery, infrastructure automation, troubleshooting, and operational efficiency.
- AWS, Terraform, Kubernetes, SRE, networking, security, or cloud certifications.
Benefits & conditions
Pulled from the full job description
- Referral program
- 401(k)
- Health insurance
- Vision insurance
- Health savings account
- Dental insurance
- Employee assistance program, Estimated Min Rate: $56.00 Estimated Max Rate: $80.00
What’s In It for You? We welcome you to be a part of the largest and legendary global staffing companies to meet your career aspirations. Yoh’s network of client companies has been employing professionals like you for over 65 years in the U.S., UK and Canada. Join Yoh’s extensive talent community that will provide you with access to Yoh’s vast network of opportunities and gain access to this exclusive opportunity available to you. Benefit eligibility is in accordance with applicable laws and client requirements. Benefits include:
- Medical, Prescription, Dental & Vision Benefits (for employees working 20+ hours per week)
- Health Savings Account (HSA) (for employees working 20+ hours per week)
- Life & Disability Insurance (for employees working 20+ hours per week)
- MetLife Voluntary Benefits
- Employee Assistance Program (EAP)
- 401K Retirement Savings Plan
- Direct Deposit & weekly epayroll
- Referral Bonus Programs
- Certification and training opportunities
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
The Best X (Twitter) Accounts for Developers
Dev Digest 120 - Apple and peers
Why Upskilling And Reskilling is Important For Developers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again