> Markdown version of [/jobs/ext/1891317-platform-infrastructure-engineer-sre-core](https://www.wearedevelopers.com/jobs/ext/1891317-platform-infrastructure-engineer-sre-core). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Infrastructure Engineer (SRE Core) - **Company:** Menlo Security - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $168,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Applicant Tracking Systems, Bash Shell, Cloud Computing, Configuration Management, Communications Protocols, Disaster Recovery, Domain Name System (DNS), Network Topologies, Identity and Access Management, Python (Programming Language), Software Architecture, Role-Based Access Control, Prometheus, TCP/IP, Scripting, Google Cloud, Istio, Large Language Models, Grafana, Mttr, Software Troubleshooting, Amazon Virtual Private Cloud (VPC), Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Golang - **Published:** August 2, 2026 - **Apply:** https://www.workingnomads.com/job/go/1767557/ ## About the Role * Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience * Proficiency in common programming and scripting languages, particularly Python, Bash, and Go * Understanding of network topologies, communication protocols (e.g., TCP/IP, HTTP/S, UDP, TLS), and enterprise-grade connectivity solutions * Kubernetes expertise, including cluster administration, RBAC, networking, workload management, and troubleshooting in production environments * Proven experience with Terraform for infrastructure provisioning and management * Knowledge of Google Cloud Platform services including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload Identity * Clear understanding of how to use LLM-based code-assist tools to effectively build and troubleshoot software Preferred / Nice to Have: * Experience with GitOps methodologies and tools ## Description Platform Infrastructure Engineering builds and operates Menlo Security's Infrastructure Platform, enabling our customers to connect to the Internet without compromise. As a Platform Infrastructure Engineer, you'll join a globally distributed team of experienced engineers building and managing the company's core infrastructure services on a cloud-native platform built on Google Kubernetes Engine and VMs spanning multiple regions and environments. The team manages infrastructure as code with Terraform and Spacelift, deploys with Helm, and emphasizes security-first design, comprehensive observability, and multi-region resilience. The team also uses AI-assisted development and code-review tools, including Gemini Code Assist, as part of the standard engineering workflow, and this role is expected to use LLM-based tooling to build and troubleshoot infrastructure code efficiently., * Reliable, secure, and scalable infrastructure across GCP and AWS supporting Menlo's platform globally. * Reduced operational toil and incident recurrence through automation and Infrastructure as Code practices. * Comprehensive, end-to-end observability framework providing deep platform visibility, proactive health monitoring, and accelerated incident detection and resolution. Success Metrics / KPIs: * Infrastructure uptime/availability across regions (e.g., 99.9%+) * Mean time to detect (MTTD) and mean time to resolve (MTTR) for incidents * Percentage of infrastructure changes deployed via IaC (Terraform) vs. manual changes * On-call incident volume and reduction in repeat/preventable incidents * Lead time for provisioning new infrastructure, * Implement, deploy, and maintain VM and Kubernetes infrastructure on GCP and AWS across dozens of clusters spanning development, staging, and production environments in multiple regions * Build and maintain Infrastructure as Code using Terraform modules and Spacelift (or equivalent TACOS), provisioning networking, compute, storage, and security components, and implementing multi-layer configuration management workflows * Implement and maintain observability solutions using Grafana Cloud, Prometheus/Mimir, and OTel collectors, designing dashboards and alerting rules across all platform components * Manage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with Cilium * Partner with peers and across Engineering, Product, Compliance, and Security teams to align on requirements and consult on capacity planning, disaster recovery, and architectural decisions * Identify and eliminate toil through automation - writing scripts, building CI/CD pipelines, and using AI-assisted coding tools to move faster * Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews, TO ALL AGENCIES: Please, no phone calls or emails to any employee of Menlo Security outside of the Talent organization. Menlo Security's policy is to only accept resumes from agencies via Ashby (ATS). Agencies must have a valid services agreement executed and must have been assigned by the Talent team to a specific requisition. Any resume submitted outside of this process will be deemed the sole property of Menlo Security. In the event a candidate submitted outside of this policy is hired, no fee or payment will be paid. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [You can’t hack what you can’t see](https://www.wearedevelopers.com/videos/41-you-can-t-hack-what-you-can-t-see) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology)