> Markdown version of [/jobs/ext/512292-cloud-operations-engineer-infrastructure](https://www.wearedevelopers.com/jobs/ext/512292-cloud-operations-engineer-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Operations Engineer - Infrastructure - **Company:** TP-Link Systems Inc. - **Location:** Irvine, CA, United States - **Experience:** Experienced - **Salary:** $100,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Microsoft Azure, Cloud Computing, Cloud Computing Security, Configuration Management, Cyber Security, Continuous Integration, Linux, Distributed Systems, Identity and Access Management, Python (Programming Language), Linux System Administration, Routing, Reliability Engineering, Service Discovery, Software Engineering, Autoscaling, Istio, Delivery Pipeline, Amazon Virtual Private Cloud (VPC), Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Golang - **Published:** June 6, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=2b11366ecba08979 ## About the Role Do you have experience in Technology security practices?, Do you have a Bachelor's degree?, * Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related field. * 2+ years of hands-on experience in cloud infrastructure, Kubernetes operations, platform engineering, SRE, or related areas. * Strong knowledge of AWS services, including EKS, IAM, VPC, EC2, S3, and related networking and security capabilities. * Hands-on experience operating Kubernetes in production environments, including cluster architecture, workload orchestration, networking, autoscaling, and troubleshooting. * Familiarity with Kubernetes ecosystem tools such as CRDs, Helm, Cluster API, HPA, Cluster Autoscaler, and CoreDNS. * Experience with GitOps tools such as FluxCD or ArgoCD. * Solid Linux administration and troubleshooting skills, including systemd, networking, and performance analysis. * Experience with CI/CD pipelines and infrastructure automation using Terraform, Go, Python, or similar tools. * Good understanding of reliability engineering practices, including SLOs, incident response, monitoring, alerting, and post-mortems. * Strong problem-solving skills and ability to diagnose and resolve complex infrastructure issues in distributed systems. * Good communication skills and ability to collaborate effectively with cross-functional engineering teams. * Willingness to participate in a scheduled on-call rotation., * Experience with NVIDIA device plugins, GPU scheduling, or GPU workload operations in Kubernetes environments. * Experience with additional public cloud platforms such as Azure or Alibaba Cloud. * Kubernetes certifications such as CKA, CKAD, or CKS are a plus. ## Description * Design, build, and maintain reliable, scalable, and secure cloud-native infrastructure platforms supporting large-scale production workloads. * Operate and optimize multi-account AWS environments, ensuring infrastructure is secure, repeatable, and auditable through Infrastructure as Code tools such as Terraform. * Manage production Kubernetes clusters, including provisioning, upgrades, autoscaling, networking, observability, capacity planning, and day-to-day operations. * Build and operate Kubernetes ecosystem components such as CRDs, Helm, HPA, Cluster Autoscaler, CoreDNS, and Cluster API. * Operate and improve GitOps-based deployment workflows using tools such as FluxCD or ArgoCD. * Manage and enhance Istio service mesh capabilities, including traffic routing, service discovery, resilience, security, and service-to-service communication. * Define and improve reliability practices, including SLOs, Error Budgets, monitoring, alerting, incident response, and post-mortems. * Participate in a scheduled on-call rotation to support production cloud infrastructure and Kubernetes platforms. * Troubleshoot complex production issues across cloud infrastructure, Kubernetes, Linux systems, networking, and distributed services. * Drive automation for infrastructure provisioning, configuration management, CI/CD pipelines, observability, and operational workflows using Terraform, Go, Python, or similar technologies. * Collaborate with application engineering, architecture, security, and platform teams to improve infrastructure reliability, scalability, and operational efficiency. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)