AWS Cloud Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+33 more
Job description
Deploy, manage, and troubleshoot workloads across AWS EKS and Azure AKS environments. Design and maintain Kubernetes clusters throughout their lifecycle. Manage Kubernetes networking, ingress controllers, service meshes, pod communication, and traffic routing. Troubleshoot complex issues involving pods, nodes, networking, DNS, storage, and cluster scalability. Implement high availability and disaster recovery strategies for Kubernetes platforms. GitOps & Continuous Delivery
Design and manage enterprise-scale GitOps workflows using Argo CD. Automate deployments and environment promotion strategies. Build and maintain CI/CD pipelines using GitHub Actions. Define and enforce Git branching strategies and release management practices. Improve deployment reliability, rollback capabilities, and release governance. AWS Cloud Engineering Strong hands-on experience with:
EKS EC2 Route 53 IAM CloudFront/CDN ALB & NLB VPC & Networking Security Groups & NACLs Secrets Manager ElastiCache (Redis) S3 CloudWatch RDS Autoscaling & High Availability Architectures Responsibilities include:
Designing secure and scalable cloud architectures. Implementing disaster recovery and business continuity solutions. Optimizing cloud cost, performance, and reliability. Managing multi-account AWS environments. Infrastructure as Code
Build and maintain reusable Terraform modules. Provision cloud infrastructure using Infrastructure as Code best practices. Manage environment consistency and compliance through automation. Troubleshoot Terraform state, drift, and large-scale deployments. Platform Automation & Engineering Excellence
Continuously identify manual processes that can be automated. Develop solutions for: Deployment automation Infrastructure automation Log monitoring automation Incident response automation Self-service engineering capabilities Build internal developer tooling to improve engineering productivity. Observability & Monitoring
Design monitoring and alerting solutions using platforms such as Datadog. Implement automated alert correlation and incident reduction mechanisms. Create dashboards, SLOs, and observability standards. Drive proactive monitoring practices across production environments. Disaster Recovery & Resiliency
Design highly resilient cloud architectures. Lead recovery planning for major production outages. Demonstrate expertise in scenarios such as: Region failures Kubernetes cluster failures Account compromises Infrastructure loss events Define recovery strategies using Infrastructure as Code, backups, replication, and automation.
Requirements
4+ years in DevOps, SRE, Platform Engineering, or Cloud Engineering. Expert-level Kubernetes administration. Strong experience with EKS and/or AKS. Deep understanding of Argo CD and GitOps principles. Strong GitHub Actions experience. Advanced Git branching and release management knowledge. Strong AWS architecture and operational expertise. Terraform expertise. Containerization experience (Docker/Kubernetes). Linux system administration. Cloud networking and security. Preferred
Azure experience. Service Mesh technologies. Multi-cloud deployments. Datadog or similar observability platforms. Python, Go, Shell scripting, or automation development. Platform Engineering experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers