> Markdown version of [/jobs/ext/1443631-cloud-systems-engineer](https://www.wearedevelopers.com/jobs/ext/1443631-cloud-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Systems Engineer - **Company:** CubeSmart - **Location:** Malvern, PA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Bash Shell, Ubuntu (Operating System), Cloud Computing Security, Cloud Engineering, Computer Programming, Computer Networks, Continuous Integration, Software Design Patterns, Linux, DevOps, Distributed Systems, Github, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, PostgreSQL, Octopus Deploy, OpenID, Redis, Reliability Engineering, Ansible, Prometheus, Zero Trust Network Access, Datadog, Scripting, Load Balancing, Cloud Platform System, Amazon ElastiCache, Grafana, Caching, Reliability of Systems, Containerization, Gitlab-ci, Kubernetes, Hashicorp, AWS Fargate, Cloudwatch, Terraform, Docker, Pagerduty, Jenkins, Static Application Security Testing, Microservices, Dynamic Application Security Testing - **Published:** July 25, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3331757807&tx=JR11407LFU&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role We are seeking a highly skilled Site Reliability & Cloud Systems Engineer to design, build, and operate scalable, secure, and highly automated cloud platforms in AWS. This role combines hands-on reliability engineering with cloud architecture and automation expertise, with a strong emphasis on building immutable infrastructure and improving system resilience., * 6-10+ years of experience in Site Reliability Engineering, DevOps, or Cloud Engineering roles * Deep hands-on expertise with AWS services and cloud architecture * Strong Linux systems engineering experience (Ubuntu preferred) * Proven experience with Infrastructure as Code (Terraform, Ansible, etc.) * Experience building and maintaining CI/CD pipelines * Proficiency in scripting/programming (Python, Bash) * Hands-on experience with monitoring and observability platforms * Solid understanding of cloud security principles (IAM, KMS, Secrets Management, Ansible Vault, Hashicorp Vault) * Bachelor's degree or equivalent practical experience * Candidates must be authorized to work in the U.S. without the need for current or future sponsorship., * Experience with containerization and orchestration (Docker, Kubernetes, EKS/ECS) * Familiarity with GitOps tools such as ArgoCD or Flux * Experience with SAST/DAST tools and secure SDLC practices * Knowledge of distributed systems, caching, and microservices architectures * Experience with FinOps and cost optimization strategies * Exposure to ITIL processes and service management platforms ## Description Reliability, Performance & Operations * Ensure uptime, reliability, and performance of AWS-hosted, Linux-based (Ubuntu) production systems and associated lower environments * Build and optimize observability using tools like Datadog, CloudWatch, Prometheus/Grafana, and PagerDuty * Working closely with the Dev teams, you will be diagnosing site issues, mitigating impact, and restoring system reliability while communicating clearly with stakeholders. * Lead incident response, root cause analysis, and post-incident reviews * Participate in on-call rotations and support 24/7 production environments Cloud Architecture & Automation * Architect and implement fully automated, ephemeral, and immutable AWS production and lower environments * Design scalable, resilient distributed systems using AWS best practices * Eliminate manual processes through Infrastructure as Code (Terraform, Ansible, Packer) * Build and maintain CI/CD and GitOps workflows (Jenkins, GitHub Actions, GitLab CI, ArgoCD/Flux) * Develop automation and tooling using Python and Bash to reduce operational toil Infrastructure & Platform Engineering * Deploy and manage AWS services including EKS, ECS, Fargate, Lambda, and RDS (Aurora PostgreSQL), Opensearch, Redis,Elasticache * Design and manage networking components such as Transit Gateways, load balancers, and service meshes * Implement caching, microservices, and distributed system design patterns Security & Governance * Architect and implement zero-trust security models using IAM, SCPs, and OIDC * Embed security into CI/CD pipelines using SAST/DAST tools (e.g., Snyk) * Ensure compliance through automated auditing, backup strategies, and governance controls Collaboration, Leadership & Strategy * Partner with development, security, and operations teams to build reliable, observable platforms * Document systems, runbooks, and operational procedures * Drive FinOps initiatives for cost optimization and forecasting * Integrate infrastructure changes into ITIL-compliant workflows (e.g., Freshservice) * Influence architectural decisions and promote engineering best practices across teams ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)