Cloud Systems Engineer

CubeSmart
Malvern, PA, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Bash Shell Ubuntu (Operating System) Cloud Computing Security Cloud Engineering Computer Programming Computer Networks Continuous Integration Software Design Patterns Linux DevOps Distributed Systems
+34 more
Github Monitoring of Systems Identity and Access Management Python (Programming Language) Key Management PostgreSQL Octopus Deploy OpenID Redis Reliability Engineering Ansible Prometheus Zero Trust Network Access Datadog Scripting Load Balancing Cloud Platform System Amazon ElastiCache Grafana Caching Reliability of Systems Containerization Gitlab-ci Kubernetes Hashicorp AWS Fargate Cloudwatch Terraform Docker Pagerduty Jenkins Static Application Security Testing Microservices Dynamic Application Security Testing

Job description

Reliability, Performance & Operations

  • Ensure uptime, reliability, and performance of AWS-hosted, Linux-based (Ubuntu) production systems and associated lower environments
  • Build and optimize observability using tools like Datadog, CloudWatch, Prometheus/Grafana, and PagerDuty
  • Working closely with the Dev teams, you will be diagnosing site issues, mitigating impact, and restoring system reliability while communicating clearly with stakeholders.
  • Lead incident response, root cause analysis, and post-incident reviews
  • Participate in on-call rotations and support 24/7 production environments

Cloud Architecture & Automation

  • Architect and implement fully automated, ephemeral, and immutable AWS production and lower environments
  • Design scalable, resilient distributed systems using AWS best practices
  • Eliminate manual processes through Infrastructure as Code (Terraform, Ansible, Packer)
  • Build and maintain CI/CD and GitOps workflows (Jenkins, GitHub Actions, GitLab CI, ArgoCD/Flux)
  • Develop automation and tooling using Python and Bash to reduce operational toil

Infrastructure & Platform Engineering

  • Deploy and manage AWS services including EKS, ECS, Fargate, Lambda, and RDS (Aurora PostgreSQL), Opensearch, Redis,Elasticache
  • Design and manage networking components such as Transit Gateways, load balancers, and service meshes
  • Implement caching, microservices, and distributed system design patterns

Security & Governance

  • Architect and implement zero-trust security models using IAM, SCPs, and OIDC
  • Embed security into CI/CD pipelines using SAST/DAST tools (e.g., Snyk)
  • Ensure compliance through automated auditing, backup strategies, and governance controls

Collaboration, Leadership & Strategy

  • Partner with development, security, and operations teams to build reliable, observable platforms
  • Document systems, runbooks, and operational procedures
  • Drive FinOps initiatives for cost optimization and forecasting
  • Integrate infrastructure changes into ITIL-compliant workflows (e.g., Freshservice)
  • Influence architectural decisions and promote engineering best practices across teams

Requirements

We are seeking a highly skilled Site Reliability & Cloud Systems Engineer to design, build, and operate scalable, secure, and highly automated cloud platforms in AWS. This role combines hands-on reliability engineering with cloud architecture and automation expertise, with a strong emphasis on building immutable infrastructure and improving system resilience., * 6-10+ years of experience in Site Reliability Engineering, DevOps, or Cloud Engineering roles

  • Deep hands-on expertise with AWS services and cloud architecture
  • Strong Linux systems engineering experience (Ubuntu preferred)
  • Proven experience with Infrastructure as Code (Terraform, Ansible, etc.)
  • Experience building and maintaining CI/CD pipelines
  • Proficiency in scripting/programming (Python, Bash)
  • Hands-on experience with monitoring and observability platforms
  • Solid understanding of cloud security principles (IAM, KMS, Secrets Management, Ansible Vault, Hashicorp Vault)
  • Bachelor’s degree or equivalent practical experience
  • Candidates must be authorized to work in the U.S. without the need for current or future sponsorship., * Experience with containerization and orchestration (Docker, Kubernetes, EKS/ECS)
  • Familiarity with GitOps tools such as ArgoCD or Flux
  • Experience with SAST/DAST tools and secure SDLC practices
  • Knowledge of distributed systems, caching, and microservices architectures
  • Experience with FinOps and cost optimization strategies
  • Exposure to ITIL processes and service management platforms

About the company

At CubeSmart, we’re intentional about culture. You can experience it everywhere from our mission statement of ā€œgenuine careā€ to our ā€œIt’s What’s Inside That Countsā€ tagline to calling each other ā€œteammatesā€ rather than employees. This spirit fosters a fun and collaborative environment that has resulted in our rapid growth and being recognized amongst the top in our industry.

CubeSmart’s award-winning team is made up of people who genuinely care. Teammates care about our customers and the life events and/or business needs they are facing. Teammates are passionate, responsible and understanding. The CubeSmart team is made up of people who have a can-do attitude, are committed to their own success and the success of the company, and lead by example.

If this sounds like a team and culture that matches your personal values and motivations, we want to hear from you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo Ā· LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard Ā· World Congress 2025

4:35 min

Setting up passwordless federated identity configuring OpenID Connect patterns

Marcel Lupo Ā· LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou Ā· Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello Ā· World Congress 2022

Videos

See all

Related articles

See all