> Markdown version of [/jobs/ext/3364108-principal-cloud-engineer](https://www.wearedevelopers.com/jobs/ext/3364108-principal-cloud-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal - Cloud Engineer - **Company:** Ally Financial Inc. - **Location:** Charlotte, NC, United States (Remote available) - **Experience:** Expert - **Salary:** $110,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Cloudfront, Amazon Elastic Compute Cloud, Amazon S3, User Authentication, Cloud Engineering, Information Systems, Databases, Continuous Integration, Data Migration, DevOps, Disaster Recovery, Amazon DynamoDB, Failover, Identity and Access Management, Python (Programming Language), Network Layer, OpenID, Role-Based Access Control, Reliability Engineering, Migration Manager, Backup and Restore, Load Balancing, Amazon Virtual Private Cloud (VPC), Gitlab, Cloudformation, Servicebus, Database Migration, Amazon Relational Database Service, Gitlab-ci, Information Technology, Performance Monitor, Route53, Cloudwatch, Api Gateway, Amazon Simple Queue Service (SQS), Terraform, Splunk, Dynatrace, Serverless Computing, Vulnerability Analysis - **Published:** September 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=cf5d88997af101a0 ## About the Role * Bachelor's degree in Computer Science, Information System, or other relevant field of study preferred * 5+ years of hands-on experience in AWS cloud engineering or site reliability engineering in a large-scale, regulated environment * Expert knowledge of AWS reliability services: Lambda, ECS, RDS/Aurora, DynamoDB, S3, SQS, EventBridge, CloudWatch, Route 53, CloudFront, API Gateway, and VPC networking * Hands-on experience with AWS Application Recovery Controller (ARC), AWS Resilience Hub, and AWS Fault Injection Service (FIS) - including readiness checks, routing controls, resiliency scoring, and chaos experiment design across compute, database, and network layers * Experience designing multi-region architectures across Active/Active, Warm Standby, Pilot Light, and Backup and Restore patterns * Expert-level AWS IAM knowledge including least-privilege design, permission boundaries, federation, and Service Control Policies (SCPs) * Proficiency in Terraform for infrastructure as code and strong CI/CD experience with GitLab including pipeline design, OIDC authentication, and deployment safety controls * Experience with AWS migration services: Migration Hub, Application Migration Service, Database Migration Service, and Application Discovery Service * Proficiency in Python for automation and scripting; experience with Amazon CloudWatch, Dynatrace, Splunk, and distributed tracing * Experience with FinOps practices including cost allocation, tagging, rightsizing, Savings Plans, and consumption forecasting * Ability to lead resilience reviews, production readiness assessments, and architecture decisions; strong communication skills to translate technical recovery risks into executive-ready reporting and business outcomes * Comfortable operating as both a principal individual contributor and a technical leader for distributed engineering teams * Preferred Certifications: AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, or equivalent advanced AWS certifications ## Description Seeking an innovative, hands-on Principal Cloud Engineer who excels at the intersection of cloud resilience engineering, infrastructure automation, and technical leadership. In this role, you will define and drive enterprise-wide cloud reliability strategy - designing multi-region architectures, leading chaos engineering programs, building recovery tooling, and partnering with application teams to achieve measurable, validated resiliency outcomes in a regulated financial services environment. The Work Itself * Lead architecture discovery sessions with application, infrastructure, security, and business stakeholders to define recovery objectives, dependency maps, and multi-region requirements * Design multi-region AWS architectures - Active/Active, Warm Standby, Pilot Light, and Backup and Restore - aligned to application criticality, RTO/RPO targets, and compliance requirements * Design and implement AWS Application Recovery Controller (ARC) patterns - readiness checks, routing controls, safety rules, and zonal shift - to enable deliberate, auditable traffic management and verify recovery environments are prepared before any failover event * Evaluate AWS Resilience Hub to identify resilience gaps, map recovery policies, and drive remediation to measurable objectives * Define traffic management and failover strategies using Amazon Route 53, CloudFront, load balancers, API Gateway, and health checks * Design backup and recovery patterns using AWS Backup, S3, RDS/Aurora cross-region capabilities, encryption, and restore validation procedures * Produce architecture decision records documenting service selection, alternatives considered, constraints, and expected outcomes * Define secure AWS landing-zone controls and design IAM roles, permission boundaries, federation, and least-privilege access models for applications, automation, and platform services * Establish infrastructure and application-performance monitoring using Amazon CloudWatch, Dynatrace, and distributed tracing; build log forwarding to enterprise analytics platforms including Splunk; design operational dashboards connecting infrastructure health, security findings, and recovery readiness indicators * Develop FinOps operating models covering account ownership, tagging, chargeback, budgets, and anomaly detection; forecast cloud consumption and build cost-optimization business cases covering rightsizing, Savings Plans, and Reserved Instances * Evaluate migration strategy patterns and design migration approaches using EC2, VPC, S3, RDS, CloudFormation, containers, and serverless services aligned to business criticality and RTO/RPO targets * Use AWS Migration Hub, Application Migration Service, and Database Migration Service to coordinate and execute migrations; define data migration, validation, reconciliation, rollback, and synchronization procedures aligned to RTO/RPO requirements with tested cutover plans * Validate migrated workloads through functional, performance, backup, restore, and disaster-recovery testing * Create reusable Terraform modules for VPC, IAM, ECS, RDS, S3, and monitoring; design GitLab CI/CD pipelines with validation, security scanning, plan review, approval controls, and drift detection * Implement blue-green, canary, rolling, and controlled failover deployment patterns; build deployment safety controls including pre-deployment validation, health gates, blast-radius reduction, and automated rollback procedures * Integrate cost-estimation tooling into pipelines and automate account and regional deployment workflows while preserving isolation, least privilege, and clear ownership * Design and run controlled chaos engineering experiments using AWS Fault Injection Service (FIS) targeting compute, database, network, and dependency layers with defined hypotheses, abort criteria, and observability instrumentation * Create game days and disaster recovery exercises involving application owners, infrastructure, cybersecurity, and business stakeholders; develop recovery runbooks for AZ impairment, regional degradation, database failure, and capacity exhaustion; establish a recovery-testing calendar covering backup restoration, failover, and failback * Define resilience engineering standards covering dependency mapping and failure-domain analysis; use experiment results to surface single points of failure and drive remediation with risk owners; coach application teams on resilience patterns and operational ownership * Create executive-ready reporting summarizing recovery readiness, experiment coverage, open resilience risks, and reliability trends for leadership ## Related Videos - [GitLab CI pipelines for a whole company](https://www.wearedevelopers.com/videos/143-gitlab-ci-pipelines-for-a-whole-company) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Delegating the chores of authenticating users to Keycloak](https://www.wearedevelopers.com/videos/1558-delegating-the-chores-of-authenticating-users-to-keycloak) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Seriously gaming your cloud expertise: from cloud tourist to cloud native](https://www.wearedevelopers.com/videos/373-seriously-gaming-your-cloud-expertise-from-cloud-tourist-to-cloud-native) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology)