Systems Architect, Disaster Recovery

PDS Inc.
Los Angeles, CA, United States
22 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$260,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Engineering Databases System Configuration Disaster Recovery Failover Hyper-V IBM Cloud Computing Identity and Access Management Systems Development Life Cycle Reliability Engineering
+23 more
Cloud Services Software Engineering Systems Architecture Virtualization Technology Scripting Google Cloud Cloud Platform System System Availability Mttr Multi-Cloud HybridCloud Gitlab Cloudformation Kubernetes Infrastructure Automation Frameworks Information Technology Database Mirroring Veeam Terraform Serverless Computing Legacy Systems Jenkins Vmware

Job description

Designs, develops and oversees system architectures, including complex systems and systems design activities. Serves as focal point for the development and communication of the system architecture. Ensures conversion of mission requirements into total systems solutions that account for design and technology maturity constraints of the system. Develops systems and system element architecture, design, and interface definition. Supports internal and external design reviews. Maintains knowledge of current and developing technologies and design and analysis methodologies. Develops models and architectural guidelines for current and future system development., Architectural Leadership

  • Define end to end DR and high availability (HA) architectures for enterprise wide workloads, incorporating multi region cloud, hybrid, and on prem solutions.

  • Develop architectural blueprints, reference designs, and pattern libraries that align with LM’s security, compliance, and cost optimization policies.

Solution Design & Implementation

  • Design and implement automated fail over, replication, and fail back mechanisms (e.g., Site Recovery Manager, Kubernetes based HA, database mirroring, storage level replication).

  • Evaluate and integrate emerging technologies (e.g., Immutable Infrastructure, Chaos Engineering, Serverless DR) to improve resiliency and reduce mean time to recover (MTTR).

Governance & Compliance

  • Ensure all DR solutions meet corporate policies CRX 301, CRX 302, and relevant regulatory requirements (e.g., NIST?800 34, ISO?22301, FedRAMP).

  • Create and maintain DR documentation, run books, and test plans; conduct periodic reviews and updates.

Testing & Validation

  • Lead full scale DR exercise planning, execution, and post mortem analysis for multi site, multi cloud environments.

  • Define success criteria, metrics, and KPIs; report findings to senior leadership and stakeholders.

Stakeholder Collaboration

  • Partner with IT Infrastructure, Cloud Engineering, Application Development, Security, and Governance teams to embed DR/HA considerations early in the SDLC.

  • Serve as the technical authority for DR during design reviews (SRR, PDR, CDR, TRR) and program risk assessments.

Continuous Improvement

  • Conduct risk assessments, threat modeling, and capacity planning to anticipate emerging resiliency challenges.

  • Drive adoption of Model Based Systems Engineering (MBSE) and automated documentation tools to keep architecture artefacts current.

DR Plan Modernization & Compliance

  • Review existing DR plan architectures across the enterprise, assessing their alignment with current resilience standards, best practices, and organizational Recovery Objectives.

  • Collaborate with internal teams (Application Owners, IT Service Managers, Engineering) to update and refine DR plans, ensuring that all applications and IT services meet the latest RTO/RPO targets.

  • Develop and implement remediation plans to bring legacy systems and applications up to date with modern resilience standards, ensuring compliance with corporate policies (CRX 301, CRX 302) and regulatory requirements.

  • Track progress and report status to senior leadership, providing insights into plan modernization efforts and risk mitigation strategies.

Requirements

5?+?years of experience designing and implementing DR/HA solutions for enterprise scale workloads in cloud, hybrid, and on prem environments.

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical discipline (or equivalent experience).

  • Proven hands on experience with cloud platforms (AWS, Azure, GCP) and related services (e.g., Disaster Recovery, Site Recovery Manager, cross region replication, networking, IAM).

  • Strong understanding of networking, storage, virtualization, container orchestration (Kubernetes), and database technologies as they relate to resiliency.

  • Excellent written and verbal communication skills; ability to translate complex technical concepts for both technical and non technical audiences.

  • U.S. citizenship required

Desired skills :

Familiarity with automated disaster recovery (DR) solutions, including but not limited to:

Amazon Web Services (AWS) Disaster Recovery Service (DRS): Experience with configuring and managing replication, fail over, and fail back processes for AWS workloads.

Microsoft Azure Site Recovery (ASR): Knowledge of setting up and managing site recovery between on premises environments, Azure, and other clouds.

Zerto: Hands on experience with continuous data protection (CDP) and near zero RPO replication across VMware, Hyper V, and cloud environments.

Veeam Backup & Replication: Experience with agent less backup, replication, and automated fail over testing for virtual, physical, and cloud workloads.

IBM Resiliency Services (formerly IBM Disaster Recovery as a Service): Familiarity with managed DR services for hybrid cloud environments, including integration with IBM Cloud and on premises infrastructure.

Experience with DR automation, including:

Scripting and integration with IaC tools (Terraform, CloudFormation) and CI/CD pipelines (Jenkins, GitLab) to automate DR workflows.

DR exercise planning and execution, including defining success criteria, metrics, and KPIs for recovery processes.

Strong analytical skills: ability to perform risk assessments, impact analysis, and cost benefit modeling for DR solutions.

Hybrid/multi-cloud deployments, with ability to manage DR across multiple cloud providers and on premises environments.

Advanced certifications (e.g., AWS Certified Solutions Architect - Professional, Azure Solutions Architect Expert, VMware VCAP DCV)., Generally has 5+ years of related experience and may have a post-secondary degree or training in a related discipline.

Benefits & conditions

Benefits offered to vary by the contract. Depending on your temporary assignment, benefits may include direct deposit, free career counseling services, 401(k), select paid holidays, short-term disability insurance, skills training, employee referral bonus, affordable medical coverage plan, and DailyPay (in some locations). For a full description of benefits available to you, be sure to talk with your recruiter.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:46 min

Navigating a career in cloud transformation consulting

Piet Van Dongen · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all