Platform Resiliency Lead (DR and BCRM)

Mars
Austin, TX, United States
about 2 months ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Software as a Service Programming Tools Disaster Recovery Secure Coding Okta System Availability

Job description

Ensure compliance with Mars BTR and DR goals and that platform reliability governance is embedded into dev and deploy processes and operating model, and continuously improve it. Ensure working are consistent, measurable, audit-ready, and friction-minimized across engineering and platform/tooling teams to support system availability required by business at minimal cost and maximum output. Help establish guardrails not gates by embedding governance into all workflows so teams can move at speed while maintaining strong control posture when complying with Mars and industry frameworks. Ensure development tools and pipelines align to Mars standards and secure coding/shift-left principles. Disaster Recovery & Preparedness

Ensure Application and Platform Disaster Recovery Plans (A/PDRPs) are created, maintained, and reviewed in line with platform criticality and policy expectations. Govern execution of disaster recovery testing, including scenario-based exercises and validation of achieved RTO/RPOs, with clear evidence capture and remediation tracking. Coordinate with Infrastructure, Network, Security, and Platform teams to confirm dependencies, recovery sequencing, and operational readiness during DR events. Incident & Recovery Leadership

Provide resilience leadership during major incidents and disaster recovery events, supporting command-and-control execution and decision making. Ensure lessons learned from incidents, DR tests, and near misses are translated into measurable improvements in platform design, tooling, and processes.

Continuous Improvement & Enablement

Track and report resilience maturity, adherence, and gaps across platforms using standardized metrics and compliance reporting. Partner with engineering and SRE teams to strengthen reliability practices such as observability, failover, backup, and controlled recovery mechanisms. Promote resilience awareness and enablement across platform and delivery teams through guidance, templates, and training. I5

Requirements

10+ years, MarTech and/or DR / BCRM experience. One of the two required. Key Tools: Salesforce Marketing/Loyalty/Service, Aprimo, TreasureData, Okta, Strong experience in platform reliability, disaster recovery, or resilience engineering within large-scale digital or cloud environments. Proven delivery of resilience or DR programs across multiple platforms or applications with differing criticality. Experience operating in regulated or audit-intensive environments.

Technical & Functional Skills

Deep understanding of DR concepts (RTO, RPO, backup strategies, recovery sequencing). Familiarity with cloud and SaaS recovery patterns and dependency management. Ability to translate technical resilience topics into clear risk, impact, and investment discussions for non technical stakeholders.

Leadership & Communication

Influences without direct authority; able to align diverse teams around resilience outcomes. Strong documentation and governance discipline suitable for audit and regulatory scrutiny. Comfortable operating in high pressure incident or recovery scenarios.

Success Measures

Platforms meet or exceed defined resilience and recovery objectives. Disaster recovery plans are complete, tested, and evidenced for critical platforms. Reduction in unmanaged resilience risks and audit findings. Improved recovery readiness and confidence across platform delivery teams.

Domain (Industry)*

CPG, Consumer Goods, Manufacturing

Minimum years of experience: 10+ Years

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Pitching technical resiliency initiatives to business decision makers

Mihaela-Roxana Ghidersa · LIVE

2:30 min

Evaluating and selecting an AI pair programming tool

Alexander Trusheim Alexander Trusheim +1 · WWC 2025

2:59 min

Applying secure coding practices and proactive system monitoring

Mihaela-Roxana Ghidersa · LIVE

2:33 min

Introduction to security advocacy and automation testing

Chris Heilmann +2 · LIVE

1:38 min

Empowering developers through continuous learning and modern programming tools

Katrin Lehmann Katrin Lehmann +1 · WWC 2025

2:47 min

Securing code provenance with digital identity signatures

Marcus Ross Marcus Ross · WWC Europe 2026

Videos

See all

Related articles

See all