GCP Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
As a GCP Engineer, you will play a critical role in designing, implementing, and supporting enterprise disaster recovery and business continuity solutions within Google Cloud Platform environments. You will be responsible for ensuring that cloud architectures are resilient, highly available, and capable of meeting demanding business recovery requirements. This position places a strong emphasis on technical leadership, stakeholder collaboration, and maintaining operational readiness for recovery scenarios., * Designing and implementing disaster recovery and business continuity architectures across GCP environments to support enterprise objectives.
- Defining and documenting recovery strategies that align with business requirements, including Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO).
- Developing multi-region and multi-zone solutions to mitigate failure domains and enhance system resiliency.
- Creating and maintaining disaster recovery runbooks, operational procedures, and thorough documentation.
- Leading disaster recovery drills, failover testing, and recovery validation activities to ensure readiness.
- Implementing backup, replication, and recovery processes leveraging native GCP services and industry best practices.
- Designing resilient networking architectures, including automated routing and failover capabilities.
- Working closely with infrastructure, application, security, and operations teams to ensure coordinated recovery readiness across platforms.
- Establishing proactive monitoring, alerting, and operational visibility for disaster recovery environments.
- Partnering with business stakeholders to evaluate risk, assess downtime tolerances, and guide investment in recovery strategies.
- Providing technical leadership and guidance on cloud resiliency, availability, and disaster recovery best practices.
- Evaluating and optimizing DR solutions to ensure recovery capabilities while balancing infrastructure and operational costs., * Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Cloud SQL, Cloud Storage, Backup and DR Service
- Cloud Monitoring, Cloud Logging, Cloud DNS, Cloud Load Balancing
- Cloud Interconnect, VPN
- Terraform / OpenTofu, GitLab
- IAM, SRE practices, Disaster Recovery/Business Continuity frameworks
Requirements
- Strong knowledge of cloud architecture, particularly Google Cloud Platform fundamentals-regions, zones, multi-region design, and failure domain mapping.
- Deep experience with GCP redundancy and recovery services, including:
- Persistent Disk snapshots and policies
- Cloud Storage multi-region and dual-region buckets with versioning
- Cloud SQL cross-region replication
- Google Kubernetes Engine (GKE) multi-cluster and multi-region deployments
- GCP Backup and DR Service
- Expertise designing resilient network architectures and failover strategies, including:
- Global HTTP(S) Load Balancing
- Cloud DNS with health checks
- Cross-region Virtual Private Cloud (VPC) setups
- Cloud Interconnect and VPN for hybrid environments
- Proven Infrastructure as Code experience, especially with Terraform (OpenTofu experience is also suitable) and GitLab.
- Site Reliability Engineering (SRE) and cloud operations experience, including:
- Defining and achieving RTO/RPO objectives
- Conducting failover and disaster recovery drills
- Setting up monitoring and alerting with Cloud Monitoring and Logging
- Managing IAM and security controls across multiple regions
- Demonstrated ability to balance resiliency requirements with infrastructure and operational costs.
- Strong experience leading disaster recovery planning and facilitating business continuity engagements with stakeholders.
- Excellent documentation skills for clear and actionable recovery procedures and operational runbooks.
- Exceptional communication and interpersonal skills, with an ability to translate highly technical recovery strategies into understandable business recommendations, especially for non-technical stakeholders.
- Google Professional Cloud Architect certification is required.
Nice to Have:
- Technical foundation and experience with Microsoft Azure cloud services.
- Experience designing multi-cloud resiliency and disaster recovery solutions.
- Additional certifications related to cloud, security, or SRE disciplines., * Bachelor’s degree in Computer Science, Information Technology, or a related technical field (or equivalent experience).
- Google Professional Cloud Architect certification (required).
- Advanced knowledge of GCP disaster recovery architecture and multi-region deployments.
- 5+ years of hands-on experience with cloud technologies, with significant emphasis on Google Cloud Platform.
- Strong problem-solving, analytical, and troubleshooting skills.
- Ability to work effectively in a remote, distributed team environment.
About the company
ECCO Select is a talent acquisition and consulting company specializing in people, process and technology solutions. We provide the talent behind the technology enabling our clients to achieve their goals. For more information about ECCO Select, visit us at www.eccoselect.com.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again