Senior CloudOps Engineer

VANHACK TECHNOLOGIES INC.
United States
20 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Languages
English
Job source

Tech stack

Agile Methodology Artificial Intelligence Amazon Elastic Compute Cloud Computing Platforms Cloud Computing Cloud Engineering Cloud Storage Databases Continuous Integration DevOps Disaster Recovery Domain Name System (DNS)
+40 more
Identity and Access Management Subnetting Key Management Network Security MongoDB MySQL Networking Basics Routing Network Segmentation OAuth OpenID Performance Tuning Role-Based Access Control Redis Reliability Engineering Openid Connect Ansible Prometheus Runbook Security Assertion Markup Language (SAML) Single Sign-On Software Engineering Data Logging Google Cloud Load Balancing DevOps Tools - Open-source Istio System Availability Delivery Pipeline Grafana Software Troubleshooting Database Performance Generative AI Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Microservices

Job description

We are looking for a Senior CloudOps Engineer to design, build, and operate scalable, secure, and highly available cloud infrastructure on Google Cloud Platform (GCP) and Kubernetes. In this role, you will lead infrastructure automation, platform reliability, observability, and CI/CD initiatives while partnering closely with software engineering teams to deliver resilient production systems.

As a senior member of the Platform Engineering team, you will drive infrastructure best practices, improve developer productivity, and leverage AI-powered DevOps tooling to enhance operational efficiency. This role also participates in a rotating on-call schedule to support our production platform., Cloud Infrastructure

  • Design, deploy, and manage scalable, secure, and highly available infrastructure on Google Cloud Platform (GCP).
  • Architect and optimize GCP services including GKE, Compute Engine, VPC, IAM, Cloud Load Balancing, Cloud DNS, and Cloud Storage.
  • Continuously improve infrastructure reliability, scalability, performance, and cost efficiency.
  • Apply infrastructure security best practices, including least-privilege access, encryption, network segmentation, and compliance controls.
  • Partner with engineering teams to design cloud-native solutions and improve platform architecture.

Kubernetes & Platform Engineering

  • Design, deploy, and operate Kubernetes clusters running stateless and stateful workloads.
  • Manage Kubernetes networking, including Ingress Controllers, Service Mesh technologies, Virtual Services, and traffic management.
  • Build reusable deployment patterns using Helm, Terraform, and infrastructure automation tools.
  • Troubleshoot and optimize Kubernetes performance, networking, and cluster health., * Develop and maintain Infrastructure as Code (IaC) using Terraform and Ansible.
  • Build, maintain, and optimize CI/CD pipelines for automated application and infrastructure deployments.
  • Standardize deployment workflows and improve developer experience through automation.
  • Continuously improve release reliability, deployment speed, and operational efficiency., * Design and maintain monitoring, logging, tracing, and alerting solutions using Prometheus, Grafana, and OpenTelemetry.
  • Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and platform health metrics.
  • Implement proactive alerting, capacity planning, and performance optimization.
  • Improve platform reliability through root cause analysis, incident reviews, and operational excellence.

Database & Platform Operations

  • Deploy, administer, and optimize MongoDB, Redis, and MySQL environments.
  • Ensure backup, disaster recovery, replication, and high-availability strategies are implemented.
  • Monitor database performance and optimize resource utilization.

Security & Identity

  • Implement secure infrastructure and cloud security best practices.
  • Manage Identity and Access Management (IAM), Role-Based Access Control (RBAC), and secrets management.
  • Implement Single Sign-On (SSO) and identity federation using OAuth 2.0, SAML, and OpenID Connect (OIDC).
  • Participate in security reviews, audits, and compliance initiatives., * Integrate AI-assisted tools into infrastructure management, CI/CD pipelines, observability, and operational workflows.
  • Evaluate and implement AI agents, Model Context Protocol (MCP) servers, and emerging AI technologies to improve platform operations.
  • Apply AIOps techniques such as anomaly detection, predictive alerting, automated remediation, and intelligent incident response.
  • Leverage generative AI to improve documentation, runbooks, troubleshooting, and operational productivity.

Collaboration & Leadership

  • Participate in architecture discussions and technical planning.
  • Mentor engineers on infrastructure, Kubernetes, cloud architecture, and operational best practices.
  • Collaborate closely with Software Engineering, Security, and Product teams.
  • Participate in a rotating on-call schedule supporting production systems.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in CloudOps, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering.
  • Strong hands-on experience with Google Cloud Platform (GCP), including GKE, Compute Engine, VPC, IAM, Cloud Load Balancing, Cloud DNS, and cloud networking.
  • Advanced expertise with Kubernetes and containerized microservices.
  • Strong experience with Infrastructure as Code using Terraform and/or Ansible.
  • Experience designing and maintaining CI/CD pipelines.
  • Experience implementing monitoring, logging, and observability solutions using Prometheus, Grafana, and OpenTelemetry.
  • Experience administering production databases including MongoDB, Redis, and MySQL.
  • Solid understanding of networking fundamentals, including VPCs, subnets, load balancing, DNS, routing, firewalls, and network security.
  • Experience implementing Identity and Access Management (IAM), RBAC, OAuth 2.0, SAML, and OpenID Connect (OIDC).
  • Experience with AI-assisted DevOps tools, AIOps platforms, or infrastructure automation using generative AI.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Excellent written and verbal communication skills in English.
  • Ability to work effectively in Agile, cross-functional, and distributed teams

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on vanhack.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:49 min

Adopting OAuth best practices and removing outdated grants

Alexander Schwartz Alexander Schwartz · WWC Europe 2026

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

1:34 min

Analyzing vulnerabilities in standard OAuth 2.0 authorization flows

Alexander Schwartz Alexander Schwartz · WWC Europe 2026

Videos

See all

Related articles

See all