DevOps / Infrastructure Operations Engineer

Spectraforce
United States
3 days ago
Apply on leoforce.us
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services CentOS Cyber Security Continuous Integration DevOps Disaster Recovery Domain Name System (DNS) Monitoring of Systems Identity and Access Management Virtual Private Networks (VPN) Lightweight Directory Access Protocols (LDAP) Linux System Administration
+20 more
Networking Basics OpenShift Red Hat Enterprise Linux Ansible Prometheus Security Assertion Markup Language (SAML) User Provisioning Software Software Vulnerability Management Datadog Load Balancing System Availability Grafana Gitlab Containerization Gitlab-ci Kubernetes Deployment Automation Firewall Services Module Terraform Docker

Job description

Client is seeking a Senior Platform Engineer to join the operations team for gitlab.cee - Client’s self-managed GitLab instance. This is a C1 (Mission-Critical) service serving ~10,000+ engineers with high availability architecture across multiple AWS regions. The platform is mature and operational. Your primary focus will be maintaining reliability, performing upgrades, managing compliance, and improving automation. You will provide US timezone coverage alongside existing team members, ensuring round-the-clock operational resilience for this critical platform.

What You’ll Do

  • Operate and maintain production and pre-production GitLab environments
  • Perform GitLab version upgrades through the Stage-to-Production pipeline
  • Execute system patching, vulnerability remediation, and compliance tasks
  • Manage GitLab Shared Runner infrastructure
  • Manage GitLab Geo replication across primary and secondary sites
  • Conduct and maintain disaster recovery exercises and documentation
  • Automate secret rotation via Ansible Automation Platform (AAP)
  • Maintain and improve Infrastructure as Code (Ansible/Terraform + GitLab CI)
  • Handle SNOW tickets - access requests, pipeline issues, configuration changes
  • Monitor service health using Prometheus, Grafana, and Datadog
  • Participate in on-call rotation with peer engineers
  • Create and maintain runbooks, documentation, and post-incident reviews, Key Stakeholders
  • ALM/DEP - Platform ownership, priority alignment
  • InfoSec - SOC monitoring, incident response, vulnerability remediation
  • IT-IAM - User provisioning, SSO integration
  • Engineering teams - Thousands of users relying on platform availability
  • PCO/Ops - Infrastructure, networking, AWS account management

Platform State

  • GitLab 10k reference architecture with high availability
  • Geo replication across multiple AWS regions
  • Automated deployment via Ansible/Terraform + GitLab CI
  • Monitoring: Prometheus + Grafana + Datadog
  • C1 Mission-Critical service

This Role Is NOT

  • A build-from-scratch project - the infrastructure is mature and well-documented
  • A pure development role - this is infrastructure operations
  • A solo position - you join an existing team of engineers
  • A user support role - you manage the platform, not individual project workflows

Requirements

  • 4+ years of experience in Platform Engineering, DevOps, or production platform operations
  • Strong GitLab administration experience - installation, configuration, upgrades, Geo replication, backup/restore at scale
  • Linux systems administration (RHEL/CentOS)
  • Infrastructure as Code proficiency - Ansible, Terraform, and CI/CD pipelines
  • Monitoring and observability experience - Prometheus, Grafana, or equivalent
  • Containerization and orchestration - Docker/Podman and Kubernetes/OpenShift
  • Incident management experience - on-call, incident response, root cause analysis
  • Networking fundamentals - DNS, load balancing, VPN, firewall rules
  • Strong documentation skills, * GitLab Geo replication operations and troubleshooting
  • Enterprise compliance frameworks (SOC2, ISO 27001, or equivalent; Red Hat ESS/PIA/SIA a plus)
  • IAM integration (SSO/SAML, LDAP)
  • High-availability architecture design and operations
  • Red Hat or IBM enterprise environment experience
  • Datadog monitoring platform
  • AWS infrastructure operations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on leoforce.us
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

9:06 min

Questions on career paths and continuous delivery orchestration platforms

Zan Markan Zan Markan · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all