DevOps / Infrastructure Operations Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+20 more
Job description
Client is seeking a Senior Platform Engineer to join the operations team for gitlab.cee - Client’s self-managed GitLab instance. This is a C1 (Mission-Critical) service serving ~10,000+ engineers with high availability architecture across multiple AWS regions. The platform is mature and operational. Your primary focus will be maintaining reliability, performing upgrades, managing compliance, and improving automation. You will provide US timezone coverage alongside existing team members, ensuring round-the-clock operational resilience for this critical platform.
What You’ll Do
- Operate and maintain production and pre-production GitLab environments
- Perform GitLab version upgrades through the Stage-to-Production pipeline
- Execute system patching, vulnerability remediation, and compliance tasks
- Manage GitLab Shared Runner infrastructure
- Manage GitLab Geo replication across primary and secondary sites
- Conduct and maintain disaster recovery exercises and documentation
- Automate secret rotation via Ansible Automation Platform (AAP)
- Maintain and improve Infrastructure as Code (Ansible/Terraform + GitLab CI)
- Handle SNOW tickets - access requests, pipeline issues, configuration changes
- Monitor service health using Prometheus, Grafana, and Datadog
- Participate in on-call rotation with peer engineers
- Create and maintain runbooks, documentation, and post-incident reviews, Key Stakeholders
- ALM/DEP - Platform ownership, priority alignment
- InfoSec - SOC monitoring, incident response, vulnerability remediation
- IT-IAM - User provisioning, SSO integration
- Engineering teams - Thousands of users relying on platform availability
- PCO/Ops - Infrastructure, networking, AWS account management
Platform State
- GitLab 10k reference architecture with high availability
- Geo replication across multiple AWS regions
- Automated deployment via Ansible/Terraform + GitLab CI
- Monitoring: Prometheus + Grafana + Datadog
- C1 Mission-Critical service
This Role Is NOT
- A build-from-scratch project - the infrastructure is mature and well-documented
- A pure development role - this is infrastructure operations
- A solo position - you join an existing team of engineers
- A user support role - you manage the platform, not individual project workflows
Requirements
- 4+ years of experience in Platform Engineering, DevOps, or production platform operations
- Strong GitLab administration experience - installation, configuration, upgrades, Geo replication, backup/restore at scale
- Linux systems administration (RHEL/CentOS)
- Infrastructure as Code proficiency - Ansible, Terraform, and CI/CD pipelines
- Monitoring and observability experience - Prometheus, Grafana, or equivalent
- Containerization and orchestration - Docker/Podman and Kubernetes/OpenShift
- Incident management experience - on-call, incident response, root cause analysis
- Networking fundamentals - DNS, load balancing, VPN, firewall rules
- Strong documentation skills, * GitLab Geo replication operations and troubleshooting
- Enterprise compliance frameworks (SOC2, ISO 27001, or equivalent; Red Hat ESS/PIA/SIA a plus)
- IAM integration (SSO/SAML, LDAP)
- High-availability architecture design and operations
- Red Hat or IBM enterprise environment experience
- Datadog monitoring platform
- AWS infrastructure operations
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DevOps Engineer Salary [2023]
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs