GitLab Infrastructure Operations Engineer

Nexturn Inc.
United States
24 days ago
Apply on nexturn.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Linux DevOps Disaster Recovery Monitoring of Systems OpenShift Red Hat Enterprise Linux Ansible Prometheus Datadog Grafana Infrastructure as Code (IaC) Gitlab
+5 more
Containerization Kubernetes Information Technology Terraform Docker

Job description

  • Perform GitLab administration, including installation, configuration, upgrades, patching, and ongoing maintenance.
  • Manage GitLab Geo replication, backup, restore, and disaster recovery processes.
  • Design and maintain infrastructure using Infrastructure as Code (IaC) principles.
  • Develop and maintain automation using Ansible and Terraform.
  • Administer Linux/RHEL environments supporting GitLab infrastructure.
  • Manage containerized environments using Docker and Podman.
  • Support and troubleshoot Kubernetes and OpenShift environments.
  • Build, maintain, and troubleshoot GitLab Shared Runner infrastructure.
  • Configure and maintain monitoring and observability using Prometheus, Grafana, and Datadog.
  • Monitor system health, capacity, availability, and performance of GitLab infrastructure.
  • Participate in on-call rotations and provide timely response to production incidents.
  • Perform incident investigation, troubleshooting, and Root Cause Analysis (RCA).
  • Implement preventive measures to minimize recurring incidents and improve platform reliability.
  • Collaborate with DevOps, Platform Engineering, Security, and Development teams.
  • Automate operational tasks and continuously improve infrastructure reliability and efficiency.
  • Maintain technical documentation, operational procedures, and disaster recovery runbooks.

Requirements

  • Strong hands-on experience in GitLab administration, including upgrades, Geo replication and backup/restore.
  • Strong experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Linux/RHEL, Docker/Podman and Kubernetes/OpenShift environments.
  • Experience with Prometheus, Grafana and Datadog for monitoring, observability and alerting.
  • Experience managing shared GitLab Runner infrastructure and CI/CD execution environments.
  • Strong experience in on-call support, incident response, troubleshooting and Root Cause Analysis (RCA). Qualifications: Bachelor’s degree or Master’s degree in computer science, Information Technology, Engineering, or equivalent.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nexturn.com
Prepare application

Good distractions

Loading talks and stories from around this role…