Senior DevOps Platform Engineer

Insight Global
Dallas, TX, United States
4 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Cloud Computing Cloud Engineering Databases Continuous Integration DevOps Distributed Systems Github Python (Programming Language) Node.Js OpenShift
+28 more
Performance Tuning Release Management Reliability Engineering Prometheus SQL Databases Datadog Data Logging Google Cloud Cloud Platform System System Availability Delivery Pipeline Grafana Reliability of Systems Backend Kotlin Event Driven Architecture Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Enterprise Integration Restful APIs Terraform New Relic (SaaS) Docker Service Stack Jenkins

Job description

The Senior DevOps Platform Engineer is responsible for designing, building, and maintaining the cloud infrastructure, deployment pipelines, and backend platform services that support a proprietary handheld device used across restaurant operations. This role focuses on platform reliability, scalability, automation, and operational excellence, ensuring applications and services remain highly available in high-volume, real-time restaurant environments. The ideal candidate brings strong experience with cloud-native architectures, CI/CD, infrastructure as code, observability, and production support while partnering closely with application engineering teams to accelerate delivery and improve system resilience., Design, build, and maintain cloud infrastructure and platform services supporting critical restaurant operations.

-Develop and optimize CI/CD pipelines, deployment automation, and release management processes.

-Manage and improve cloud environments, ensuring scalability, security, performance, and cost efficiency.

-Drive infrastructure as code (IaC) practices using Terraform and other automation tools.

-Own platform reliability, availability, and operational readiness across production environments.

-Lead root cause analysis and resolution of complex production incidents and performance issues.

-Partner with application engineering teams to support backend services, APIs, and distributed systems.

Implement observability best practices including monitoring, logging, tracing, and alerting.

-Participate in on-call rotations and provide L3 support for critical platform issues.

-Contribute to platform standards, security best practices, and engineering excellence initiatives.

Technology Stack

Cloud: Google Cloud Platform (GCP), including GKE, Cloud Run, Cloud SQL, Pub/Sub

Infrastructure as Code: Terraform

Containerization & Orchestration: Docker, Kubernetes

CI/CD: GitHub Actions, Jenkins, GitLab CI, or similar

Backend Services: Java, Kotlin, Node.js, Python, or Go

Observability: New Relic, Datadog, Grafana, Prometheus, OpenTelemetry

APIs & Integrations: REST APIs, event-driven architectures, service integrations

Data: SQL, transactional databases, troubleshooting and performance tuning

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

7+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure roles.

-Strong experience building and supporting cloud-native platforms in GCP.

-Expertise with Terraform, Kubernetes, containerized applications, and CI/CD automation.

-Experience supporting high-volume, customer-facing or operationally critical applications.

-Proven ability to diagnose and resolve complex production incidents in distributed systems.

-Strong understanding of reliability engineering, scalability, security, and performance optimization.

-Experience implementing observability, monitoring, and incident management practices.

-Ability to work closely with application engineering teams to improve developer experience and delivery velocity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Understanding Kotlin Multiplatform and its compiler targets

Petar Marijanović · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:18 min

Recommended resources and frameworks for learning Kotlin natively

Iris Hunkeler · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all