Senior DevOps Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+28 more
Job description
The Senior DevOps Platform Engineer is responsible for designing, building, and maintaining the cloud infrastructure, deployment pipelines, and backend platform services that support a proprietary handheld device used across restaurant operations. This role focuses on platform reliability, scalability, automation, and operational excellence, ensuring applications and services remain highly available in high-volume, real-time restaurant environments. The ideal candidate brings strong experience with cloud-native architectures, CI/CD, infrastructure as code, observability, and production support while partnering closely with application engineering teams to accelerate delivery and improve system resilience., Design, build, and maintain cloud infrastructure and platform services supporting critical restaurant operations.
-Develop and optimize CI/CD pipelines, deployment automation, and release management processes.
-Manage and improve cloud environments, ensuring scalability, security, performance, and cost efficiency.
-Drive infrastructure as code (IaC) practices using Terraform and other automation tools.
-Own platform reliability, availability, and operational readiness across production environments.
-Lead root cause analysis and resolution of complex production incidents and performance issues.
-Partner with application engineering teams to support backend services, APIs, and distributed systems.
Implement observability best practices including monitoring, logging, tracing, and alerting.
-Participate in on-call rotations and provide L3 support for critical platform issues.
-Contribute to platform standards, security best practices, and engineering excellence initiatives.
Technology Stack
Cloud: Google Cloud Platform (GCP), including GKE, Cloud Run, Cloud SQL, Pub/Sub
Infrastructure as Code: Terraform
Containerization & Orchestration: Docker, Kubernetes
CI/CD: GitHub Actions, Jenkins, GitLab CI, or similar
Backend Services: Java, Kotlin, Node.js, Python, or Go
Observability: New Relic, Datadog, Grafana, Prometheus, OpenTelemetry
APIs & Integrations: REST APIs, event-driven architectures, service integrations
Data: SQL, transactional databases, troubleshooting and performance tuning
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.
Requirements
7+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure roles.
-Strong experience building and supporting cloud-native platforms in GCP.
-Expertise with Terraform, Kubernetes, containerized applications, and CI/CD automation.
-Experience supporting high-volume, customer-facing or operationally critical applications.
-Proven ability to diagnose and resolve complex production incidents in distributed systems.
-Strong understanding of reliability engineering, scalability, security, and performance optimization.
-Experience implementing observability, monitoring, and incident management practices.
-Ability to work closely with application engineering teams to improve developer experience and delivery velocity.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers