BizOps SRE - Production Engineering

The Hire
St. Louis, MO, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Microsoft Azure Bash Shell Cloud Computing DevOps Monitoring of Systems Python (Programming Language) Release Management Reliability Engineering Prometheus Runbook
+8 more
Datadog Scripting Google Cloud Mttr Kubernetes Terraform Splunk Dynatrace

Job description

  • Own production operations, service stability, availability, and continuity for critical services.
  • Lead Major Incidents / Technical Response Team (TRT) activities and drive timely resolution and reduced MTTR.
  • Provide clear and timely communication to stakeholders, customers, and leadership during production incidents.
  • Drive operational readiness for application releases, infrastructure migrations, and peak business events.
  • Implement and improve monitoring, observability, alerting, and service health dashboards.
  • Lead Root Cause Analysis (RCA) and problem management activities.
  • Identify reliability risks and implement proactive solutions to improve system resilience.
  • Support capacity planning, performance improvement, and resilience initiatives.
  • Automate repetitive operational activities and reduce manual operational toil.
  • Develop and maintain reusable runbooks, SOPs, operational procedures, and readiness standards.
  • Partner with Engineering, Infrastructure SRE, Operations, Product, and other technical teams.
  • Convert production insights and incident learnings into engineering improvements and roadmap initiatives.
  • Participate in global production support and operational activities as required.

Requirements

We are seeking an experienced BizOps SRE / Production Engineer to support critical production services and drive reliability, automation, and operational excellence across a global environment.

The ideal candidate will have strong hands-on experience with SRE, DevOps, production operations, cloud infrastructure, Kubernetes, Terraform, observability, incident management, and automation. This role will focus on improving service stability, reducing MTTR and operational toil, strengthening production readiness, and partnering closely with Engineering, SRE, Operations, and Product teams., * Strong experience in SRE, DevOps, Production Engineering, or Site Reliability Engineering.

  • Hands-on experience supporting production environments.
  • Strong knowledge of Linux administration and troubleshooting.
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with Kubernetes.
  • Experience with Terraform and Infrastructure as Code.
  • Strong scripting/programming experience with one or more of:
  • Python
  • Go
  • Java
  • Bash
  • Experience with observability and monitoring tools such as:
  • Splunk
  • Dynatrace
  • Prometheus
  • Grafana
  • Datadog
  • Strong experience with Incident Management, Major Incident Management, RCA, Problem Management, and Change/Release Management.
  • Experience with CI/CD pipelines and automation.
  • Strong troubleshooting and production support skills.
  • Excellent communication and stakeholder management skills., * Experience supporting 24x7/global production environments.
  • Experience working with distributed/global teams.
  • Experience in financial services, banking, payments, or other high-availability environments.
  • Experience with production readiness reviews and release/migration readiness.
  • Experience developing operational runbooks, SOPs, and reliability standards.
  • Demonstrated experience reducing incidents, MTTR, or operational toil through automation and process improvements., The ideal candidate is a hands-on SRE/Production Engineer who can operate effectively in a high-availability production environment, lead critical incidents, troubleshoot complex infrastructure and application issues, automate repetitive processes, and collaborate effectively with engineering and business stakeholders.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all