Site Reliability Engineer - Production Support

Bayside Solutions
Cupertino, CA, United States
18 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JIRA Microsoft Azure Command-Line Interface Continuous Integration Software Debugging Linux Domain Name System (DNS) Intrusion Detection and Prevention Python (Programming Language) Network Troubleshooting Ansible
+15 more
Shell Script Datadog SSL Certificate Management Transport Layer Security Load Balancing Git Containerization Kubernetes Low Latency Puppet Terraform Splunk Docker Jenkins Servicenow

Job description

  • Production support / application support / NOC / service reliability / incident management. Screen out: security analysts whose Splunk work is log-based threat detection.
  • Watch service health across production; catch, triage, and drive issues to resolution.
  • Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team.
  • Monitor and investigate using Splunk (and Datadog, where in play).
  • Support the application stack in cloud and datacenter environments.
  • Non-prod and prod deployments using our in-house CD tool; monitor deployments and roll back / debug when they fail.
  • Keep Git repos current; debug deployment scripts.

Requirements and Qualifications:

  • Hands-on production support/application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
  • Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
  • Splunk is required. They will be asked what commands/searches they run.
  • SOC/production monitoring
  • Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
  • Linux/Unix fluency.
  • Containerized workloads: Docker + Kubernetes (managing, not just using)
  • Cloud: AWS primary; Azure/AKS acceptable and precedented. Datadog - Naveen called it out as a plus point. New to this team’s stack vs. prior reqs in Lighthouse is likely a growing surface.
  • Payments/banking production environment (card networks, Visa/Mastercard, wallets, issuer/acquirer processing).
  • Java application troubleshooting.

Requirements

  • CI/CD tooling - Jenkins, specifically, can debug the scripts behind a deployment, not just trigger one.
  • Python and/or shell scripting for automation and runbook work.
  • Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
  • Config management (Ansible/Chef/Puppet), autoscaling design in Kubernetes.
  • Terraform / infrastructure-as-code exposure.
  • ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:58 min

Teaching software skills inside real production environments

Alex Goldman Alex Goldman · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all