Site Reliability Engineer - Production Support
Bayside Solutions
Cupertino, CA, United States
18 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
JIRA
Microsoft Azure
Command-Line Interface
Continuous Integration
Software Debugging
Linux
Domain Name System (DNS)
Intrusion Detection and Prevention
Python (Programming Language)
Network Troubleshooting
Ansible
+15 more
Shell Script
Datadog
SSL Certificate Management
Transport Layer Security
Load Balancing
Git
Containerization
Kubernetes
Low Latency
Puppet
Terraform
Splunk
Docker
Jenkins
Servicenow
Job description
- Production support / application support / NOC / service reliability / incident management. Screen out: security analysts whose Splunk work is log-based threat detection.
- Watch service health across production; catch, triage, and drive issues to resolution.
- Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team.
- Monitor and investigate using Splunk (and Datadog, where in play).
- Support the application stack in cloud and datacenter environments.
- Non-prod and prod deployments using our in-house CD tool; monitor deployments and roll back / debug when they fail.
- Keep Git repos current; debug deployment scripts.
Requirements and Qualifications:
- Hands-on production support/application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
- Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
- Splunk is required. They will be asked what commands/searches they run.
- SOC/production monitoring
- Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
- Linux/Unix fluency.
- Containerized workloads: Docker + Kubernetes (managing, not just using)
- Cloud: AWS primary; Azure/AKS acceptable and precedented. Datadog - Naveen called it out as a plus point. New to this team’s stack vs. prior reqs in Lighthouse is likely a growing surface.
- Payments/banking production environment (card networks, Visa/Mastercard, wallets, issuer/acquirer processing).
- Java application troubleshooting.
Requirements
- CI/CD tooling - Jenkins, specifically, can debug the scripts behind a deployment, not just trigger one.
- Python and/or shell scripting for automation and runbook work.
- Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
- Config management (Ansible/Chef/Puppet), autoscaling design in Kubernetes.
- Terraform / infrastructure-as-code exposure.
- ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
24 days ago
LM
Luis Minvielle
The 8 Best Code Testing Tools
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
AB
Andre Braun, GitLab
Now is the time for industrialized software development
about 1 year ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago