> Markdown version of [/jobs/ext/2120157-site-reliability-engineer-production-support](https://www.wearedevelopers.com/jobs/ext/2120157-site-reliability-engineer-production-support). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Production Support - **Company:** Bayside Solutions - **Location:** Cupertino, CA, United States (Remote available) - **Salary:** $124,800.0 - $145,600.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JIRA, Microsoft Azure, Command-Line Interface, Continuous Integration, Software Debugging, Linux, Domain Name System (DNS), Intrusion Detection and Prevention, Python (Programming Language), Network Troubleshooting, Ansible, Shell Script, Datadog, SSL Certificate Management, Transport Layer Security, Load Balancing, Git, Containerization, Kubernetes, Low Latency, Puppet, Terraform, Splunk, Docker, Jenkins, Servicenow - **Published:** August 19, 2026 - **Apply:** https://www.dice.com/job-detail/cf4eccbd-b5f2-4aa1-ab4d-7052f64abfa5 ## About the Role * CI/CD tooling - Jenkins, specifically, can debug the scripts behind a deployment, not just trigger one. * Python and/or shell scripting for automation and runbook work. * Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting. * Config management (Ansible/Chef/Puppet), autoscaling design in Kubernetes. * Terraform / infrastructure-as-code exposure. * ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity. ## Description * Production support / application support / NOC / service reliability / incident management. Screen out: security analysts whose Splunk work is log-based threat detection. * Watch service health across production; catch, triage, and drive issues to resolution. * Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team. * Monitor and investigate using Splunk (and Datadog, where in play). * Support the application stack in cloud and datacenter environments. * Non-prod and prod deployments using our in-house CD tool; monitor deployments and roll back / debug when they fail. * Keep Git repos current; debug deployment scripts. Requirements and Qualifications: * Hands-on production support/application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident. * Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team. * Splunk is required. They will be asked what commands/searches they run. * SOC/production monitoring * Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here. * Linux/Unix fluency. * Containerized workloads: Docker + Kubernetes (managing, not just using) * Cloud: AWS primary; Azure/AKS acceptable and precedented. Datadog - Naveen called it out as a plus point. New to this team's stack vs. prior reqs in Lighthouse is likely a growing surface. * Payments/banking production environment (card networks, Visa/Mastercard, wallets, issuer/acquirer processing). * Java application troubleshooting. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)