> Markdown version of [/jobs/ext/3078289-site-reliability-engineer-hybrid-w2-self-corp-local-to-florida](https://www.wearedevelopers.com/jobs/ext/3078289-site-reliability-engineer-hybrid-w2-self-corp-local-to-florida). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer[Hybrid]-( W2 ,Self-Corp)-[Local to Florida] - **Company:** Save Mart Supermarkets LLC - **Location:** Miami, FL, United States - **Experience:** Expert - **Salary:** $68,500.0 - $92,600.0 - **Contract:** Temporary to permanent - **Skills:** Artificial Intelligence, HP Systems Insight Manager, Reliability Engineering, Web Platforms, Large Language Models, Grafana, Software Troubleshooting, BIG-IP Access Policy Manager (APM), Splunk, Dynatrace - **Published:** September 25, 2026 - **Apply:** https://www.careerjet.com/jobad/usf6373e4975a6e6b81c5cfcab730c7f32 ## About the Role 8 - 12 years of hands-on Site Reliability Engineering experience supporting production environments. 5+ years of experience with observability and monitoring, including: Splunk Dynatrace APM tools Dashboards Alerting System/application monitoring Strong experience with production incident management, incident triage, outage troubleshooting, and root cause analysis. Ability to assess incident severity, scope, customer impact, and business impact. 2+ years of experience using AI/LLMs for operational automation, monitoring enhancements, reliability activities, or intelligent guardrails. Strong understanding of SRE principles and reliability engineering. Experience with high-availability and customer-facing production applications. Experience with chaos engineering, resiliency practices, and synthetic monitoring. Strong troubleshooting and problem-solving skills. Experience partnering with development and engineering teams to resolve complex production issues. Strong communication and stakeholder-management skills. Preferred Experience AI agents and LLM-powered operational automation Automated incident response and remediation Reliability guardrails Customer-journey synthetic monitoring Release health monitoring and automated rollback strategies Performance and availability engineering Hospitality, travel, e-commerce, or other high-volume customer-facing environments ## Description We are seeking an experienced Site Reliability Engineer (SRE) to support the reliability, availability, performance, and resiliency of high-volume, customer-facing digital platforms and applications. The ideal candidate will have strong hands-on experience in SRE, observability, production incident management, resiliency engineering, automation, and AI/LLM technologies. This role will act as a first responder for production issues while proactively identifying reliability risks and implementing intelligent automation and guardrails. Key Responsibilities Monitor health, availability, and performance of production applications using Splunk, Dynatrace, APM, dashboards, alerting, and observability tools. Build and enhance dashboards to provide visibility into application and platform health. Identify trends, recurring issues, performance degradation, and potential reliability risks. Proactively detect and address issues before they impact customers. Serve as a first responder for production application and platform incidents. Triage incidents and determine severity, scope, root cause, and business impact. Troubleshoot production outages and coordinate resolution with engineering teams. Collaborate with technical and business stakeholders during critical incidents. Drive follow-up actions and permanent remediation for recurring issues. Support release evaluations and recommend pause/rollback decisions when releases negatively affect stability. Apply SRE principles to improve reliability, availability, scalability, and operational performance. Support chaos engineering, resiliency testing, and high-availability initiatives. Identify system weaknesses and develop strategies to reduce future incidents. Partner with development and engineering teams on long-term reliability improvements. Develop automation and operational guardrails using AI, LLMs, and AI agents. Build and enhance synthetic monitoring to simulate customer journeys and validate application health. Use automated monitoring to proactively identify failures and reliability risks. Leverage AI/LLMs to automate repetitive operational and incident-management activities. Continuously identify opportunities to improve monitoring, automation, and incident response. Required Qualifications ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [It's Not Vibe Coding If You Know What You're Doing](https://www.wearedevelopers.com/videos/100119-it-s-not-vibe-coding-if-you-know-what-you-re-doing) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)