Site Reliability Engineer[Hybrid]-( W2 ,Self-Corp)-[Local to Florida]
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We are seeking an experienced Site Reliability Engineer (SRE) to support the reliability, availability, performance, and resiliency of high-volume, customer-facing digital platforms and applications. The ideal candidate will have strong hands-on experience in SRE, observability, production incident management, resiliency engineering, automation, and AI/LLM technologies. This role will act as a first responder for production issues while proactively identifying reliability risks and implementing intelligent automation and guardrails. Key Responsibilities Monitor health, availability, and performance of production applications using Splunk, Dynatrace, APM, dashboards, alerting, and observability tools. Build and enhance dashboards to provide visibility into application and platform health. Identify trends, recurring issues, performance degradation, and potential reliability risks. Proactively detect and address issues before they impact customers. Serve as a first responder for production application and platform incidents. Triage incidents and determine severity, scope, root cause, and business impact. Troubleshoot production outages and coordinate resolution with engineering teams. Collaborate with technical and business stakeholders during critical incidents. Drive follow-up actions and permanent remediation for recurring issues. Support release evaluations and recommend pause/rollback decisions when releases negatively affect stability. Apply SRE principles to improve reliability, availability, scalability, and operational performance. Support chaos engineering, resiliency testing, and high-availability initiatives. Identify system weaknesses and develop strategies to reduce future incidents. Partner with development and engineering teams on long-term reliability improvements. Develop automation and operational guardrails using AI, LLMs, and AI agents. Build and enhance synthetic monitoring to simulate customer journeys and validate application health. Use automated monitoring to proactively identify failures and reliability risks. Leverage AI/LLMs to automate repetitive operational and incident-management activities. Continuously identify opportunities to improve monitoring, automation, and incident response. Required Qualifications
Requirements
8 - 12 years of hands-on Site Reliability Engineering experience supporting production environments. 5+ years of experience with observability and monitoring, including: Splunk Dynatrace APM tools Dashboards Alerting System/application monitoring Strong experience with production incident management, incident triage, outage troubleshooting, and root cause analysis. Ability to assess incident severity, scope, customer impact, and business impact. 2+ years of experience using AI/LLMs for operational automation, monitoring enhancements, reliability activities, or intelligent guardrails. Strong understanding of SRE principles and reliability engineering. Experience with high-availability and customer-facing production applications. Experience with chaos engineering, resiliency practices, and synthetic monitoring. Strong troubleshooting and problem-solving skills. Experience partnering with development and engineering teams to resolve complex production issues. Strong communication and stakeholder-management skills. Preferred Experience AI agents and LLM-powered operational automation Automated incident response and remediation Reliability guardrails Customer-journey synthetic monitoring Release health monitoring and automated rollback strategies Performance and availability engineering Hospitality, travel, e-commerce, or other high-volume customer-facing environments
About the company
Jones Lang LaSalle
-
Miami, FL JLL empowers you to shape a brighter way. Our people at JLL are shaping the future of real estate for a better world by combining world class services, advisory and technology fo…, IBA Worldwide
- Miami, FL
- $68,500-92,600 per year Life at IBA At IBA, we’re not just building technology - we’re shaping the future of Cancer care. Headquartered in Belgium and powered by over 2,200 passionate professionals worl…
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Fully Remote Software Engineer Jobs
Dev Digest 120 - Apple and peers
Where To Find Software Engineering Jobs