> Markdown version of [/jobs/ext/1321586-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1321586-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** The Best - **Location:** Wilmington, DE, United States - **Salary:** $120,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, JIRA, Bash Shell, Cloud Computing, Continuous Delivery, Continuous Integration, Linux, File Transfer, Python (Programming Language), Windows PowerShell, Reliability Engineering, Datadog, Scripting, Reliability of Systems, Information Technology, Servicenow - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a310cd072075de5d ## About the Role * Hands-on familiarity with production support, monitoring, alerting, and incident response practices. * Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts. * Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry. * Exposure to AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools. * Clear communication skills during incidents, service requests, and post-incident follow-through. * Strong experience leading production incident recovery and cross-system reliability investigations. * Ability to mentor engineers and influence technical decisions without direct authority., * Datadog, AWS, ITIL, Linux, or automation certification. * Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments. * Exposure to AIOps, anomaly detection, operational automation, or reliability engineering. * Familiarity with financial services controls, secure file transfer, or regulated operations. ## Description As Site Reliability Engineer, you will serve as a reliability subject matter expert who leads major incident recovery, drives observability and reliability improvements, mentors associate engineers, reduces operational toil, influences technical decisions, and improves resiliency standards. This role requires depth across production systems, telemetry, batch operations, automation, and incident response. You will be expected to guide technical direction for reliability improvements and help teams prevent recurring failures. Employees joining Best Egg's Information Technology organization can expect a culture centered on Continuous Delivery, Total Quality Management, Knowledge Sharing, Personal and Career Advancement, Empowerment, Innovation, and Collective Ownership., * Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation. * Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services. * Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk. * Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment. * Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns. * Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards. * Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure., * Serves as a reliability SME across observability, incident response, batch operations, and operational platforms. * Leads major incident recovery with calm command of telemetry, dependencies, impact, and restoration options. * Reduces operational toil through automation, standards, better alerting, and durable remediation. * Mentors engineers and improves the quality of technical support practices across the team. * Influences design and readiness decisions that improve resiliency and operational resilience. Success Metrics in First 90 Days * Lead a major incident or complex reliability investigation with clear recovery and follow-through. * Deliver an observability or automation improvement that measurably reduces alert noise, toil, or repeat issues. * Mentor associate engineers through troubleshooting reviews, runbook improvements, or incident debriefs. * Identify and influence remediation of a meaningful resiliency or supportability gap. * Improve standards or patterns for telemetry, escalation, batch support, or operational validation. ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)