> Markdown version of [/jobs/ext/2719695-software-engineer-site-reliability](https://www.wearedevelopers.com/jobs/ext/2719695-software-engineer-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Site Reliability - **Company:** BLOOMERANG, LLC - **Location:** Indianapolis, IN, United States (Remote available) - **Experience:** Expert - **Salary:** $115,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** .NET Framework, PHP (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Build Automation, Static Program Analysis, Databases, Relational Databases, Software Debugging, DevOps, PostgreSQL, Node.Js, Reliability Engineering, Site Reliability Engineering Practices, Standard Sql, Software Engineering, Scripting, Grafana, Break Fix, Reliability of Systems, Codebase, Cloudwatch, Kibana, New Relic (SaaS) - **Published:** September 4, 2026 - **Apply:** https://www.builtincolorado.com/job/sr-software-engineer-site-reliability/11013492?handler=ApplyRedirect ## About the Role SRE Experience & Transformation * Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices. * Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction. Observability & Incident Management * Experience building monitoring, dashboards, alerts, and telemetry using tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar. * Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through. Technical Depth * Strong programming and scripting skills to navigate and troubleshoot application code and build automation and operational tooling. * Strong SQL and relational database skills for production troubleshooting and safe data correction; PostgreSQL experience preferred. * Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying potential reliability issues. * Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases. Comfort navigating application stacks across technologies such as PHP, .NET, and Node.js; deep expertise in each is not required. AI, Ownership & Collaboration * Demonstrated experience using and embracing AI-assisted tools in day-to-day engineering workflows, including troubleshooting, code analysis, scripting, and automation. * A collaborative self-starter who tackles difficult problems, adapts to changing priorities, and challenges the status quo. * Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams. ## Description We are evolving our Tier 3 Support Engineering team into a Site Reliability Engineering (SRE) organization focused on improving product reliability, observability, and operational efficiency. We're looking for an experienced SRE who brings strong software engineering fundamentals and is excited to help shape this transformation and mature our SRE practices. This is a highly collaborative, hands-on role working across application code, telemetry, databases, APIs, and infrastructure to diagnose complex production issues and improve reliability. Production support and product defects remain part of today's work as you help reduce reactive effort through observability, SLOs, automation, permanent fixes, and proactive reliability engineering. What Success Looks Like Success isn't measured solely by issues resolved, but by issues that no longer require manual intervention. You'll help detect problems earlier, reduce recurring issues and toil, strengthen incident response, and create more capacity for proactive reliability engineering. What You Will Do * Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work. * Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions to recurring problems and defects. * Bring proven SRE practices to the team and foster proactive reliability, continuous improvement, automation, and shared ownership. * Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews, turning lessons learned into reliability improvements. * Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts. * Define and mature SLIs and SLOs that measure system reliability and customer experience. * Develop synthetic monitoring for critical customer journeys to detect failures before they impact customers. * Identify sources of recurring operational toil and drive automation, tooling, process improvements, or permanent fixes that reduce manual effort. * Use AI-assisted tools and source code repositories to accelerate triage, troubleshooting, code analysis, automation, and technical investigation. * Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)