> Markdown version of [/jobs/ext/2166026-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2166026-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** ISO New England - **Location:** United States - **Experience:** Expert - **Salary:** $134,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Bash Shell, Cloud Computing, Cyber Security, DevOps, Disaster Recovery, Distributed Systems, Monitoring of Systems, Information Technology Operations, Python (Programming Language), Paessler Router Traffic Grapher, Windows PowerShell, Reliability Engineering, Runbook, Systems Integration, Data Logging, Performance Testing, Data Ingestion, System Availability, Indexer, Infrastructure as Code (IaC), Information Technology, Terraform, Splunk, Dynatrace, Pagerduty, Service Stack - **Published:** August 21, 2026 - **Apply:** https://www.dice.com/job-detail/84e7c975-90f1-48c4-ac95-16b9900f26b4 ## About the Role This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role., * 5+ years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations * Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation. * Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices * Strong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments * Demonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environment * Ability to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms * Knowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms * Familiarity with SRE principles including SLOs, error budgets, observability, and toil reduction * Ability to analyze and troubleshoot complex technical systems * Preferred Qualifications * Experience in mission-critical, highly available, or regulated environments * Experience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activities * Knowledge of ITIL processes and/or SRE best practices * Experience with performance testing, capacity planning, resilience testing, or disaster recovery validation This employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.). ## Description The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions., * Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage) * Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows * Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead * Develop meaningful KPIs and dashboards for business and IT service health * Engineer and implement resilience patterns including HA, DR, and automated failover * Partner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage * Conduct performance testing, capacity modeling, forecasting, and right-sizing * Participate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements * Identify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions * Design, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency * Design, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability * Identify gaps in observability coverage and drive engineering solutions to close them * Collaborate with architecture and application teams to ensure production readiness * Leverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness * Reduce repeat incidents by engineering permanent fixes and driving continuous improvement ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs)