> Markdown version of [/jobs/ext/2723058-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2723058-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** MSD - **Location:** United States - **Experience:** Expert - **Salary:** $230,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Software as a Service, Cloud Computing, DevOps, Domain Name System (DNS), Python (Programming Language), Load Testing, Routing, Network Administration, Systems Development Life Cycle, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, TCP/IP, Datadog, Data Logging, Load Balancing, System Availability, Grafana, Firewalls (Computer Science), Kubernetes, Terraform, Splunk - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-forward-networks-9702951 ## About the Role * 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment * Proven experience building or significantly maturing an SRE function - not just operating within one someone else built * Strong fundamentals in networking - TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus * Hands-on experience with Kubernetes and container orchestration in production environments * Deep proficiency with observability tooling - Prometheus, Grafana, Datadog, Splunk, or similar * Strong scripting and automation skills in Python, Bash, or similar * Experience with cloud platforms - AWS, GCP, or Azure - including infrastructure as code (Terraform, Ansible, or equivalent) * Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements * Ability to communicate clearly with both engineering teams and non-technical stakeholders - you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them Nice to Have * Experience supporting enterprise or federal government customers with high availability requirements * Experience in a foundational or early SRE hire capacity at a growth stage company ## Description About the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward - defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand. If you thrive in environments where you're handed a problem rather than a playbook this role is for you. What You'll Own * Define and drive SRE practices from the ground up - SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use * Drive the reliability and operational excellence of the Forward SaaS platform * Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the team always knows what's happening before customers do * Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice * Partner with engineering teams to embed reliability thinking into the SDLC - capacity planning, load testing, chaos engineering, and production readiness reviews * Help define and build the SRE team as the company scales - this is a foundational hire with a path to leadership ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)