> Markdown version of [/jobs/ext/573929-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/573929-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Fmr LLC - **Location:** Durham, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Adobe InDesign, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Software Applications, Systems Engineering, Automation of Tests, Baselining, Bash Shell, Big Data, Cloud Computing, Cloud Engineering, Computer Programming, Data Visualization, Query Languages, DevOps, Disaster Recovery, Distributed Systems, Fault Tolerance, Identity and Access Management, Apache JMeter, Python (Programming Language), Load Testing, Node.Js, Systems Development Life Cycle, Reliability Engineering, Prometheus, Software Engineering, Web Applications, Datadog, Data Logging, Scripting, Performance Testing, System Availability, Delivery Pipeline, Grafana, Scalability Testing, Amazon Relational Database Service, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Kibana, Terraform, Splunk, Golang - **Published:** June 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=88e56d3a3b1f5f32 ## About the Role Do you have experience in Validation design?, Do you have a Master's degree?, The team comes from diverse technical backgrounds, and the responsibilities provide the opportunity for a variety of challenges. Ideal candidates will have a background in either software engineering or systems engineering with a desire to learn the other or previous experience as an SRE. We are looking for a Systems Thinking, SRE Engineer who has helped teams scale through production insights, operational automation, developer guidance, real-time metrics, automation., * Bachelor's degree or higher in a technology related field (e.g. Engineering, Computer Science, etc.) required, master's degree is a plus. * Minimum 5 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at a scale. * 3 plus years of experience in Cloud development (AWS) and migration skills; Experience with building and operating highly resilient platforms in AWS cloud environments. * 3-5 years of experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation * Ensure platforms meet high availability, scalability, fault tolerance, and disaster recovery requirements. * Hands on experience with one or more observability tools (Datadog, Splunk, Kibana, Prometheus, Grafana, ELK/OpenSearch, Open Telemetry). * Hands on experience in designing, developing, and executing performance tests using K6/JMeter and other performance testing tools to ensure comprehensive performance testing. * Define Performance Test Strategy Document: set approach, metrics, benchmarks, baseline, user response requirements environments, technical environment and data conditions, and toolsets to use in executing the performance testing. * Experience in performance testing types: Load testing, Stress testing, Scalability testing, Spike testing, Volume testing, Chaos testing, Endurance/Soak testing * Hands-on experience with container orchestration, preferably with Kubernetes * Experience identifying memory leakage, connection issues and throughput bottlenecks in various technologies such as web application(s), infrastructure, and Cloud. * Strong knowledge of CI/CD pipelines and DevOps practices. * Familiarity with chaos engineering and resilience testing tools (e.g., Chaos Monkey, Gremlin). * Experience working in high-availability, large-scale production environments. * Strong programming/scripting skills in one or more: + Python, Java, Go, or Bash * Expertise in automation frameworks and tools for performance validation. * Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef) * Solid understanding of Cloud Computing and DevOps concepts including CI/CD pipelines. * Experienced in Instrumentation with systems skills on building and operating, monitoring, logging, alerting services of distributed systems at scale. * Proven experience in maintaining scalability and resiliency of complex environments. * Proven experience in implementing advanced observability practices and techniques at scale. * Ability to triage, execute root cause analysis, and be decisive under pressure. * Experience managing and interpreting large datasets using query languages and visualization tools. * Proficient communication skills with an ability to reach both technical and non-technical audience. * Ability to work with a variety of individuals and groups, both in person and virtually, in a constructive and collaborative manner and build and maintain effective relationships. * Experience in design, implement, and maintain performance test frameworks, which will validate to a high degree of confidence, the production readiness of software applications and infrastructure for stability and performance. Solid understanding of AWS services and experience setting up test environments on AWS (S3, EC2, RDS, etc.). ## Description Our Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability with resilience by using automation and Infrastructure Code. We build reliability into our ecosystem by applying best practices in Resiliency Engineering, Automation, Observability, Performance testing and Chaos testing. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Add Location-based Searching to Site with ElasticSearch](https://www.wearedevelopers.com/videos/77-add-location-based-searching-to-site-with-elasticsearch) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)