> Markdown version of [/jobs/ext/2712384-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/2712384-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** nOCD, Inc. - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $160,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Cloud Computing, Cloud Engineering, Databases, Computer Engineering, DevOps, Programming Tools, Disaster Recovery, Distributed Systems, Github, Identity and Access Management, Python (Programming Language), Key Management, Octopus Deploy, Reliability Engineering, Prometheus, Software Engineering, TypeScript, Datadog, Data Logging, Grafana, Software Security, Reliability of Systems, Event Driven Architecture, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Production Code, Cloudwatch, Terraform, Splunk, Devsecops, Docker, Jenkins, Programming Languages, Microservices - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-site-reliability-engineer-nocd-company-8993170 ## About the Role * 7+ years of professional software engineering experience, with significant experience in SRE, platform engineering, DevOps, or cloud infrastructure. * Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field, or equivalent professional experience. * Strong software engineering fundamentals and experience writing production-quality code. * Strong hands-on experience with AWS and cloud architecture. * Strong experience with Terraform or other infrastructure-as-code tools. * Strong experience with Docker and Kubernetes. * Proficiency in Python, TypeScript, or a similar programming language. * Experience designing and operating CI/CD pipelines and deployment automation. * Strong understanding of distributed systems, APIs, networking, databases, and cloud architecture. * Experience with monitoring, logging, observability, and production troubleshooting. * Experience participating in or leading incident response and root-cause analysis. * Strong understanding of software reliability, scalability, availability, and performance. * Demonstrated ability to own technical initiatives and work effectively across engineering teams., * Experience working in healthcare, fintech, or another regulated environment. * Experience supporting HIPAA, SOC 2, HITRUST, or similar compliance frameworks. * Experience with Datadog, CloudWatch, Prometheus, Grafana, Splunk, or similar observability platforms. * Experience with GitHub Actions, Jenkins, ArgoCD, or GitOps. * Experience designing highly available or multi-region AWS architectures. * Experience with microservices and event-driven architectures. * Experience with infrastructure security, DevSecOps, IAM, secrets management, and encryption. * Experience establishing SLIs, SLOs, SLAs, and error budgets. * Experience building internal developer platforms or developer tooling. * Experience with disaster recovery, capacity planning, and performance engineering. * Experience working in a high-growth startup environment. ## Description * Develop and maintain APIs, microservices, automation, and internal engineering tools. * Contribute to architecture decisions, technical design reviews, code reviews, and engineering standards. * Apply software engineering principles to infrastructure, automation, and reliability challenges. Cloud & Platform Engineering * Design, build, and operate AWS infrastructure supporting production applications. * Manage infrastructure as code using Terraform. * Build and maintain containerized environments using Docker and Kubernetes. * Design and improve CI/CD pipelines, deployment automation, and release processes. * Build internal tooling and platform capabilities that improve developer productivity. * Help establish scalable infrastructure patterns that can support continued company growth. Reliability & Observability * Own and improve the reliability, availability, performance, and scalability of production systems. * Develop monitoring, alerting, logging, and observability strategies across our infrastructure and applications. * Lead incident response, troubleshooting, and root-cause analysis for production issues. * Establish and improve operational practices around incident management, postmortems, and preventative remediation. * Identify reliability risks and proactively improve system resilience, capacity, and performance. * Help define and monitor appropriate SLIs, SLOs, and operational metrics. Security & Compliance * Implement cloud and application security best practices across infrastructure and production systems. * Partner with Security and Engineering to support HIPAA, SOC 2, and other compliance requirements. * Implement appropriate controls around access management, secrets, encryption, logging, and infrastructure security. * Help identify and remediate infrastructure and application security risks. Technical Leadership * Own technical initiatives from design through production. * Partner closely with Software Engineering, Product, Security, and other teams to solve complex technical problems. * Mentor engineers and contribute to a strong engineering culture. * Help establish engineering and operational best practices as the company scales. * Balance reliability, security, engineering velocity, and business priorities when making technical decisions. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)