> Markdown version of [/jobs/ext/165489-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/165489-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - SRE - **Company:** Open Practice Solutions, Ltd. - **Location:** Hudson, OH, United States - **Experience:** Expert - **Salary:** $85,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Amazon Elastic Compute Cloud, Software as a Service, Databases, Continuous Integration, Data Stores, Data Warehousing, Linux, DevOps, Fuzz Testing, Memcached, MySQL, Nagios, Performance Tuning, Redis, Reliability Engineering, Web Applications, Data Logging, Java Application Server, Amazon ElastiCache, Grafana, Gitlab, Graphite, Cloudwatch, Terraform, Amazon Redshift - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e829e35262902277 ## About the Role * 3+ years of experience in SRE, DevOps, or production operations roles * Strong understanding of AWS infrastructure and cloud-native scaling patterns * Experience supporting Java applications in production * Solid knowledge of MySQL performance, replication, and scaling strategies * Experience operating cache layers and data stores at scale * Understanding of multi-tenant architectures, including isolation, noisy-neighbor issues, and capacity planning * Strong Linux fundamentals and troubleshooting skills * Ability to stay calm, think clearly, and prioritize during incidents * A "get-things-done" mindset - pragmatic, resourceful, and comfortable with imperfect systems Nice to Have * Experience scaling multi-tenant SaaS platforms * Familiarity with Redshift performance tuning and data workflows * Infrastructure-as-code experience (Terraform) * CI/CD and GitLab pipeline experience * Prior ownership of on-call rotations and incident processes * Experience improving reliability without large architectural rewrites What We Value * Engineers who work within reality, not just ideal architectures * Incremental improvements that reduce risk and improve uptime * Clear communication during incidents * Ownership, accountability, and practical problem-solving ## Description We're looking for a Site Reliability Engineer to help operate and scale a multi-tenant, web-based application running on AWS. This is a hands-on role for someone who's comfortable jumping into an already-established architecture, making incremental improvements, and solving real production problems. You'll work closely with engineering and product teams to keep our platform reliable, performant, and scalable as customer usage grows. This is not a "greenfield rewrite" role - we need someone scrappy, practical, and effective inside real-world constraints. What You'll Do * Ensure the reliability, availability, and performance of a multi-tenant production system * Scale and operate AWS-based infrastructure supporting a Java web application * Monitor and troubleshoot issues across application, database, cache, and data warehouse layers * Improve observability through metrics, logging, and alerting * Participate in on-call rotations and lead incident response and root cause analysis * Identify performance bottlenecks and scaling limits in a shared-tenant environment * Automate operational tasks and reduce toil where it matters most * Work within existing frameworks and tooling to make systems safer and more scalable * Partner with developers to improve deployments, capacity planning, and failure handling * Implement automated load and fuzz testing * Define key service level objectives (SLO) Technologies You'll Work With * AWS (EC2, ECS, RDS, ElastiCache, Redshift, and related services) * Java-based web applications * MySQL (performance tuning, scaling, reliability) * Amazon ElastiCache (Redis/Memcached) * Amazon Redshift * Monitoring and alerting tools (Graphite, Grafana, Cloudwatch) ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)