> Markdown version of [/jobs/ext/2073456-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2073456-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** OneStream Software LLC - **Location:** Birmingham, MI, United States (Remote available) - **Experience:** Experienced - **Salary:** $114,000.0 - $148,000.0 - **Contract:** Permanent contract - **Skills:** .NET Framework, Agile Methodology, Amazon Web Services, JIRA, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Ubuntu (Operating System), Command-Line Interface, Software as a Service, Cloud Computing, Cloud Computing Security, Configuration Management, Debian Linux, Linux, Github, Identity and Access Management, Integrated Development Environments, Python (Programming Language), Log Analysis, OpenShift, Windows PowerShell, Reliability Engineering, Cloud Services, Ansible, OneStream, Prometheus, Software Engineering, SQL Databases, Datadog, Scripting, Google Cloud, Grafana, Git, Cloudformation, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Production Code, Bicep, Azure AKS, Puppet, Restful APIs, Terraform, New Relic (SaaS), Software Version Control, Dynatrace, Azure Resource Manager - **Published:** August 15, 2026 - **Apply:** https://jobs.mitalent.org/job-seeker/job-details/JobCode/401978984 ## About the Role BS/BA in computer science, engineering, or technology-related field (or equivalent work experience). Proven work experience as a Site Reliability Engineer or in a similar role. 6+ years of cloud infrastructure and software development experience. 2+ years hands on experience of Azure Kubernetes Services (AKS) with container-based deployment skills or other platforms such as OpenShift, GKS, EKS. Advanced understanding of APM and observability tools such as Dynatrace, AppInsights, DataDog, Log Analytics, New Relic, Prometheus and Grafana. Advanced understanding of Infrastructure-as-Code (IaC) concepts and tooling (Terraform, CloudFormation templates, Bicep or ARM templates) on Microsoft Azure, Amazon Web Services (AWS), or Google Cloud Platform (GCP). Deep knowledge of Configuration Management/Orchestration utilities such as Ansible, PowerShell DSC, Chef, and Puppet. Advanced understanding of cloud concepts including elasticity, security, and identity management. Well versed familiarity with Agile Development methodologies utilizing Jira or Azure DevOps Boards. 6+ years of hands-on experience with the following technologies, tools, and concepts: Automating processes using PowerShell, Bash, CLI, REST APIs, python, ARM Templates or other scripting languages. Comfortable leveraging source control tools such as Git, Azure DevOps, or GitHub. Knowledge of container orchestration platforms such as Kubernetes, OpenShift, AKS, GKS or helm. Microsoft Azure, Amazon Web Services (AWS) or Google Cloud (GCP). Preferred Education and Experience Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider. 6+ years of relevant Azure experience deploying and managing leveraging Infrastructure-as-Code (IAC) concepts. Experience with Microsoft and .NET (.NET, C#, SQL). Experience writing efficient and reliable code in a development environment. Debian, Ubuntu, Alpine or other distributions of the Linux operating systems. Deep knowledge and understanding of containerized applications, with special attention to reliability and monitoring of those containerized applications. Knowledge, Skills, and Abilities Deal well with ambiguous/undefined problems. Ability to self-motivate and work independently. Strong organizational and prioritization skills. Ability to find and apply effective solutions to emerging problems and challenges. Strong attention to detail. Comfortable communicating with all levels of management and engineering. Ability to get up to speed quickly with modern technologies and services. Ability to multitask on a variety of projects. ## Description As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If you enjoy staying at the forefront of technology and automating infrastructure deployments, then this is the job for you. This vital role within Cloud Services requires knowledge and experience designing, implementing, and monitoring scalable and secure cloud services. The employee is expected to work well in a small team and willing to share responsibilities with other team members as needed. You will interact with internal staff, managers, and customers to implement and maintain operations. A passion for technology and learning, and the ability to grow others are vital for success in this role. Primary Duties and Responsibilities Implement application/infrastructure observability solutions to ensure desired application availability, reliability, and performance. Participate in regular On-Call rotations and share details related to incidents and their resolution through post-mortem reports and regular review meetings. Proactively partner with Product and Engineering teams to identify, develop, deploy, and maintain reliable systems and services. Influence and create new designs, architectures, standards, and methods for large-scale systems. Sustain a high level of reliability for key services and automated systems. Automate processes to improve reliability, performance, and availability. Update technical documentation, workflows, and knowledge base articles. Provide feedback in pull requests and peer coding reviews. Implement codified automated solutions that build integrations between Dynatrace, Azure DevOps and Jira. Solid knowledge in focused areas of OneStream Software. Ability to mentor others in several technical areas. Understanding practical use of SOC/FedRAMP controls to assist Compliance and Security teams. Required Education and Experience, Transparency around corporate structure, salary, and benefits. Core value of customer success. Variety of project work (not industry-specific). Strong culture andcamaraderie. Multiple training opportunities. ## Related Videos - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)