> Markdown version of [/jobs/ext/733762-senior-cloud-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/733762-senior-cloud-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Site Reliability Engineer - **Company:** Solace Corporation - **Location:** Indianapolis, IN, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Software as a Service, Cloud Computing, Software Debugging, Linux, Groovy, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Cloud Services, Prometheus, Datadog, Google Cloud, Multi-Cloud, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Azure AKS, Kibana, Terraform, Golang - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8d7f74cc9cbfec3b ## About the Role Do you have experience in Tooling?, * Highly technical, excited by technology, and eager to stay up to date in a rapidly evolving environment. * Expert-level knowledge in Cloud Networking Solutions * Knowledgeable in demonstrating the ability to debug at a system level and resolve incidents in complex cloud-based environments * Expert in Site reliability engineering and Incident response * A strong communicator who can articulate complex technical issues clearly and concisely & get on the phone with customers. * Experienced in SaaS operations and customer-facing technical support, * Proven expertise with public cloud providers (AWS, Azure, GCP) services & features * Proven expertise with cloud Kubernetes infrastructure platforms such as AWS Elastic Kubernetes Service, Azure Kubernetes Service, Google Kubernetes Service * Hands-on experience with Monitoring tools like Datadog, Kibana, Prometheus etc. * Hands-on experience with Infrastructure Automation using Terraform, Cloud Formation * Hands-on expertise in debugging production alerts * Expert-level understanding of Linux Operating Systems * Programmer in languages such as Groovy, Python, and Go * Certified Kubernetes Administrator * Certified Cloud Administrator (AWS, Azure, or GCP) ## Description This position is for a Senior Cloud Site Reliability Engineer. You will be responsible for the daily operations of Solace Cloud, our market-leading SaaS offering, across leading cloud providers and platforms such as Amazon Web Services, Microsoft Azure, Google Cloud Platform, Kubernetes, etc. What You Will Do: * Ensuring that the Solace Cloud Services are healthy and reliable, and that SLAs are being met * Design and implement our infrastructure tooling, observability, and automation * Contribute to making the production operations more efficient, less error-prone, etc. * Expert-level knowledge in handling production Incidents in production-grade multi-cloud environments according to industry-standard Incident management process * Process handling service requests and provisioning by the customers. * Proven ability to manage customer escalations and drive resolution in mission-critical, high-impact production environments * Work directly with customers to identify, troubleshoot, and resolve operational issues. * Expert debugging knowledge in Linux and Kubernetes to detect operational issues. * Be on-call rotation and provide 24x7 off-hours support ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)