> Markdown version of [/jobs/ext/3473003-manager-of-infrastructure-cloud-operations-it-chaos-conductor](https://www.wearedevelopers.com/jobs/ext/3473003-manager-of-infrastructure-cloud-operations-it-chaos-conductor). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager of Infrastructure, Cloud Operations & IT (Chaos Conductor) - **Company:** Ron Turley Associates, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Cloud Computing Security, Cloud Engineering, Collaborative Software, Databases, Continuous Integration, Data Infrastructure, DevOps, Identity and Access Management, Information Technology Operations, Reliability Engineering, Prometheus, Software Vulnerability Management, AI Infrastructure, Datadog, Application Enhancement Tool, System Availability, Grafana, Mttr, Database Performance, Cloudformation, Solid Principles, Information Technology, Build Tools, Machine Learning Operations, Terraform, Splunk, New Relic (SaaS), Docker, Microservices - **Published:** September 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=01866618bc3a55d8 ## About the Role Our ideal candidate is a hands-on technical operations leader who has grown beyond simply being the strongest technical person in the room. You know the technology, but you also know how to lead people, establish operational discipline, prioritize competing demands, communicate with executives, and build systems that scale. You'll likely bring: * 7+ years of progressive experience in cloud infrastructure, DevOps, SRE/platform engineering, infrastructure operations, or a closely related technical discipline. * Meaningful hands-on experience supporting AWS production environments. * Demonstrated experience leading and developing technical teams, including setting expectations, creating accountability, coaching performance, and helping strong technical people grow. * Experience owning or significantly influencing production reliability, availability, monitoring, and incident management. * Strong knowledge of modern DevOps, CI/CD, observability, infrastructure-as-code, and automation practices. * Strong understanding of modern cloud architecture and software design principles, including microservices and how infrastructure decisions affect application reliability and scalability. * Experience with technologies such as AWS, Docker, Terraform/CloudFormation, Grafana, Prometheus, Datadog, New Relic, Splunk, or comparable platforms. * The ability to translate complex technical problems into clear priorities, risks, tradeoffs, and recommendations for business and executive stakeholders. * Experience using AI and/or automation to improve technical, engineering, infrastructure, or operational workflows. * Strong organizational and prioritization skills. * A leadership style that combines confidence with humility. You're comfortable making decisions, receiving feedback, engaging in healthy conflict, and changing direction when the facts tell you to., * You've led across multiple disciplines such as Cloud, DevOps/SRE, Infrastructure, Security, Database Engineering, or IT. * You've successfully led remote or distributed technical teams. * You've implemented AIOps, intelligent observability, predictive scaling, automated remediation, AI-assisted incident response, or other AI-powered operational workflows. * You have experience with security frameworks, vulnerability management, cloud security, or compliance environments. * You understand database resiliency, high availability, backup/recovery, and performance considerations. * You're familiar with AI infrastructure concepts such as model serving, compute-resource management, vector databases, or LLM integration patterns. * You hold relevant certifications such as AWS Solutions Architect, ITIL, or other cloud, DevOps, security, infrastructure, or AI-related credentials. A bachelor's degree in Computer Science, MIS, Engineering, or a related discipline is welcomed but not required. We care much more about what you've built, led, learned, and improved., We're looking for a rare combination: the technical credibility to understand the systems, the leadership ability to make multiple technical disciplines better, the operational discipline to create reliability and accountability, and the curiosity to keep pushing RTA forward. You don't have to personally be the deepest expert in Cloud, DevOps, Security, Databases, Infrastructure, and IT. You do need to know enough to ask the right questions, recognize risk, establish priorities, develop strong technical people, make good decisions, and know when it's time to roll up your sleeves. If that sounds like you-and this sounds like your type of company-keep reading., * Ability to sit or stand for extended periods. * Ability to work at a computer for prolonged periods. * Ability to communicate effectively through video, phone, messaging, and other remote collaboration tools. * Occasional travel to RTA's corporate headquarters in Glendale, Arizona and/or other company or team gatherings based on business needs. ## Description Own the operational health of RTA's cloud infrastructure and the teams and systems that support it. Continuously improve reliability, availability, scalability, observability, capacity planning, incident response, database resiliency, security posture, and operational readiness. As RTA grows, our technology foundation needs to grow with us-without reliability, security, or performance becoming an afterthought. 2. Build & Lead a High-Performing Infrastructure and Operations Team Create clarity, accountability, strong operating rhythms, and professional growth across Cloud Engineering, Infrastructure, DevOps, Database Engineering, Security Engineering, and IT. You won't just coordinate technical specialists. You'll develop people, establish priorities, remove roadblocks, drive accountability, and help multiple technical disciplines operate as one team. Because this is a remote role, you'll know how to create connection, visibility, accountability, and strong communication across a distributed team. 3. Turn AI & Automation Into Measurable Operational Improvement Move AI and automation from experimentation to actual business and operational value. Reduce manual toil, improve monitoring and observability, accelerate incident response, strengthen root-cause analysis, automate repeatable work, and increase team productivity. Develop and maintain an AI and automation roadmap for Infrastructure and IT Operations, prioritizing opportunities based on business impact, reliability, efficiency, risk, and measurable ROI. You don't need to be an ML engineer. You do need to experiment, learn quickly, challenge old ways of working, and turn emerging technology into meaningful operational improvement., Guide the architecture, performance, scalability, availability, monitoring, alerting, and capacity planning of RTA's AWS environment. Use strong observability practices, predictive analytics, and intelligent capacity planning to identify potential issues before they become incidents. DevOps & Engineering Operations Strengthen CI/CD, infrastructure-as-code, deployment reliability, automation, observability, and the operational practices that help Engineering move quickly without sacrificing stability. Production Incident Management Provide calm, structured leadership when things go sideways. Improve incident response and resolution times, and ensure post-incident reviews turn lessons learned into lasting improvements. Infrastructure & Platform Operations Maintain a forward-looking view of RTA's technical infrastructure, identify risks and capacity needs, and build roadmaps that support company and product growth. Database & Data Platform Reliability Guide database performance, resiliency, capacity planning, backups, recovery, tuning, and high availability. Security Engineering Guide vulnerability management, cloud security hardening, security engineering priorities, compliance initiatives, and incident-response preparedness. Ensure security risks are identified, prioritized, communicated, and addressed as RTA scales. IT Strategy & Oversight Set strategic direction for endpoint management, identity and access management, networking, internal tooling, and the employee technology experience while empowering the IT Manager to lead day-to-day execution. AI & Intelligent Automation Evaluate AI-powered tools, AIOps, intelligent monitoring, predictive analytics and scaling, automated remediation, AI-assisted development tools, automated runbooks, and other technologies that can improve reliability and team effectiveness. Hands-On Technical Leadership Stay close enough to the technology to jump in when needed-whether that's troubleshooting an AWS configuration, reviewing a deployment issue, investigating a security alert, evaluating observability data, refining alerting thresholds, or helping the team through a difficult production incident. People Leadership Lead, coach, develop, and create accountability across a multidisciplinary technical team. Establish clear goals and operating rhythms and make sure people understand not only what matters, but why. Cross-Functional Partnership Work closely with Engineering, Product, Support, Security, and other teams so reliability, scalability, security, and operational considerations are built into decisions early-not discovered after something breaks., * Reliability & Scale: RTA's production environment is reliable, observable, resilient, and prepared for growth, with capacity and infrastructure needs anticipated before they become emergencies. * Incident Performance: Operational issues are detected earlier and resolved faster, with measurable improvement in incident response, MTTR, root-cause analysis, and post-incident learning. * AI & Automation: RTA has a clear AI and automation roadmap, manual toil is decreasing, and initiatives are producing measurable improvements in efficiency, productivity, reliability, or response time-not just interesting demos. * Security & Resilience: Security vulnerabilities, compliance priorities, database resiliency, backup/recovery, and infrastructure risks are proactively identified, prioritized, and managed. * Team Performance: Cloud, Infrastructure, DevOps, Database, Security, and IT teams have clear priorities, ownership, accountability, and development expectations-and your people are growing. * Executive Visibility: Senior leadership understands system health, technical risk, investment priorities, and where RTA is headed next. Who Thrives at RTA? In general, someone who: * Is passionate about serving others. * Thinks of themselves less, while not thinking less of themselves. You're confident, yet humble. * Is comfortable being part of a team that thrives on healthy conflict. We challenge ideas, ask hard questions, and speak the kind truth. * Passionately cares about our clients. Our clients are fleet managers, parts clerks, and automotive technicians who maintain everything from squad cars to school buses-so everyone comes home safely at the end of the day. * Believes no job is beneath them. Sometimes leadership means setting strategy. Sometimes it means jumping in and helping solve the problem. * Takes ownership and initiative, identifying how to make processes, systems, and teams better without waiting for permission. * Loves to read, learn, grow, and stretch themselves. * Is curious about AI and emerging technology and looks for practical ways to use them to work smarter and create better outcomes., As part of our commitment to creating a fair, efficient, and consistent hiring process, we may use artificial intelligence (AI) to help our recruiting teams organize, summarize, and analyze information provided by candidates, including resumes, application responses, and other materials submitted during the application process. AI may be used to identify patterns, highlight relevant skills and experience, and assist in comparing a candidate's qualifications with the requirements of a specific role. These tools are used to improve efficiency and consistency while supporting more informed hiring decisions, which will ultimately be made by the hiring team. Notice Regarding the Use of Artificial Intelligence in Application Review ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons)