> Markdown version of [/jobs/ext/3038671-cloud-engineersite-reliability-architect](https://www.wearedevelopers.com/jobs/ext/3038671-cloud-engineersite-reliability-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud EngineerSite Reliability Architect - **Company:** Slalom, LLC - **Location:** Austin, TX, United States - **Salary:** $149,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Cloud Computing, Continuous Integration, Software Debugging, Python (Programming Language), Network Interface, Reliability Engineering, Cloud Services, Azure Machine Learning, Google Cloud, Multi-Cloud, AWS ECS, Kubernetes, Hardware Infrastructure, Terraform, Oracle Cloud Infrastructure - **Published:** September 23, 2026 - **Apply:** https://www.careerbuilder.com/job-details/cloud-engineersite-reliability-architect-austin-tx--df4fb5c1-4729-4b54-8dfa-d6efc9cba27b ## About the Role * Hands-on Kubernetes experience including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across fleets. * Strong Terraform proficiency and experience writing and reviewing production infrastructure as code on a daily basis. * Practical Python skills for tooling and automation, including the ability to write code without relying on artificial intelligence assistance. * Strong critical-thinking and debugging skills, with a disciplined approach to forming hypotheses, validating evidence, and isolating root causes. * Ability to read unfamiliar code, understand its behavior, and diagnose issues effectively. * Resourcefulness and a consistent ability to self-unblock using documentation, telemetry, code, peers, and available tooling. * A proactive working style: you arrive with proposals, formulate solutions, and communicate informed technical opinions. * Clear communication and effective collaboration across engineering teams and technical stakeholders. * Working knowledge of AWS or similar Cloud services relevant to compute and data-intensive platforms, including Amazon EC2, Amazon S3, Amazon EFS, and Amazon FSx for Lustre. * Experience designing or operating CI/CD pipelines and automated infrastructure provisioning workflows. * Knowledge of cloud networking, storage, security, observability, reliability engineering, and platform governance. Nice-to-have: * Experience using AI coding tools responsibly: you remain accountable for the solution, validate generated code, identify edge cases, and reject unnecessary or incorrect output. ## Description This is a Kubernetes-heavy platform engineering role supporting large-scale, multi-cloud GPU infrastructure. Success requires deep operational judgment, disciplined debugging, strong automation skills, and the ability to protect platform stability while capacity grows rapidly. What You'll Do * Operate Kubernetes platforms at significant scale across providers, including Amazon Elastic Kubernetes Service (EKS), CoreWeave Kubernetes Service (CKS), and Google Kubernetes Engine (GKE). * Own the Kubernetes cluster lifecycle, including node-pool design and management, scheduler troubleshooting, container network interface (CNI) and network-policy troubleshooting, capacity planning, and safe rolling upgrades across large fleets. * Develop, review, and maintain infrastructure as code with Terraform, along with Python tooling and automation that improve reliability, repeatability, and operational efficiency. * Define and maintain service-level indicators (SLIs) and service-level objectives (SLOs); build monitoring and alerting that surface meaningful risks before they affect workloads. * Debug complex distributed-system failures by forming hypotheses, testing them methodically, and separating temporal correlation from causation. * Provision high-performance computing (HPC) and GPU infrastructure through the Conveyor CI/CD system across AWS, CoreWeave, Google Cloud Platform (GCP), and Oracle Cloud Infrastructure (OCI), with additional providers as the platform expands. * Coordinate daily with Networking, Storage, Security, and AI/ML platform teams to resolve cross-system dependencies and improve the end-to-end developer and researcher experience. * Bring forward well-reasoned proposals, solutions, and informed opinions; use available tools and resources to self-unblock and drive issues to resolution. ## Related Videos - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Celery on AWS ECS - the art of background tasks & continuous deployment](https://www.wearedevelopers.com/videos/561-celery-on-aws-ecs-the-art-of-background-tasks-continuous-deployment) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)