> Markdown version of [/jobs/ext/1452636-software-engineer-cluster-deployment](https://www.wearedevelopers.com/jobs/ext/1452636-software-engineer-cluster-deployment). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Cluster Deployment - **Company:** Cerebras Systems - **Location:** Sunnyvale, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Intelligent Platform Management Interface, Bash Shell, Border Gateway Protocol, Big Data, Client Server Models, Command-Line Interface, Code Review, Data Centers, Dynamic Host Configuration Protocol, Software Debugging, Linux, File Systems, Monitoring of Systems, Networking Hardware, Python (Programming Language), Network Configuration and Change Management, Networking Basics, Ansible, Prometheus, Software Systems, Virtual Local Area Networks, AI Infrastructure, Delivery Pipeline, Grafana, Git, Kubernetes, Deployment Automation, Bare Metal, Free and Open-Source Software, Api Design, Terraform, Network Server - **Published:** July 26, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1ddf4b4bbc68ec98 ## About the Role * 2+ years of mid- to large-scale data center deployment * Strong fundamentals in Python and Bash, with the ability to write scripts and small programs. * Basic Linux experience, including command-line usage, processes, filesystems, and disk troubleshooting. * Working knowledge of Git, including branching, commits, pull requests, and code review. * CS, ECE, or related technical degree, or equivalent practical experience. * Curiosity, strong problem-solving instincts, and willingness to work hands-on with real infrastructure. Preferred Qualifications * Networking fundamentals, including VLANs and routing basics; exposure to BGP, switch configuration, or Arista EOS automation is a plus. * Kubernetes experience or familiarity. * Infrastructure-as-code and GitOps experience, including Terraform, Ansible, and PR-based change control. * Bare-metal provisioning concepts such as PXE, DHCP, iPXE, Redfish, IPMI, and BMC management. * Observability experience with Prometheus or Grafana. * API design and client-server architecture. * Automation side projects or open-source contributions. ## Description We build and operate the software systems that deploy, validate, and manage AI compute clusters across data centers worldwide for the world's fastest AI Inference. Our work turns complex bare-metal infrastructure into repeatable, automated deployment flows spanning server provisioning, network configuration, Kubernetes bring-up, health validation, and operational handoff. As a Software Engineer on the Cluster Deployment Automation team, you will help build the pushbutton tooling that makes large-scale cluster deployments faster, safer, and more reproducible. This role is designed for talented engineers passionate about learning through hands-on work with Python, Bash, Ansible, Linux, bare-metal servers, networking equipment, Kubernetes, and observability systems while learning how production AI infrastructure is built and operated at scale. Responsibilities * Develop and maintain automation for deployment workflows, including provisioning, configuration, validation, and operational handoff. * Turn manual deployment steps into tested, repeatable pushbutton workflows. * Participate in hands-on cluster deployments to build practical debugging and operational expertise. * Troubleshoot issues across Linux systems, bare-metal servers, networking, storage, Kubernetes, and connectivity. * Contribute to infrastructure-as-code and GitOps workflows using tools such as Terraform, Ansible, pull requests, and code review. * Add health checks, observability, dashboards, and validation logic to improve deployment reliability. * Partner with networking, infrastructure, security, and operations teams to deliver secure and reproducible data center deployments. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters)