> Markdown version of [/jobs/ext/2557499-ai-cloud-senior-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2557499-ai-cloud-senior-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Cloud Senior DevOps Engineer - **Company:** Bitdeer Technologies Group - **Location:** San Jose, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Software Applications, Application Performance Management, Build Automation, Microsoft Azure, Cloud Computing, Cloud Engineering, Computer Clusters, Computer Programming, Continuous Integration, DevOps, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Key Management, Machine Learning, Performance Tuning, Reliability Engineering, Ansible, Prometheus, Zero Trust Network Access, Security Software, TCP/IP, AI Infrastructure, Data Logging, Scripting, Google Cloud, Load Balancing, Computer Network Technologies, Real Time Systems, System Availability, Grafana, Multi-Cloud, HybridCloud, Infrastructure as Code (IaC), Containerization, AI Platforms, Kubernetes, Information Technology, Machine Learning Operations, Terraform, Devsecops, Docker - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1202ac8547e11bcd ## About the Role * Experience & Education: Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles. * Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs). * Containerization & Orchestration: Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices. * Cloud Platforms: Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies. * Programming Skills: Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development. * Domain Knowledge: Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles. * Soft Skills: Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment., * AI/ML Infrastructure Experience: Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads. * Large-Scale Systems: Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms). * Platform Engineering: Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity. * Advanced Security: Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001). * Leadership: Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams. ## Description We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role, you will be the backbone of our deployment and infrastructure operations, ensuring that our AI-powered products and platforms are delivered with speed, security, and exceptional reliability. You will act as a crucial bridge between our research/development teams and real-world deployment, driving automation, optimizing cloud-native architectures, and establishing best practices for MLOps and traditional DevOps workflows., * CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production. * Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing. * High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs. * Infrastructure as Code (IaC): Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments. * Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics. * Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency through Internal Developer Platforms (IDP) and Platform Engineering initiatives. * Governance, Security & Compliance: Establish and enforce robust system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance with industry frameworks (e.g., SOC2, ISO27001). * Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, conduct thorough root cause analysis (RCA), and implement preventative remediation plans. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)