Cloud & DevOps Engineer (ML Infrastructure)

Tanisha Systems Inc
Austin, TX, United States
22 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Microsoft Azure Big Data Cloud Computing Cloud Computing Security Computer Programming Computer Networks Continuous Integration Linux DevOps Distributed Data Store
+31 more
Distributed Systems Github Monitoring of Systems Identity and Access Management Key Management Linux Kernel Machine Learning Enterprise Messaging Systems Reliability Engineering Prometheus Software Engineering Data Streaming AI Infrastructure Google Cloud Cloud Platform System System Availability Delivery Pipeline Grafana Apache Spark Skills Cloud Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Apache Flink Apache Kafka Machine Learning Operations Video Streaming Terraform Docker Jenkins

Job description

We are seeking a highly skilled Senior Cloud & DevOps Engineer to design, build, and operate scalable cloud-native platforms that support modern data, machine learning, and application workloads. The ideal candidate will have strong expertise in Kubernetes, cloud infrastructure, Linux systems, and distributed data platforms, along with hands-on experience building secure and reliable CI/CD and platform automation solutions., Design, deploy, and manage cloud-native infrastructure on AWS or other public cloud platforms. Build and maintain Kubernetes-based platforms for scalable application and ML workload deployments. Develop infrastructure automation and operational tooling using Go or Java. Administer and troubleshoot Linux systems, networking, and distributed environments. Implement and manage cloud security controls, IAM policies, and platform governance. Support machine learning teams by enabling scalable ML infrastructure and deployment pipelines. Deploy and operate big data and streaming technologies such as Kafka, Spark, and Flink. Build and enhance CI/CD pipelines, Infrastructure-as-Code (IaC), and platform observability solutions. Monitor platform reliability, performance, security, and cost optimization. Collaborate with engineering, data, and ML teams to improve developer productivity and operational excellence.

Requirements

Bachelor’s degree in computer science, Engineering, or a related technical field. 5-7 years of experience in Software Engineering, Platform Engineering, SRE, or DevOps roles. 3-4 years of hands-on experience with AWS or other cloud platforms (Azure/Google Cloud Platform). 3-4 years of experience deploying and operating Kubernetes in production environments. Strong programming experience in Go and/or Java. Deep understanding of Linux internals, system administration, troubleshooting, and networking concepts. Experience with cloud IAM, security best practices, secrets management, and access control. Hands-on experience with at least one big data technology such as Apache Spark, Apache Flink, and messaging platforms like Apache Kafka. Experience with containerization technologies such as Docker. Strong understanding of CI/CD pipelines, automation, and Infrastructure as Code.

Preferred Qualifications Exposure to Machine Learning platforms, MLOps, or AI infrastructure. Experience with Terraform, Helm, ArgoCD, GitHub Actions, or Jenkins. Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience supporting large-scale distributed systems in production. Cloud certifications (AWS, Kubernetes, or equivalent) are a plus.

Key Skills Cloud: AWS, Azure, Google Cloud Platform Containers & Orchestration: Kubernetes, Docker, Helm Languages: Go, Java Operating Systems: Linux Administration & Internals Data & Streaming: Kafka, Spark, Flink Security: IAM, Cloud Security, Secrets Management DevOps: CI/CD, IaC, Automation, Monitoring ML Exposure: MLOps, Model Deployment, ML Infrastructure

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all