Senior AI/ML Platform Engineer

Vsg Business Solutions
Plano, TX, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$86,800.0 - $165,200.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Bash Shell Cloud Computing Cloud Engineering Computer Programming DevOps Groovy Monitoring of Systems Python (Programming Language) Machine Learning
+24 more
Automation of Marketing Site Reliability Engineering Practices Cloud Services Ansible Tensorflow Azure Machine Learning Shell Script Scripting Pytorch System Availability Large Language Models Grafana Generative AI Infrastructure as Code (IaC) Cloudformation Scikit Learn Kubernetes Information Technology Machine Learning Operations Virtual Agents Terraform Splunk Dynatrace Docker

Job description

  • Design, develop, and manage AWS-based cloud infrastructure and platform services
  • Build and maintain Infrastructure as Code (IaC) solutions using Terraform
  • Develop AI/ML-powered automation solutions for infrastructure, operations, and platform engineering workflows
  • Implement and optimize CI/CD pipelines to support efficient software delivery and infrastructure deployments
  • Enhance observability and monitoring capabilities using Dynatrace, Grafana, Splunk, and related tools
  • Build intelligent automation for alert triage, incident response, root cause analysis, and operational workflows
  • Collaborate with DevOps, SRE, Engineering, Product, and Operations teams to identify automation opportunities and improve platform reliability
  • Develop and support MLOps processes including model deployment, monitoring, governance, and lifecycle management
  • Create automation scripts using Python, Bash, Shell, or Groovy
  • Support cloud modernization initiatives and platform transformation efforts

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field
  • 7+ years of experience in Cloud Engineering, DevOps, Platform Engineering, or Infrastructure Engineering
  • Strong hands-on experience with AWS cloud services and infrastructure management
  • Expertise in Terraform and Infrastructure as Code (IaC)
  • 3+ years of experience with observability and monitoring tools such as Dynatrace, Grafana, and Splunk
  • Strong experience building and maintaining CI/CD pipelines
  • 2 3+ years of hands-on AI/ML engineering experience in production environments
  • Strong programming and automation skills using Python
  • Experience with scripting languages such as Bash, Shell Script, or Groovy
  • Experience with AI/ML frameworks such as TensorFlow, PyTorch, or Scikit-learn
  • Experience implementing MLOps practices and AI-driven automation solutions
  • Strong understanding of cloud architecture, automation, monitoring, and operational excellence

Preferred Qualifications

  • Experience with Generative AI, LLMs, Agentic AI, RAG, or AI-assisted development tools
  • Experience with Kubernetes, Docker, Ansible, or CloudFormation
  • Java development experience
  • Experience building AI-powered observability or operational automation platforms
  • Experience working in large-scale enterprise or financial services environments
  • Knowledge of SRE practices, incident management, and platform reliability engineering

Required Technical Skills

  • AWS Cloud Services
  • Terraform
  • CI/CD Pipelines
  • DevOps & Platform Engineering
  • Python
  • AI/ML Engineering
  • MLOps
  • Dynatrace
  • Grafana
  • Splunk
  • Bash/Shell/Groovy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all