Principal Machine Learning Engineer (MLE)

Equinix
Dallas, TX, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Google App Engines Microsoft Azure Continuous Integration Data Transformation Software Design Patterns Python (Programming Language) Machine Learning Tensorflow Azure Machine Learning Software Engineering
+15 more
Google Cloud Cloud Platform System Feature Engineering Chatbots Data Ingestion Pytorch Large Language Models Multi-Cloud Generative AI Build Management Containerization Kubernetes Machine Learning Operations Software Version Control Docker

Job description

Experteer Overview As a Principal Machine Learning Engineer, you design, build, and scale ML and generative AI systems powering real-world products. You collaborate with AI and business teams to translate advanced ML/LLM capabilities into production-grade solutions across GCP, AWS, and Azure. The role blends ML, software engineering, and MLOps to deliver robust, scalable cloud-native systems. You will work on end-to-end pipelines and multi-cloud deployments that drive measurable impact, shaping how AI enables business outcomes. Compensation / Benefits * Design, develop, and deploy ML and LLM-based solutions for production use cases * Collaborate with Generative AI Center of Excellence leaders and stakeholders to evaluate buy vs. build for generative AI * Develop end-to-end ML pipelines (data ingestion, feature engineering, training, evaluation, deployment, monitoring) * Architect and implement LLM-powered systems across multiple clouds into a unified solution * Optimize ML workflows for performance, scalability, reliability, and cost in cloud environments * Implement and maintain MLOps best practices (CI/CD, model versioning, experiment tracking, retraining) * Work with PyTorch and TensorFlow for model development * Containerize ML services and deploy via Docker, Kubernetes, App Engine, or VMs * Apply NLP fundamentals (transformers, attention, embeddings, preprocessing) * Deploy and manage models in production, conduct A/B testing, measure performance with statistics * Develop features, run experiments, translate insights into improvements * Build and deploy classical ML models and NLP/vision applications (sentiment, summarization, Q&A, chatbots, CV tasks) Tasks * PhD with 5+ years, Master with 6+ years, or Bachelor with 7+ years in ML/CS/Data Science or related field * Strong Python proficiency for ML and production systems * Solid software engineering fundamentals, system design, and design patterns * Hands-on experience with at least one major cloud platform (GCP, Azure, AWS) * Experience building and deploying production-grade ML systems * Strong communication skills to explain technical concepts to diverse stakeholders * Excellent time management, collaboration, and organizational skills Key requirements * Employee Assistance Program * Health, life, disability insurance * Retirement plans * Paid Time Off (PTO) and holidays * Equity may be offered * Paid vacations and holidays

Requirements

for performance, scalability, reliability, and cost in cloud environments * Implement and maintain MLOps best practices (CI/CD, model versioning, experiment tracking, retraining) * Work with PyTorch and TensorFlow for model development * Containerize ML services and deploy via Docker, Kubernetes, App Engine, or VMs * Apply NLP fundamentals (transformers, attention, embeddings, preprocessing) * Deploy and manage models in production, conduct A/B testing, measure performance with statistics * Develop features, run experiments, translate insights into improvements * Build and deploy classical ML models and NLP/vision applications (sentiment, summarization, Q&A, chatbots, CV tasks) Tasks * PhD with 5+ years, Master with 6+ years, or Bachelor with 7+ years in ML/CS/Data Science or related field * Strong Python proficiency for ML and production systems * Solid software engineering fundamentals, system design, and design patterns * Hands-on experience with at least one major cloud platform (GCP, Azure, AWS) * Experience building and deploying production-grade ML systems * Strong communication skills to explain technical concepts to diverse stakeholders * Excellent time management, collaboration, and organizational skills Key requirements * Employee Assistance Program * Health, life, disability insurance * Retirement plans * Paid Time Off (PTO) and holidays * Equity may be offered * Paid vacations and holidays

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:00 min

Introduction to chatbot infrastructure and cloud challenges

Stan Girard Stan Girard · World Congress 2024

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all