ML Platform / Infrastructure Engineer

Hays Specialist Recruitment LLC
Pembroke Pines, FL, United States
7 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Compensation
$215,000.0 - $280,000.0
Working hours
Regular working hours
Job source

Tech stack

Airflow Cloud Engineering Computer Clusters Continuous Integration Distributed Computing Environment Python (Programming Language) Machine Learning Prometheus Azure Machine Learning Delivery Pipeline Large Language Models Grafana
+4 more
Containerization Kubernetes Infrastructure Automation Frameworks Machine Learning Operations

Job description

Build and maintain the infrastructure that powers large-scale AI model training, evaluation, deployment, and monitoring. You’ll partner closely with research engineers to enable reliable, reproducible, and scalable ML workflows. Design and maintain ML training and deployment infrastructure Build scalable MLOps pipelines for model training, evaluation, and release Deploy and operate ML services in production Improve experiment reproducibility, CI/CD, and model lifecycle management Monitor system health, performance, and model observability Optimize infrastructure for reliability, scalability, and cost

Requirements

The final salary or hourly wage, as applicable, paid to each candidate/applicant for this position is ultimately dependent on a variety of factors, including, but not limited to, the candidate’s/applicant’s qualifications, skills, and level of experience as well as the geographical location of the position.

Applicants must be legally authorized to work in the United States. Visa sponsorship not available., Strong Python and software engineering fundamentals Deep experience with AWS and cloud-native architectures Production experience with Kubernetes and containerized workloads Experience building MLOps platforms and ML deployment pipelines Familiarity with distributed training infrastructure Experience with experiment tracking and observability tools (e.g., Weights & Biases, MLflow, Prometheus, Grafana) Experience with CI/CD, Infrastructure as Code, and automation Experience supporting large-scale LLM or VLM training Familiarity with GPU clusters and distributed training frameworks Experience with model serving and inference optimization Knowledge of workflow orchestration tools (e.g., Argo, Airflow, Kubeflow)

About the company

This position is a contract/temporary role where Hays offers you the opportunity to enroll in full medical benefits, dental benefits, vision benefits, 401K and Life Insurance ($20,000 benefit).

Why Hays?

You will be working with a professional recruiter who has intimate knowledge of the industry and market trends. Your Hays recruiter will lead you through a thorough screening process in order to understand your skills, experience, needs, and drivers. You will also get support on resume writing, interview tips, and career planning, so when there’s a position you really want, you’re fully prepared to get it.

Nervous about an upcoming interview? Unsure how to write a new resume?

Visit the Hays Career Advice section to learn top tips to help you stand out from the crowd when job hunting.

Hays is committed to building a thriving culture of diversity that embraces people with different backgrounds, perspectives, and experiences. We believe that the more inclusive we are, the better we serve our candidates, clients, and employees. We are an equal employment opportunity employer, and we comply with all applicable laws prohibiting discrimination based on race, color, creed, sex (including pregnancy, sexual orientation, or gender identity), age, national origin or ancestry, physical or mental disability, veteran status, marital status, genetic information, HIV-positive status, as well as any other characteristic protected by federal, state, or local law. One of Hays’ guiding principles is ‘do the right thing’. We also believe that actions speak louder than words. In that regard, we train our staff on ensuring inclusivity throughout the entire recruitment process and counsel our clients on these principles. If you have any questions about Hays or any of our processes, please contact us.

In accordance with applicable federal, state, and local law protecting qualified individuals with known disabilities, Hays will attempt to reasonably accommodate those individuals unless doing so would create an undue hardship on the company. Any qualified applicant or consultant with a disability who requires an accommodation in order to perform the essential functions of the job should call or text .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all