ML engineer

Whiz Global LLC
Jersey City, NJ, United States
14 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Computer Vision Big Data Program Optimization Software Quality Code Review Computer Programming Continuous Integration Data Structures Distributed Computing Environment
+18 more
Distributed Systems Python (Programming Language) Machine Learning Natural Language Processing Performance Tuning Tensorflow Software Engineering User-Centered Design Pytorch Large Language Models Deep Learning Containerization Scikit Learn Kubernetes Information Technology Machine Learning Operations Data Pipelines Docker

Job description

Lead the end-to-end design, development, and deployment of scalable machine learning models and systems in production.

  • Architect robust, high-performance ML infrastructure and data pipelines that support training, validation, and real-time inference.
  • Drive technical decision-making, setting standards for code quality, model governance, and MLOps practices.
  • Mentor and provide technical guidance to junior and mid-level engineers through code reviews, pairing, and knowledge sharing.
  • Partner with data scientists to operationalize advanced research into reliable, production-ready services.
  • Define and implement monitoring, observability, and automated retraining strategies to ensure model reliability and detect drift.
  • Collaborate with product, engineering, and business leadership to shape ML roadmaps and translate strategic goals into technical deliverables.
  • Evaluate and introduce emerging ML technologies, frameworks, and methodologies to keep the organization at the forefront of innovation.
  • Own the technical health of ML systems, including performance optimization, cost efficiency, and scalability.

Requirements

Bachelor’’s or Master’’s degree in Computer Science, Data Science, Engineering, or a related field.

  • 10+ years of hands-on experience deploying machine learning models in production, with a track record of delivering large-scale systems.
  • Expert-level programming skills in Python, with basic understanding in additional languages such as Java, Scala.
  • Advanced understanding of data structures, algorithms, distributed systems, and software engineering principles.
  • Extensive experience with cloud platforms (AWS) and containerization/orchestration (Docker, Kubernetes).

Deep experience with MLOps tooling (MLflow, Kubeflow, SageMaker, Vertex AI) and CI/CD for ML.

  • Proven experience designing and maintaining large-scale data processing systems.
  • Demonstrated experience leading technical projects and mentoring engineers.

Preferred Qualifications

– Advanced knowledge of deep learning, NLP, computer vision, or large language models.

  • Experience with distributed training, model optimization, and high-throughput serving infrastructure.
  • Understanding of ML frameworks such as TensorFlow, PyTorch, or scikit-learn.

Core Competencies

  • Strategic thinking with strong analytical and problem-solving capabilities.
  • Exceptional communication skills, with the ability to influence both technical and non-technical stakeholders.
  • Proven technical leadership and mentorship abilities.
  • Ability to navigate ambiguity, drive initiatives independently, and manage competing priorities.
  • Strong ownership mindset and commitment to engineering excellence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all