Machine Learning Platform - Associate - New York
Role details
Job location
Tech stack
Job description
-
Deliver scalable, efficient, secure and automated processes for building, deploying and monitoring Machine Learning models
-
Enable solutions that provide business customers with the ability to leverage the latest and greatest AI/ML infrastructure, frameworks, and tooling to deliver high impact outcomes
-
Develop and demonstrate deep subject matter expertise on how to optimize machine learning model deployments to scale to the specific needs of each business customer
-
Deliver high quality, production ready code leveraging CI/CD best practices
-
Author and maintain high quality documentation for both the engineering team as well as for business customers
Requirements
-
2 years of experience in software engineering (backend, platform, or infrastructure).
-
2 years of experience in Python or a similar backend programming language.
-
1 year of experience supporting production ML systems (MLOps, platform or inference-related work)
-
Basic understanding of APIs (REST or similar) and service-to-service communication.
-
Experience working with containers (e.g., Docker).
-
Familiarity with Unix-based systems.
-
Exposure to public cloud environments (e.g., AWS or GCP), including core concepts such as compute, storage, and basic IAM.
-
Experience working with databases (SQL or NoSQL).
-
Solid grasp of software engineering fundamentals, including debugging, testing, and maintainable code design.
-
Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environment.
-
Curiosity and a strong desire to keep learning-especially in the model inference and LLM platform space.
Preferred Qualifications:
-
4 years of experience in software engineering (backend, platform, or infrastructure)
-
4 years of experience supporting production ML systems (MLOps, platform or inference-related work)
-
4 years of experience in Python or a similar backend programming language.
-
Strong understanding of the end-to-end Model Development Lifecycle (MDLC)
-
Basic understanding of distributed systems concepts and exposure to observability concepts (logging, metrics, tracing).
-
Experience building containerized runtime environments for model serving (e.g. vLLM, SGLang, TensorRT, Triton, AWS Multi Model Server)
-
Experience with infrastructure-as-code tools, such as Terraform or CloudFormation
-
Experience with Kubernetes and other container orchestration platforms in the public cloud (e.g. AWS, GCP)
-
Experience building Machine Learning models with frameworks such as PyTorch and TensorFlow
-
Excellent communication skills and the ability to articulate complex technical concepts to both technical and non-technical stakeholders.