ML Platform Engineer

OpenKyber LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Databases Information Engineering Extract Transform Load (ETL) Distributed Systems Monitoring of Systems Identity and Access Management Python (Programming Language)
+22 more
Machine Learning Performance Tuning Azure Machine Learning Search Technologies Systems Integration Workflow Management Systems AWS Cdk Cloud Platform System Large Language Models Informatica Cloud Generative AI Backend Cloudformation Information Technology Low Latency AWS Glue AWS Data Analytics Machine Learning Operations Cloudwatch Terraform GPT Data Pipelines

Job description

We are seeking a ML Engineer LLM Platforms & Assistants who will design, build, and operate production-grade large language model (LLM) pipelines primarily within AWS-based environments. This role focuses on integrating OpenAI models into modular Python services, implementing Retrieval-Augmented Generation (RAG) and semantic search, and deploying scalable, secure, and observable AI assistants., * Design and maintain LLM integrations using OpenAI APIs within AWS environments.

  • Build Python-based LLM services deployed on AWS compute platforms (ECS, EKS, Lambda, or EC2).
  • Implement RAG workflows and semantic search using AWS data and storage services.
  • Develop LangChain or agentic workflows supporting reasoning and tool use.
  • Integrate LLM pipelines with ETL/ELT workflows and enterprise data systems.
  • Deploy and integrate MCP servers and emerging orchestration tools.
  • Apply AWS security best practices using IAM, KMS, and Secrets Manager.
  • Implement monitoring and observability using CloudWatch and related tools.
  • Migrate custom GPT solutions into production-grade AWS-hosted assistants.

Requirements

12+ years of overall IT development experience, with a strong background in backend and distributed systems. 7+ years of experience in Machine Learning, Data Engineering, or Applied AI engineering. Strong proficiency in Python, with experience building modular, production-grade services. Proven experience implementing Retrieval-Augmented Generation (RAG) and semantic search architectures. Hands-on experience integrating and operationalizing OpenAI LLM APIs in production environments. Solid experience deploying and managing systems within AWS environments, including services such as S3, Lambda, ECS/EKS, and IAM. Experience building scalable, secure, and observable AI/ML systems in production., * Experience working with Amazon SageMaker and/or Amazon Bedrock for model development, deployment, or managed LLM services.

  • Strong familiarity with AWS data services, including AWS Glue, Amazon Athena, Amazon OpenSearch Service, and Amazon Aurora.
  • Hands-on experience designing and implementing ETL/ELT data pipelines in cloud environments.
  • Experience building LLM orchestration pipelines, including reasoning workflows, tool usage, and multi-step agent architectures.
  • Knowledge of LLM benchmarking, evaluation frameworks, and performance optimization (latency, cost, quality metrics).
  • Experience integrating enterprise systems using SnapLogic.
  • Exposure to Craxel Black Forest Time-Series Database (or similar time-series platforms); willingness to learn/train if not previously experienced.
  • Experience implementing Infrastructure as Code (IaC) using AWS CDK, CloudFormation, or Terraform.

Key Skills: Machine Learning, LLM, AWS, Amazon SageMaker / Amazon Bedrock, RAG, Python

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou Β· Coffee With Developers

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 Β· WWC 2024

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 Β· Coffee With Developers

3:35 min

Defining a serverless architecture using AWS CDK

Raphael Manke Raphael Manke Β· WWC 2023

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas Β· LIVE

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky Β· WWC 2025

Videos

See all

Related articles

See all