> Markdown version of [/jobs/ext/1481642-senior-ml-engineer](https://www.wearedevelopers.com/jobs/ext/1481642-senior-ml-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Engineer - **Company:** INTELEPEER INC - **Location:** Dania Beach, FL, United States - **Experience:** Expert - **Salary:** $180,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Analysis, Artificial Neural Networks, Microsoft Azure, Distributed Computing Environment, Machine Learning, Open Source Technology, Raw Data, Search Technologies, Reinforcement Learning, Large Language Models, Multi-Agent Systems, Model Validation, Fastapi, Information Technology, Low Latency, Optimization Algorithms, ONNX (Open Neural Network Exchange) Format, HuggingFace, Machine Learning Operations, Data Generation - **Published:** July 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e5b80a256065f1a9 ## About the Role Bachelors in computer science or statistics * 3-8+ years of hands-on ML engineering experience with a strong production track record. * Deep understanding of core ML concepts: neural network architectures (transformers, attention mechanisms), loss functions, optimization algorithms, regularization, and model evaluation. * Practical experience fine-tuning LLMs (LoRA, QLoRA, PEFT, instruction tuning, DPO) on custom datasets using frameworks such as Hugging Face Transformers, TRL, or Axolotl. * Hands-on experience with RL-based alignment techniques - specifically PPO and GRPO - for reward modeling, preference optimization, and RLHF pipelines. * Experience hosting and serving LLMs: vLLM, TGI, Triton, or similar; understanding of model quantization (GPTQ, AWQ, int4/int8), batching strategies, and throughput optimization. * Working knowledge of major inference vendors and cloud AI APIs; ability to evaluate and select providers based on cost, latency, and capability benchmarks. * Proficiency in embedding models (sentence-transformers, OpenAI embeddings, or equivalent) and vector search infrastructure for RAG pipelines. * Understanding of Model Context Protocol (MCP) and how to expose ML functionality as structured tools for agentic systems. Key Competencies: * Experience with distributed training frameworks (DeepSpeed, FSDP, Megatron-LM) for multi-GPU or multi-node training runs. * Familiarity with MLOps tooling: MLflow, Weights & Biases, DVC, or similar for experiment tracking, model registry, and pipeline orchestration. * Knowledge of synthetic data generation techniques for augmenting fine-tuning datasets. * Exposure to multimodal models (vision-language, speech-language) or voice/speech AI systems. * Contributions to open-source ML projects or published research (papers, blog posts, or technical write-ups). Physical Requirements: * Sedentary work lifting no more than 10 pounds. * Occasional lifting, carrying, and standing. * Frequent hand/eye coordination to operate office equipment. * Vision sufficient to read computer screens, reports, and related department documents. * Dexterity to operate computer keyboards and other related office equipment. * Endurance sufficient to sit and work at a computer for extended periods of time. * Frequent speech communication and hearing. ## Description IntelePeer is building AI-native communications products and we need an ML engineer who gets their hands dirty. This is not a research role - you will own the full lifecycle of machine learning systems: designing training pipelines, fine-tuning and aligning large language models, optimizing inference, and shipping models that run reliably in production. You will work alongside our AI Engineering team to push the capabilities of our platform and deliver measurable impact., * Design, implement, and maintain end-to-end ML training pipelines - from raw data ingestion and preprocessing through model training, evaluation, and deployment. * Fine-tune large language models using techniques such as LoRA, QLoRA, and full fine-tuning; apply PEFT strategies to balance performance and compute cost. * Implement and experiment with reinforcement learning from human feedback (RLHF) workflows, including PPO (Proximal Policy Optimization) and GRPO (Group Relative Policy Optimization) for model alignment and preference optimization. * Host, serve, and optimize LLMs in production using inference frameworks such as vLLM, Text Generation Inference (TGI), Triton Inference Server, or ONNX Runtime. * Evaluate, benchmark, and select inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on latency, cost, throughput, and model capability trade-offs. * Build and maintain embedding pipelines - generate, index, and retrieve dense embeddings using vector databases (Pinecone, pgvector, Weaviate, or similar) for RAG and semantic search applications. * Implement and expose ML capabilities via Model Context Protocol (MCP) - enabling AI agents to call model-backed tools in a structured, context-aware manner. * Perform rigorous data analysis and processing: clean, transform, and curate datasets for training, fine-tuning, and evaluation; build data quality and validation pipelines. * Develop robust model evaluation frameworks - define metrics, build eval harnesses, run A/B experiments, and track regressions across model versions. * Collaborate with software engineers to integrate ML systems into product features via FastAPI services; ensure models are observable, versioned, and maintainable in production. ## Related Videos - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) - [Bringing Clarity to Event Streams: Enabling Analytics and AI Through Rich Metadata](https://www.wearedevelopers.com/videos/1616-bringing-clarity-to-event-streams-enabling-analytics-and-ai-through-rich-metadata) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)