> Markdown version of [/jobs/ext/2728325-principal-ai-engineer-ai-model-training-data-strategy](https://www.wearedevelopers.com/jobs/ext/2728325-principal-ai-engineer-ai-model-training-data-strategy). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI Engineer - AI Model Training & Data Strategy - **Company:** Maxonic, Inc. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Big Data, Information Engineering, Data Governance, Distributed Computing Environment, Machine Learning, Regression Testing, Management of Software Versions, Feature Engineering, Pytorch, Large Language Models, Apache Spark, Data Strategy, Data Lakes, Scikit Learn, Storage Technologies, Information Technology, HuggingFace, Data Management, Machine Learning Operations, Data Pipelines, Data Generation - **Published:** September 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pho3etyr3u ## About the Role * Strong hands-on experience training and fine-tuning both traditional AI/ML models and LLMs in production * Deep experience with parameter-efficient fine-tuning (QLoRA, LoRA, PEFT), quantization, and the tradeoffs versus full fine-tuning * Proficiency with ML/DL frameworks and libraries (e.g., PyTorch, Hugging Face Transformers/PEFT/TRL, scikit-learn) * Experience building and operating large-scale data pipelines and platforms (e.g., Spark, Ray, dbt, or equivalents) * Strong grasp of data management: dataset storage architecture, versioning, lineage, governance, and PII handling * Experience with experiment tracking and reproducible ML (e.g., MLflow, Weights & Biases) * Understanding of distributed training and GPU/compute optimization Education and Experience * Bachelor's degree in Computer Science, Engineering, or related discipline; advanced degree in ML, AI, or Data Science preferred * 12 or more years of experience in AI/ML engineering, applied ML, or data engineering, with significant hands-on model training and fine-tuning ## Description * Define and own the end-to-end model training strategy across CAI products, spanning traditional AI/ML models and large language models * Fine-tune large language models using parameter-efficient techniques (e.g., QLoRA, LoRA, PEFT) and full fine-tuning where warranted * Train, evaluate, and tune traditional AI/ML models (classification, regression, ranking, clustering, and similar) * Work with large volumes of data - design and optimize pipelines for ingestion, cleaning, labeling, and feature engineering * Define standards for how and where training data from Commercial AI products is stored, versioned, and accessed (data lakes/warehouses, feature stores, dataset registries) * Establish data governance, lineage, quality, licensing/consent, and PII-handling practices for training data * Build reproducible training pipelines and experiment tracking (datasets, hyperparameters, checkpoints, and metrics) * Define evaluation methodology and benchmarks for model quality, including offline evaluation and regression testing * Curate and clean training, validation, and test datasets, including synthetic data generation where appropriate * Optimize training cost and compute utilization (GPU efficiency, distributed training, quantization) ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Unlocking the Power of AI: Accessible Language Model Tuning for All](https://www.wearedevelopers.com/videos/951-unlocking-the-power-of-ai-accessible-language-model-tuning-for-all) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)