> Markdown version of [/jobs/ext/1918884-principal-software-engineer-ai-data-platform](https://www.wearedevelopers.com/jobs/ext/1918884-principal-software-engineer-ai-data-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer, AI & Data Platform - **Company:** ELEMYNT LLC - **Location:** San Diego, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Application Programming Interfaces (APIs), Artificial Intelligence, Software Applications, Batch Processing, Big Data, Cloud Computing, Data Deduplication, Data Infrastructure, Data Recovery, Data Systems, Distributed Computing Environment, Graph Database, Python (Programming Language), Machine Learning, Open Source Technology, Operational Databases, Systems Development Life Cycle, Scientific Computating, Software Engineering, Data Processing, Scripting, Model Validation, Generative AI, Information Technology, Production Code, Data Management, Machine Learning Operations, Data Pipelines - **Published:** August 4, 2026 - **Apply:** https://www.careerbuilder.com/job-details/principal-software-engineer-ai-data-platform-xora-portfolio-company-san-diego-ca--00212839-d213-49a1-ade5-ccb09aa01ef8 ## About the Role * Bachelor's or Master's degree in Computer Science or a related engineering field, with 10 plus years building and shipping production software. * Expert Python and a strong record of shipping systems end to end. * Deep experience with large scale data systems, including object storage, analytical processing, training optimized formats, and production data pipelines. * Hands on experience building data pipelines for model training, fine tuning, evaluation, and continuous improvement. * Direct experience training or fine tuning models for structured outputs, tool use, workflow automation, or domain specific applications. * Strong understanding of relational, document, and columnar data models, with judgment about where each belongs. * Comfort operating in cloud, enterprise, and technical compute environments, including distributed training or large scale batch processing. * Ability to set technical direction in ambiguous early stage environments and carry it through implementation. NICE TO HAVE * Experience applying machine learning to scientific data, such as property prediction, generative models, graph based methods, or simulation data. * Experience with atomistic, materials, chemistry, or engineering data systems. * Experience with retrieval over structured data, knowledge graphs, or hybrid search systems. * Experience designing APIs or tool interfaces that intelligent systems can call reliably. * Experience building complex data and machine learning workflows on production orchestrators. * Contributions to open source machine learning, data infrastructure, or scientific computing tools., Analysis Skills, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Benchmarking, Cloud Computing, Computer Science, Computer Software, Computer Systems, Continuous Improvement, Data Formats, Data Management, Data Modeling, Data Processing, Data Recovery, Enterprise Computing, Large-Scale Systems, Machine Learning, Open Source, Predictive Modeling, Python Programming/Scripting Language, Scalable System Development, Simulation, Software Engineering, Structured Data, Systems Scalability, Traceability, Training Data Sets ## Description Elemynt's platform turns scientific and engineering data into reusable assets for analysis, model training, and automated workflows. This role owns the data and AI engineering foundation that makes those systems reliable, scalable, and measurable. You will define core patterns for data modeling, training pipelines, evaluation systems, and intelligent workflow interfaces, then prove those patterns in production code. This is a hands on principal role for someone who can set technical direction and still build the hardest parts themselves. WHAT YOU WILL DO * Architect the data foundation for large scale scientific and engineering output, keeping results clean, queryable, reusable, and ready for model training. * Model domain specific scientific data so the same datasets can support interactive analysis, automation, and downstream machine learning workflows. * Build scalable data processing patterns across object storage, analytical stores, and training optimized formats. * Create machine learning data pipelines for curation, deduplication, formatting, evaluation sets, and regression tracking. * Build and operate training and fine tuning pipelines for models used in scientific and workflow driven products. * Develop intelligent workflow interfaces that connect user intent, structured platform capabilities and executable workflows without exposing unnecessary complexity to users. * Own model evaluation, benchmarking, automated scoring, and quality tracking so each iteration is measurable. * Set data and AI engineering standards for the team and turn them into code, documentation, and reusable patterns. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)