Principal Software Engineer, AI & Data Platform

ELEMYNT LLC
San Diego, CA, United States
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Training Data Application Programming Interfaces (APIs) Artificial Intelligence Software Applications Batch Processing Big Data Cloud Computing Data Deduplication Data Infrastructure Data Recovery Data Systems Distributed Computing Environment
+17 more
Graph Database Python (Programming Language) Machine Learning Open Source Technology Operational Databases Systems Development Life Cycle Scientific Computating Software Engineering Data Processing Scripting Model Validation Generative AI Information Technology Production Code Data Management Machine Learning Operations Data Pipelines

Job description

Elemynt’s platform turns scientific and engineering data into reusable assets for analysis, model training, and automated workflows.

This role owns the data and AI engineering foundation that makes those systems reliable, scalable, and measurable. You will define core patterns for data modeling, training pipelines, evaluation systems, and intelligent workflow interfaces, then prove those patterns in production code.

This is a hands on principal role for someone who can set technical direction and still build the hardest parts themselves.

WHAT YOU WILL DO

  • Architect the data foundation for large scale scientific and engineering output, keeping results clean, queryable, reusable, and ready for model training.
  • Model domain specific scientific data so the same datasets can support interactive analysis, automation, and downstream machine learning workflows.
  • Build scalable data processing patterns across object storage, analytical stores, and training optimized formats.
  • Create machine learning data pipelines for curation, deduplication, formatting, evaluation sets, and regression tracking.
  • Build and operate training and fine tuning pipelines for models used in scientific and workflow driven products.
  • Develop intelligent workflow interfaces that connect user intent, structured platform capabilities and executable workflows without exposing unnecessary complexity to users.
  • Own model evaluation, benchmarking, automated scoring, and quality tracking so each iteration is measurable.
  • Set data and AI engineering standards for the team and turn them into code, documentation, and reusable patterns.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related engineering field, with 10 plus years building and shipping production software.
  • Expert Python and a strong record of shipping systems end to end.
  • Deep experience with large scale data systems, including object storage, analytical processing, training optimized formats, and production data pipelines.
  • Hands on experience building data pipelines for model training, fine tuning, evaluation, and continuous improvement.
  • Direct experience training or fine tuning models for structured outputs, tool use, workflow automation, or domain specific applications.
  • Strong understanding of relational, document, and columnar data models, with judgment about where each belongs.
  • Comfort operating in cloud, enterprise, and technical compute environments, including distributed training or large scale batch processing.
  • Ability to set technical direction in ambiguous early stage environments and carry it through implementation.

NICE TO HAVE

  • Experience applying machine learning to scientific data, such as property prediction, generative models, graph based methods, or simulation data.
  • Experience with atomistic, materials, chemistry, or engineering data systems.
  • Experience with retrieval over structured data, knowledge graphs, or hybrid search systems.
  • Experience designing APIs or tool interfaces that intelligent systems can call reliably.
  • Experience building complex data and machine learning workflows on production orchestrators.
  • Contributions to open source machine learning, data infrastructure, or scientific computing tools., Analysis Skills, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Benchmarking, Cloud Computing, Computer Science, Computer Software, Computer Systems, Continuous Improvement, Data Formats, Data Management, Data Modeling, Data Processing, Data Recovery, Enterprise Computing, Large-Scale Systems, Machine Learning, Open Source, Predictive Modeling, Python Programming/Scripting Language, Scalable System Development, Simulation, Software Engineering, Structured Data, Systems Scalability, Traceability, Training Data Sets

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:37 min

Classifying and anonymizing data during system design

Reto Kaeser · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

4:15 min

Bridging operational and analytical systems using formal data contracts

Matthias Niehoff Matthias Niehoff · World Congress 2024

Videos

See all

Related articles

See all