Principal AI Engineer & Data Analyst

Artificial Intelligence
United States
4 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Airflow Amazon Web Services Business Analytics Applications Microsoft Azure Big Data Cloud Database Encodings Information Engineering Data Infrastructure Extract Transform Load (ETL) Data Transformation
+20 more
Database Queries Python (Programming Language) Machine Learning Object-Oriented Software Development Query Optimization Software Deployment Software Engineering Data Logging Google Cloud Azure Data Factory Large Language Models Apache Spark Generative AI Pytest Apache Flink Luigi AWS Glue Data Analytics Software Version Control Devsecops

Job description

Shape Safer Skies with Artificial Intelligence! Concepts Beyond is seeking a hands-on AI / Data Engineer & Analyst to serve as our AI technical lead, powering aviation safety and air traffic management programs through enterprise-scale data engineering, applied machine learning, and intelligent system design. This is a high-ownership, individual-contributor role - you will personally architect pipelines, train and fine-tune models on mission-critical corpora, and deploy analytics solutions that drive real decisions in the National Airspace System. We welcome candidates from a range of backgrounds, including recent Ph.D. graduates, postdoctoral researchers, government or military technologists moving into industry, and engineers from academic or startup environments - sharp, eager, fast-learning technologists ready to grow into a highly visible, public-facing technical leadership role.

Requirements

  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or a related technical field from a US accredited institution. Recent Ph.D. graduates and postdoctoral researchers with strong applied AI/ML or NLP research backgrounds are strongly encouraged to apply. \n

  • 5+ years of relevant experience in data engineering, applied AI/ML research, or a related technical discipline. Equivalent experience - via a Ph.D./postdoctoral research, federal or military technical service, or startup engineering roles - will be considered in place of traditional industry tenure. \n

  • Proficient in Python with strong software engineering practices: OOP, testing frameworks (pytest), logging, error handling, and version control. \n

  • Expertise in ETL/ELT orchestration (Apache Airflow, Prefect, Luigi, AWS Glue, or Azure Data Factory); deep SQL proficiency including query optimization and index tuning. \n

  • Data modeling expertise: normalization, star/snowflake schemas, slowly changing dimensions (Type 1/2). \n

  • Experience with big data processing frameworks (Apache Spark, Flink) and cloud data ecosystems (AWS, Azure, GCP). \n

  • Hands-on experience custom-developing AI/ML solutions (LLMs/SLMs, real-time voice/speech), and predictive data analytics; pipelines, preprocessing, embedding, grounding, and production deployment. \n

  • Working knowledge of Generative AI and RAG architectures, vector databases, and enterprise data infrastructure integration with model versioning, monitoring, and rollback strategies. \n

  • Fast learner with strong communication skills - able to explain complex technical work clearly to both technical and non-technical audiences, and eager to grow into public speaking and industry engagement. \n

  • Must be located in the U.S. or authorized to work in the U.S.; demonstrable portfolio of production AI/ML or data engineering work required. \n, * FAA domain or Aviation Safety systems exposure a plus but not required (e.g., ASIAS, SWIM, Foundry) - strong candidates from adjacent regulated domains (finance, healthcare, defense) with deep AI/ML skills will also be considered; domain training is provided.

Benefits & conditions

n \n

  • Develop, train, and operationalize NLP/ML models for low-latency real-time voice pipelines using streaming speech-to-text and text-to-speech, diarization, classification, and named entity recognition over controller-pilot voice and text. \n

  • Develop Retrieval-Augmented Generation (RAG) pipelines over enterprise vector stores with hybrid retrieval, re-ranking, and grounded evaluation. \n

  • Custom-develop, fine-tune, and deploy large and small language models (LLMs and SLMs) for real-time operational analysis and decision support; build streaming NLP and agentic architectures that integrate with enterprise aviation platforms. \n

  • Develop AI/ML solutions for predictive analytics in aviation safety; probabilistic modeling, time-series and anomaly detection, and causal-factor analysis on ASIAS/FOQA/ASRS and related data. \n

\n

Data Engineering & Infrastructure

\n \n

  • Architect, build, and maintain scalable data pipelines for structured, semi-structured, and unstructured data using orchestration tools (e.g., Apache Airflow, Prefect, AWS Glue, or Azure Data Factory). \n

  • Design and implement robust ETL/ELT processes with strong error handling, idempotency, monitoring, and dependency management across cloud and hybrid environments. \n

  • Integrate and manage data across enterprise platforms (e.g., Palantir Foundry, AWS, Azure, GCP); process high-volume data using distributed frameworks (Apache Spark, Flink). \n

\n

Thought Leadership & Innovation

\n \n

  • Analyze state-of-the-art technologies, drive cutting-edge AI strategies and architectures, identify emerging trends, gaps, and innovation opportunities. \n

  • Promote thought leadership through publications, conference presentations, and industry collaboration. \n

  • Contribute ideas that support growth and new business opportunities. \n

\n

\n, n

  • DevSecOps in regulated or safety-critical environments; experience leading technical architecture discussions \n

  • LLM/SLM fine-tuning using PyTorch, TensorFlow, Hugging Face Transformers (LoRA, QLoRA); MLOps practices: model drift detection, retraining pipelines, deployment monitoring. \n

  • Real-time STT/TTS, NLP, and streaming voice (Whisper, WhisperX, faster-whisper, Wispr Flow, Google Cloud Speech, Azure Speech) with custom models, accent/ATC phraseology adaptation; real-time inference and AI agents/agentic architectures. \n

  • Speech technologies: STT/TTS systems (Whisper, Google Cloud Speech, Azure Speech Services), custom voice models, accent adaptation. \n

  • Probabilistic modeling and Bayesian inference (pgmpy, PyMC, Pyro, Stan); causal inference and graphical models applied to safety precursor analysis. \n

  • Vector databases (Pinecone, Weaviate, ChromaDB, FAISS, pgvector); API integrations (RESTful/GraphQL) and streaming platforms (Kafka, Kinesis, Pulsar). \n

  • Containerization (Docker, Kubernetes), infrastructure-as-code (Terraform, CloudFormation); data visualization \n

  • Computer vision frameworks (PyTorch Vision, OpenCV, Detectron2, Ultralytics YOLO, SAM) and multimodal models (CLIP, LLaVA, GPT-4o Vision) for surveillance, surface, and document imagery.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

5:30 min

Extending testing workflows using popular pytest plugins

Florian Bruhin · World Congress 2021

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · World Congress 2021

1:11 min

Deploying and running Airflow in cloud environments

Alan Mazankiewicz · LIVE

Videos

See all

Related articles

See all