Data Science ML/Gen AI Engineer (Mid-Level)

CORNERSTONE GLOBAL PARTNERS COLUMBUS, INC.
United States
5 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Geographic Information Systems Artificial Intelligence Amazon Web Services Data Analysis Microsoft Azure Computer Programming Continuous Integration Data Security Github Python (Programming Language) Machine Learning
+21 more
Natural Language Processing PostGIS Role-Based Access Control Power BI SQL Databases Data Streaming Google Cloud Feature Engineering Large Language Models Prompt Engineering Apache Spark Deep Learning Multi-Cloud Generative AI Change Data Capture Indexer Data Lakes Pyspark Real Time Data Machine Learning Operations Databricks

Job description

Our client is building a modern, multi-cloud, enterprise-grade data estate - a unified Databricks-based platform, larger in scope than a standard lakehouse, that centralizes data across the company’s products and cloud environments (AWS, Azure, and GCP) and feeds multiple product lines, primarily the damage prevention platform. This role will primarily support that platform by integrating data from multiple sources, with occasional work touching other product lines. This is a skill-first role: strong fundamentals in statistics, machine learning, and deep learning, combined with genuinely practical GenAI/agentic-systems experience (LangChain, agent workflows, MCPs), matter more than direct industry-domain background., * Contribute to medallion architecture pipelines (Bronze * Silver * Gold) using Databricks, and support column-level lineage and governance initiatives (targeting at least 95% lineage coverage).

  • Explore, prototype, evaluate, and productionize machine learning and GenAI solutions across use cases including forecasting, anomaly detection, NLP, Retrieval-Augmented Generation (RAG), LLM-powered assistants/copilots, and predictive analytics.
  • Package and manage models using Unity Catalog model management/registries, and design batch and streaming inference architectures where appropriate.
  • Partner with Product and business stakeholders to define success metrics, KPIs, and A/B testing strategies, and move successful experiments from prototype to production with clear SLAs, monitoring, and documentation.
  • Build production workflows, jobs, and notebooks as infrastructure-as-code using Databricks Asset Bundles (DABs), with CI/CD via GitHub Actions.
  • Contribute business metrics, definitions, and semantic models to Unity Catalog, supporting consumption through Power BI and Databricks AI/BI.
  • Implement secure data and ML architectures using RBAC/ABAC within Unity Catalog, and support compliance requirements (SOC 2, ISO 27001, GDPR, PIPEDA).

Requirements

  • 3-6 years of experience in Data Science, Machine Learning, or ML Engineering, with a track record of taking models from development through production.
  • Strong basic understanding of statistics, machine learning, and deep learning fundamentals.
  • Practical, hands-on GenAI/LLM experience: prompt engineering, RAG, vector databases/vector stores, LLM evaluation, agentic workflows, LangChain, and MCPs.
  • Strong programming and data skills in Python, SQL, and Spark/PySpark, with hands-on Databricks experience (Delta Lake, Unity Catalog, DBSQL, Jobs and Workflows, medallion architecture).
  • Bachelor’s degree is sufficient - a master’s in statistics is a plus, as is a well-regarded certification (e.g. DeepLearning.AI, Coursera courses by instructors such as Andrew Ng).
  • Strong communication and collaboration skills.

Nice-to-Have

  • Experience with Microsoft Azure, AWS, or Google Cloud Platform (GCP).
  • Experience with geospatial data and analytics (PostGIS, spatial joins, spatial indexing, GIS-based feature engineering).
  • Experience with streaming and real-time data, including Structured Streaming and Change Data Capture (CDC).
  • Hands-on experience with MLflow and Unity Catalog Model Serving.
  • Experience working in utilities, energy, infrastructure, or public works industries.

Benefits & conditions

  • Be part of a dynamic and growing company that is well-respected in its industry.
  • Competitive compensation based on experience and qualifications.
  • Health Insurance coverage.

As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.

About the company

Our client is a growing SaaS company in the damage prevention, asset integrity, and infrastructure-protection space, serving energy, utility, telecom, and infrastructure companies across North America.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all