Data Scientist - (Mid Level) US

CORNERSTONE GLOBAL PARTNERS COLUMBUS, INC.
United States
5 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$160,000.0
Working hours
Shift work
Job source

Tech stack

Geographic Information Systems Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Microsoft Azure Continuous Integration Customer Data Management Data Validation Extract Transform Load (ETL) Data Security Database Queries
+27 more
Amazon DynamoDB Github Python (Programming Language) Key Management Machine Learning Natural Language Processing PostGIS Role-Based Access Control Power BI Data Streaming Feature Engineering Retrieval-Augmented Generation Large Language Models Prompt Engineering Apache Spark Generative AI Gitlab Git Data Lakes Pyspark Real Time Data Bitbucket Machine Learning Operations Drift Detection Evaluation of Large Language Models Key Vault Databricks

Job description

Our client is launching a new project built on a unified, governed Databricks Lakehouse to power cross-product insights and customer-facing data products. They’re hiring two Data Scientists to join a new, US-based team of Data Scientists and Data Engineers. This is a greenfield project: there’s nothing to inherit, and you’ll help build ML and GenAI capabilities from the ground up. A core near-term focus is AI-assisted ETL: using AI to automatically map customer data from unfamiliar source schemas into a standard target schema and automate customer data onboarding. Industry background is not required; you’ll learn the industry context on the job., * Build ML and GenAI capabilities for the new platform, working primarily in Databricks.

  • Develop AI-assisted ETL that maps unfamiliar customer source schemas to a standard target data schema and automates data onboarding.
  • Work with Data Engineers on data modeling, feature engineering, and turning data into measurable customer value.
  • Explore, prototype, evaluate, and productionize ML and GenAI solutions, including forecasting, anomaly detection, NLP, RAG, LLM-powered assistants, and predictive analytics.
  • Contribute to medallion architecture pipelines (Bronze, Silver, Gold), including data quality checks and data contracts.
  • Package and manage models using Unity Catalog, and design batch and streaming inference where appropriate.
  • Move successful experiments to production with clear SLAs, monitoring, documentation, and runbooks.
  • Build workflows as code using Databricks Asset Bundles, with CI/CD through GitHub Actions.
  • Implement secure data and ML architectures, including RBAC/ABAC, PII masking, and secrets management.
  • Coordinate with counterparts on an international team based in India.

Requirements

  • 3-6 years of experience in Data Science, Machine Learning, or ML Engineering, with a track record of taking models into production.
  • Hands-on Databricks experience: Delta Lake, Unity Catalog, Databricks SQL, Jobs and Workflows, medallion architecture, and Lakehouse.
  • Python and Spark/PySpark.
  • Strong SQL skills.
  • Hands-on GenAI/LLM experience at any level: prompt engineering, RAG, vector stores, LLM evaluation, and guardrails.
  • Strong ML fundamentals: feature engineering, model training and evaluation, monitoring, and drift detection.
  • Some hands-on CI/CD experience for data or ML workloads.
  • Git-based development (GitHub preferred; GitLab or Bitbucket also acceptable).
  • Understanding of data security and compliance, including PII handling and access controls.
  • Based in the United States, with availability to overlap with India working hours between 8:00 and 11:00 AM Eastern Time.
  • Strong communication skills and the ability to write clear technical documentation.

Nice to Have

  • Microsoft Azure (ADLS, Entra ID, Key Vault, Fabric, Power BI).
  • Experience with schema mapping, entity matching, or automated data onboarding.
  • MLflow and Unity Catalog Model Serving.
  • AWS (S3, KMS, Secrets Manager, RDS, DynamoDB).
  • Geospatial data and analytics (PostGIS, spatial joins, GIS-based features).
  • Streaming and real-time data (Structured Streaming, CDC).
  • Experience in utilities, energy, infrastructure, or related industries.

Benefits & conditions

  • Monday to Friday, standard business hours, with overlap between 8:00 and 11:00 AM Eastern Time. East Coast is ideal; West Coast works for early starters.
  • Direct hire, full-time. Laptop provided.
  • Compensation: $160,000 USD base salary plus a 5% bonus., * Competitive Salary: A competitive compensation package based on experience and qualifications.
  • Medical, Dental, and Vision Insurance: Comprehensive insurance coverage to support you and your family.
  • 401(k) Plan with Company Match.
  • Generous Paid Time Off (PTO): Time off to support work-life balance and personal needs.
  • Company-Paid Holidays: Paid holidays throughout the year.
  • Flexible Work Options: Work-from-home opportunities are available, depending on role and business needs.

As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.

About the company

Our client is a leading provider of cloud-based SaaS software that helps energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, they serve customers across North America and continue to expand their platform with new data-driven and AI-powered capabilities.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Loading talks and stories from around this role…