> Markdown version of [/jobs/ext/3565760-data-scientist-mid-level-us](https://www.wearedevelopers.com/jobs/ext/3565760-data-scientist-mid-level-us). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - (Mid Level) US - **Company:** CORNERSTONE GLOBAL PARTNERS COLUMBUS, INC. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $160,000.0 - **Contract:** Permanent contract - **Skills:** Geographic Information Systems, Artificial Intelligence, Amazon Web Services, Amazon S3, Data Analysis, Microsoft Azure, Continuous Integration, Customer Data Management, Data Validation, Extract Transform Load (ETL), Data Security, Database Queries, Amazon DynamoDB, Github, Python (Programming Language), Key Management, Machine Learning, Natural Language Processing, PostGIS, Role-Based Access Control, Power BI, Data Streaming, Feature Engineering, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Apache Spark, Generative AI, Gitlab, Git, Data Lakes, Pyspark, Real Time Data, Bitbucket, Machine Learning Operations, Drift Detection, Evaluation of Large Language Models, Key Vault, Databricks - **Published:** October 3, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pouroob1ba ## About the Role * 3-6 years of experience in Data Science, Machine Learning, or ML Engineering, with a track record of taking models into production. * Hands-on Databricks experience: Delta Lake, Unity Catalog, Databricks SQL, Jobs and Workflows, medallion architecture, and Lakehouse. * Python and Spark/PySpark. * Strong SQL skills. * Hands-on GenAI/LLM experience at any level: prompt engineering, RAG, vector stores, LLM evaluation, and guardrails. * Strong ML fundamentals: feature engineering, model training and evaluation, monitoring, and drift detection. * Some hands-on CI/CD experience for data or ML workloads. * Git-based development (GitHub preferred; GitLab or Bitbucket also acceptable). * Understanding of data security and compliance, including PII handling and access controls. * Based in the United States, with availability to overlap with India working hours between 8:00 and 11:00 AM Eastern Time. * Strong communication skills and the ability to write clear technical documentation. Nice to Have * Microsoft Azure (ADLS, Entra ID, Key Vault, Fabric, Power BI). * Experience with schema mapping, entity matching, or automated data onboarding. * MLflow and Unity Catalog Model Serving. * AWS (S3, KMS, Secrets Manager, RDS, DynamoDB). * Geospatial data and analytics (PostGIS, spatial joins, GIS-based features). * Streaming and real-time data (Structured Streaming, CDC). * Experience in utilities, energy, infrastructure, or related industries. ## Description Our client is launching a new project built on a unified, governed Databricks Lakehouse to power cross-product insights and customer-facing data products. They're hiring two Data Scientists to join a new, US-based team of Data Scientists and Data Engineers. This is a greenfield project: there's nothing to inherit, and you'll help build ML and GenAI capabilities from the ground up. A core near-term focus is AI-assisted ETL: using AI to automatically map customer data from unfamiliar source schemas into a standard target schema and automate customer data onboarding. Industry background is not required; you'll learn the industry context on the job., * Build ML and GenAI capabilities for the new platform, working primarily in Databricks. * Develop AI-assisted ETL that maps unfamiliar customer source schemas to a standard target data schema and automates data onboarding. * Work with Data Engineers on data modeling, feature engineering, and turning data into measurable customer value. * Explore, prototype, evaluate, and productionize ML and GenAI solutions, including forecasting, anomaly detection, NLP, RAG, LLM-powered assistants, and predictive analytics. * Contribute to medallion architecture pipelines (Bronze, Silver, Gold), including data quality checks and data contracts. * Package and manage models using Unity Catalog, and design batch and streaming inference where appropriate. * Move successful experiments to production with clear SLAs, monitoring, documentation, and runbooks. * Build workflows as code using Databricks Asset Bundles, with CI/CD through GitHub Actions. * Implement secure data and ML architectures, including RBAC/ABAC, PII masking, and secrets management. * Coordinate with counterparts on an international team based in India.