Senior Data Scientist

DATA MICRO TECHNOLOGIES INC.
United States
13 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Application Programming Interfaces (APIs) Artificial Intelligence Airflow Data Analysis Big Data Software Quality Continuous Integration Data Validation Information Engineering Data Infrastructure Extract Transform Load (ETL)
+34 more
Django Web Framework Python (Programming Language) KNIME Machine Learning NumPy SQL Databases Scripting ReactJS Flask (Web Framework) Large Language Models Prompt Engineering Apache Spark Model Validation Change Data Capture Backend Fastapi Pandas Event Driven Architecture Containerization Scikit Learn Apache Flink Dask Low-code Apache Kafka Spark Streaming Data Management Celery Front End Software Development Dataiku Api Design Stream Processing Data Pipelines Docker Alteryx

Job description

We are looking for a Senior Data Scientist who can lead the end-to-end delivery of data science, analytics, and data engineering projects for our clients. You will guide a team of junior data scientists and analysts, translating client problems into working solutions built on our suite of data platforms - and rolling up your sleeves to write Python when the tooling alone isn’t enough.

This is a hands-on leadership role. You will spend part of your time mentoring and reviewing your team’s work, part of it engaging clients and managing delivery, and part of it in the weeds: designing pipelines, building models, and stitching platform components together with custom code.

What You’ll Do

Project & Technical Delivery

  • Lead the execution of client-facing data science, analytics, and data engineering projects from scoping through deployment and handover.
  • Design analytical approaches and solution architectures using our low-code data science studio, visual ETL platform, and enterprise lakehouse (with change-data-capture capabilities).
  • Write Python to extend, integrate, or bridge these platforms - custom transformations, model logic, APIs, automation, and anything the visual tools can’t handle out of the box.
  • Build and operate the harder engineering layers of a solution when projects demand it: orchestrated pipelines, worker-based and event-driven architectures, stream processing, and scaling for data volume and concurrency.
  • Own solution quality: data validation, model evaluation, reproducibility, performance, and operational reliability (monitoring, alerting, failure recovery).

Team Leadership

  • Lead and mentor a team of junior data scientists and analysts; review their work, unblock them technically, and grow their skills.
  • Break projects down into workable tasks, assign them across the team, and keep delivery on track.
  • Set and uphold standards for code quality, documentation, and analytical rigor.

Stakeholder & Client Management

  • Serve as the primary technical point of contact for client stakeholders during project delivery.
  • Translate ambiguous business problems into well-defined analytical work, and communicate findings and trade-offs clearly to non-technical audiences.
  • Manage scope, timelines, and expectations; escalate risks early and propose options.

Internal Collaboration

  • Work alongside our internal product engineering team (Python/FastAPI backend, React frontend) to feed field learnings back into the product and, occasionally, contribute to backend integrations.

Requirements

  • 5+ years of experience in data science and analytics, with at least 1-2 years leading or mentoring junior team members.
  • Strong, practical Python skills: pandas/NumPy, scikit-learn, data pipeline scripting, working with APIs, and writing maintainable code others can build on. Strong familiarity with Python backend frameworks (either or a combination of FastAPI, Django, Flask)
  • Solid grounding in statistics, machine learning, and analytical problem-solving - you can choose the right method for the problem, not just the fashionable one.
  • Proficiency in SQL and comfort working with large datasets in modern data platforms (lakehouse architectures, columnar stores, CDC-based ingestion).
  • Hands-on experience building and operating data orchestration pipelines (e.g., Airflow, Dagster, Prefect, or equivalent) - scheduling, dependency management, retries, backfills, and monitoring.
  • Working knowledge of distributed compute frameworks (e.g., Spark, Dask) and when to reach for them versus simpler approaches.
  • Demonstrated project delivery experience: scoping, planning, managing stakeholders, and shipping on time.
  • Excellent communication skills, both with technical teams and business stakeholders.

Nice-to-haves

  • Exposure to AI engineering: LLM applications, RAG pipelines, prompt engineering, or model serving. This is a very strong nice-to-have for us, as we are also going big into developing AI applications and data platforms.
  • Experience designing beyond single-machine batch jobs: distributed workers and task queues (e.g., Celery, RQ), stream processing (e.g., Kafka, Flink, Spark Structured Streaming), and scaling pipelines for throughput and reliability.
  • Experience with low-code/visual data science or ETL tools (e.g., Dataiku, Alteryx, KNIME, or similar) and a clear sense of when to use them versus custom code.
  • React/frontend literacy.
  • Experience with containerized deployments (Docker) and CI/CD practices.
  • Consulting or client-services background.

What We Expect

  • Within 3 months: take ownership of at least one active client project, know our platform stack well enough to design solutions on it, and have established a working rhythm with the team.
  • Within 6 months: Independently leading multiple concurrent projects, juniors under you are visibly leveling up, and clients ask for you by name.
  • Within 12 months: you’ve shaped how we deliver data science projects - reusable patterns, better standards, faster ramp-up for new team members.

About the company

DataMicron is a multi-award Big Data and AI company, with patented technologies in USA and China. We have extensive experience in developing products and full-scale solutions in the field of business intelligence, machine learning & AI, data engineering and front-end development. We have covered and implemented end-to-end Big Data Analytics solutions in eight countries across major public and private sector organizations, and are currently going through a large expansion in our business portfolio.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Scaling performance across multiple GPUs using specialized frameworks

Paul Graham Paul Graham · World Congress 2025

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

2:31 min

Collecting vectors into linear algebraic matrices and structures

Jodie Burchell · LIVE

Videos

See all

Related articles

See all