Data Scientist

CPMC, LLC.
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Build Automation Automation of Tests Profiling Computer Programming Continuous Integration Software Debugging DevOps Distributed Systems Statistical Hypothesis Testing Python (Programming Language)
+9 more
Scientific Computating Software Engineering Computational Statistics Data Processing Apache Spark Containerization Kubernetes Data Pipelines Docker

Job description

CPMC is seeking a Data Scientist who thrives on technically demanding problems involving large-scale computation, complex algorithms, and high-performance data processing. You will engineer the U.S. Census Bureau’s Disclosure Avoidance System (DAS), a system that executes advanced statistical and differential privacy algorithms across massive datasets.

This role is ideal for an engineer who enjoys:

  • Understanding how algorithms behave under realworld scale
  • Turning research prototypes into robust, high-performance systems
  • Diagnosing subtle numerical, performance, or correctness issues
  • Building distributed systems that must be reproducible, efficient, and scientifically trustworthy

You’ll be building the computational machinery that makes cutting-edge statistical methods run reliably at national scale., * Engineer productiongrade implementations of complex statistical and differential privacy algorithms, ensuring correctness, stability, and performance

  • Translate research code (Python/R) into optimized, maintainable systems, often requiring algorithmic insight and careful handling of numerical edge cases
  • Design and optimize largescale data processing pipelines for ingestion, transformation, validation, and output generation
  • Profile, benchmark, and optimize distributed workloads (Spark, EMR, containerized compute) to reduce runtime and cost
  • Diagnose algorithmic performance issues-from data skew to solver behavior to memory pressure
  • Collaborate deeply with statisticians to understand algorithmic assumptions, constraints, and expected behavior under scale
  • Develop reproducible experiment frameworks, including parameter tracking, environment isolation, and deterministic execution
  • Build automation and tooling that enable researchers to run large experiments safely and efficiently
  • Tune compute and solver configurations (Spark, Gurobi, storage layouts, partitioning strategies) for largescale statistical workloads
  • Support distributed execution environments and contribute to DevOps/automation where needed to keep the system reliable

Requirements

Do you have experience in Statistics?, Experience: 10+ years software engineering; 5+ years Python/R. Would consider # of years experience with PHD in a quantitative discipline., * 10+ years professional software engineering experience

  • Strong programming skills in Python (primary) and familiarity with R
  • Experience with distributed computing (Spark, EMR, or equivalent)
  • Strong background in performance engineering, profiling, and debugging complex systems
  • Experience building and maintaining largescale data pipelines
  • Handson experience with AWS (EMR, S3)
  • Experience with CI/CD, automated testing, and environment management
  • Familiarity with basic probability and statistics concepts (e.g. hypothesis testing, probability distributions, least squares, etc.)”
  • Ability to read, reason about, and improve scientific or researchoriented code, * Experience collaborating with statisticians or working in scientific computing environments
  • Familiarity with numerical methods, statistical computing, or algorithmic evaluation
  • Experience with optimization solvers (e.g., Gurobi) or largescale simulations
  • Knowledge of differential privacy or privacypreserving computation
  • Experience with containerization (Docker, Kubernetes)
  • Experience with HPC or large distributed systems, * You collaborate with researchers pushing the boundaries of statistical privacy
  • You solve problems where the bottleneck might be a numerical instability, a distributed shuffle, a solver configuration, a data partitioning strategy, or an algorithmic assumption that breaks at scale
  • You directly influence the performance and reliability of a system that protects the confidentiality of Census data, * Organizational Skills: Can plan and prioritize work. Follows tasks to their logical conclusion and makes sure that everything has been done to the right standard. Good attention to detail.
  • Team Work: Able to enthuse and maintain project interest. Comfortable working both individually and as part of a team. Prepared to challenge ideas within a group in a constructive way.
  • Communications: Ability to communicate clearly and efficiently to team members and clients, verbally and in writing. Able to present ideas in a variety of ways depending upon audience and context. Excellent active listening skills.
  • Problem Solving: Natural inclination for planning strategy and tactics. Ability to analyze problems and determine root cause, generating alternatives, evaluating and selecting alternatives and implementing solutions.
  • Results oriented: Able to drive things forward regardless of personal interest in the task.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

Videos

See all

Related articles

See all