> Markdown version of [/jobs/ext/171666-data-scientist](https://www.wearedevelopers.com/jobs/ext/171666-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** CPMC, LLC. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Build Automation, Automation of Tests, Profiling, Computer Programming, Continuous Integration, Software Debugging, DevOps, Distributed Systems, Statistical Hypothesis Testing, Python (Programming Language), Scientific Computating, Software Engineering, Computational Statistics, Data Processing, Apache Spark, Containerization, Kubernetes, Data Pipelines, Docker - **Published:** May 20, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a8b83278154c4cae ## About the Role Do you have experience in Statistics?, Experience: 10+ years software engineering; 5+ years Python/R. Would consider # of years experience with PHD in a quantitative discipline., * 10+ years professional software engineering experience * Strong programming skills in Python (primary) and familiarity with R * Experience with distributed computing (Spark, EMR, or equivalent) * Strong background in performance engineering, profiling, and debugging complex systems * Experience building and maintaining largescale data pipelines * Handson experience with AWS (EMR, S3) * Experience with CI/CD, automated testing, and environment management * Familiarity with basic probability and statistics concepts (e.g. hypothesis testing, probability distributions, least squares, etc.)" * Ability to read, reason about, and improve scientific or researchoriented code, * Experience collaborating with statisticians or working in scientific computing environments * Familiarity with numerical methods, statistical computing, or algorithmic evaluation * Experience with optimization solvers (e.g., Gurobi) or largescale simulations * Knowledge of differential privacy or privacypreserving computation * Experience with containerization (Docker, Kubernetes) * Experience with HPC or large distributed systems, * You collaborate with researchers pushing the boundaries of statistical privacy * You solve problems where the bottleneck might be a numerical instability, a distributed shuffle, a solver configuration, a data partitioning strategy, or an algorithmic assumption that breaks at scale * You directly influence the performance and reliability of a system that protects the confidentiality of Census data, * Organizational Skills: Can plan and prioritize work. Follows tasks to their logical conclusion and makes sure that everything has been done to the right standard. Good attention to detail. * Team Work: Able to enthuse and maintain project interest. Comfortable working both individually and as part of a team. Prepared to challenge ideas within a group in a constructive way. * Communications: Ability to communicate clearly and efficiently to team members and clients, verbally and in writing. Able to present ideas in a variety of ways depending upon audience and context. Excellent active listening skills. * Problem Solving: Natural inclination for planning strategy and tactics. Ability to analyze problems and determine root cause, generating alternatives, evaluating and selecting alternatives and implementing solutions. * Results oriented: Able to drive things forward regardless of personal interest in the task. ## Description CPMC is seeking a Data Scientist who thrives on technically demanding problems involving large-scale computation, complex algorithms, and high-performance data processing. You will engineer the U.S. Census Bureau's Disclosure Avoidance System (DAS), a system that executes advanced statistical and differential privacy algorithms across massive datasets. This role is ideal for an engineer who enjoys: * Understanding how algorithms behave under realworld scale * Turning research prototypes into robust, high-performance systems * Diagnosing subtle numerical, performance, or correctness issues * Building distributed systems that must be reproducible, efficient, and scientifically trustworthy You'll be building the computational machinery that makes cutting-edge statistical methods run reliably at national scale., * Engineer productiongrade implementations of complex statistical and differential privacy algorithms, ensuring correctness, stability, and performance * Translate research code (Python/R) into optimized, maintainable systems, often requiring algorithmic insight and careful handling of numerical edge cases * Design and optimize largescale data processing pipelines for ingestion, transformation, validation, and output generation * Profile, benchmark, and optimize distributed workloads (Spark, EMR, containerized compute) to reduce runtime and cost * Diagnose algorithmic performance issues-from data skew to solver behavior to memory pressure * Collaborate deeply with statisticians to understand algorithmic assumptions, constraints, and expected behavior under scale * Develop reproducible experiment frameworks, including parameter tracking, environment isolation, and deterministic execution * Build automation and tooling that enable researchers to run large experiments safely and efficiently * Tune compute and solver configurations (Spark, Gurobi, storage layouts, partitioning strategies) for largescale statistical workloads * Support distributed execution environments and contribute to DevOps/automation where needed to keep the system reliable ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)