Data Scientist
Amatriot Group, LLC
Woodlawn, MD, United States
5 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$130,000.0 - $180,000.0
Working hours
Regular working hours
Job source
Tech stack
Business Analytics Applications
Business Logic
Big Data
Software Quality
Code Review
Encodings
Continuous Integration
Data Validation
Data Cleansing
Data Discovery
Data Integrity
Data Security
+31 more
IBM DB2
Database Queries
Distributed Systems
Text Processing
Apache Hadoop
Information Extraction
Information Retrieval
Information Sciences
Python (Programming Language)
PostgreSQL
Machine Learning
Microsoft SQL Server
Natural Language Processing
Named Entity Recognition
Oracle (Applications)
Pattern Recognition
Performance Tuning
Regular Expressions
SQL Databases
Enterprise Data Management
Privacy Controls
Data Processing
Freeform SQL
Sql Optimization
Indexer
Scikit Learn
Information Technology
Spacy
Software Version Control
Data Pipelines
Jenkins
Job description
Develop and maintain advanced data processing and entity resolution solutions using Python and SQL on enterprise data platforms. Support data cleansing, validation, optimization, testing, deployment, and monitoring while maintaining data integrity, security, reproducibility, and privacy standards., * Design, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms.
- Support end-to-end delivery from data validation and testing through deployment and post-implementation monitoring.
Data Hygiene and Management
- Clean, transform, and manage large-scale datasets from diverse and complex sources.
- Ensure data integrity, reliability, and security.
Performance Optimization
- Optimize complex SQL queries and database operations to support efficient data access, processing, and scalability.
Engineering Standards
- Participate in code reviews and enforce version control.
- Uphold best practices for code quality, reproducibility, data privacy, and security.
Requirements
- Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with experience in NLP, Text Processing, Information Extraction, Python, SQL, Regex, and specialized libraries/frameworks. [Required]
- Master’s degree with 10+ years of experience, Bachelor’s degree with 12+ years of experience, or 18+ years of experience in lieu of a degree. [Required]
Experience
- 10+ years of overall experience in the IT industry. [Required]
- Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction. [Required]
- Practical knowledge of Named Entity Recognition and Address Standardization for extracting and cleaning unstructured text data. [Required]
- Deep understanding of data matching strategies, including Blocking and Indexing, String Distance Metrics, and Phonetic Encoding. [Required]
- Experience applying TF-IDF and Cosine Similarity for text comparisons and information retrieval. [Required], * Strong Python development skills for building analytics solutions and manipulating data. [Required]
- Advanced SQL proficiency for complex data querying, optimization, and database operations. [Required]
- Practical experience using Regex for advanced text processing, data cleansing, and pattern matching. [Required]
- Familiarity with specialized libraries and frameworks, including:
- Linkage libraries such as Splink/FastLink, Dedupe, or recordlinkage.
- Core Python data science libraries, specifically spaCy for NLP tasks and Scikit-Learn for general machine learning and clustering.
- Familiarity with code reviews, version control, and maintaining data security and reproducibility standards. [Required]
- Excellent communication skills. [Required]
Clearance
- Ability to obtain and maintain a public trust clearance. [Required]
Other Requirements
- Willingness to work on-site at SSA HQ in Woodlawn, MD, 5 days a week. [Required], * Prior experience delivering IT or data initiatives within federal, state, or local government environments.
- Experience retrieving, migrating, and manipulating data from legacy and distributed systems, including PostgreSQL, DB2, Oracle, SQL Server, and Hadoop, as well as unstructured flat files.
- Experience utilizing Jenkins to automate continuous integration, testing, and deployment (CI/CD) for data validation pipelines.
- Experience with pipeline automation tools to schedule and monitor complex data cleansing jobs.
Skills
- Ability to operate independently, take ownership of data pipeline architectures, and drive projects from data discovery through post-implementation.
- Strong ability to translate complex algorithmic decisions, such as probabilistic match thresholds, into clear business logic for executive leadership and non-technical stakeholders.
- Excellent problem-solving skills and proven verbal and written communication skills when collaborating across cross-functional teams.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
about 3 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
about 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
DC
Daniel Cranney
Dev Digest 166: Sycophancy, Zip bombs and AI Native Development
over 1 year ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago