Data Scientist
BestPegasus LLC
Woodlawn, MD, United States
4 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
HTML
Artificial Intelligence
Airflow
Amazon Web Services
Amazon Elastic Compute Cloud
Data Analysis
Cloudera Impala
Nvidia CUDA
Databases
Information Engineering
Data Systems
IBM DB2
+40 more
DevOps
Distributed Computing Environment
Markup Languages
Apache Hadoop
Apache Hive
Information Sciences
Python (Programming Language)
LaTeX
PostgreSQL
Machine Learning
Mathematica
Microsoft SQL Server
MySQL
Natural Language Processing
NLTK (NLP Analysis)
NumPy
Oracle (Applications)
Regular Expressions
Tensorflow
Azure Machine Learning
Search Technologies
SQLite
SQL Databases
Transaction Data
Data Processing
Pytorch
Large Language Models
Apache Spark
Parallel Computation
Generative AI
Gpu Programming
Git
Pandas
Matplotlib
Scikit Learn
Kubernetes
Information Technology
Machine Learning Operations
Spacy
Software Version Control
Job description
- Hands on experience in Python, NLP frameworks, SQL, Pandas, NLTK, SPACy and LLMs
- Well versed in SQL and analyzing trends and transactional data.
- Understand real world challenges and develop automated data solutions
- Develop, test, and deploy new techniques for NLP understanding
- Scalable development/deployment of ML and Generative AI approaches (such as Large Language Models (LLMs)
- Train and optimize NLP/LLM models and create Python based pipelines
- Experience building cloud native solutions on AWS
- Determine the nature of analytic problems, evaluate options, and offer recommendations for resolution.
- Advise on the methods and data needed and/or available to evaluate the (intelligence or data) problem.
- Collaborate with data collectors and analysts to identify and close gaps on complex monitoring problems.
- Provide accurate, timely, complex, and sophisticated data analysis.
Requirements
Note: Selected candidate must be able to obtain and maintain a public trust clearance ACCEPTING locals only - Selected candidate must be willing to work on-site in Woodlawn, MD 5 days a week
Key Required Skills-
- Solid Experience with Natural Language Processing (NLP), Python, NLP frameworks, SQL, Pandas, NLTK and SPACy.
- Experience with Generative AI and Large Language Models (LLM)
- Excellent Communication skills, Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with industry experience on Python, NLP frameworks, SQL, Pandas, NLTK and SPACy, data science, and AI/ML/LLM engineering.
- Overall 10+ years’ experience in IT industry
- Solid Experience with Natural Language Processing (NLP), Python, NLP frameworks, SQL, Pandas, NLTK and SPACy.
- Experience with Generative AI and Large Language Models (LLM)
- Evidence of true self-starter and operating independently.
- Fluency in Python Programming, version control and collaboration with GIT, standard Python packages (ex. Pandas, numpy, matplotlib) and ML frameworks
- Knowledge of TensorFlow, PyTorch, Pandas, scikit-learn, NLTK, Azure ML (optional), Amazon Web Services EC2.
- Experience with scalable data engineering frameworks such as Apache Spark and orchestration frameworks such as Airflow, and/or experience with semantic search.
- Expert knowledge in conducting data analysis and applying advanced statistical concepts and ML methods to build, train, test, and evaluate a variety of supervised and unsupervised analytic models.
- Experience with ML model deployment and operations like DevOps, MLOps, LLMOps.
- Experience with NLP and Generative AI libraries like regular expressions (e.g., spacy, langchain), text annotation tools and semantic frameworks.
- Ability to clean and process large amounts of real-world data.
- Experience retrieving and manipulating data from a variety of data sources included DB2, Oracle, SQL Server, Hadoop and flat files.
- Excellent Communication skills.
- Experience with database management systems (e.g., PostgresSQL, MySQL, SQLite, SQL, etc.)
- Excellent analytical skills to identify potential risks and propose effective solutions.
- Excellent problem-solving skills, ability to collaborate with cross-functional teams and proven communication in written and verbal formats to various audiences to include executive leadership.
- Prior experience with federal or state governments IT projects.
- Industry experience preferred
- Experience with, or the ability and willingness to learn distributed processing via the Hadoop ecosystem, i.e., Spark, Impala and Hive.
- Experience working in an analytical research environment.
- Experience in parallel processing such as GPU programming with CUDA
- Experience with Mathematica
- Experience using markup languages such as LaTeX, HTML, etc.
- Experience with Natural Language Processing for anomaly detection
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
EM
Eli McGarvie
Data Analyst Salary in the UK
about 3 years ago
KD
Krissy Davis
Best Coding Boot Camps in Germany
over 3 years ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago