Data Scientist
Role details
Job location
Tech stack
Job description
Leidos is seeking a Data Scientist to support the customer in their program that automates processing of large forensic images, extract and enrich metadata, and display resulting information in meaningful ways for analysts to conduct their assessments. This candidate will be working with a Systems Engineer to complete this work.
This team provides technical solutions and capabilities that enable a cadre of analysts to make critical assessments, including a platform to conduct enterprise search, digital forensics, and data analytics.
This Data Scientist will be building a modern cloud-based system to replace a legacy standalone system and requires an infrastructure team that can work in a quick-paced, dynamic, agile software development environment.
You will adhere to an Agile Scrum development methodology best practices and has 2 week sprint cycles.
Your team will work with a variety of individuals, including key stakeholders and other development teams. However, the customer's project manager will manage priorities.
The team will be responsible for the data stream to data storage solutions. Focus on design, implementation, and operation of data management systems and will design how data is stored, consumed, integrates, and managed by different data entities and digital systems. They will plan, design, and optimize for data throughput and query performance issues and run security scans on existing and new web applications.
Requirements
-
Must have an active TS/SCI with poly to be considered
-
Must have a Masters and at least 15 years of experience OR a doctorate and at least 13 years of experience OR a Bachelors and 18 years of experience.
-
Experience working with an AWS cloud environment.
-
Experience understanding and implementing system security requirements.
-
Experience creating a testing environment to identify and improve bugs and efficiencies.
-
Experience programming in Python, PowerShell, and Java 8+.
-
Experience using Python and the PySpark library to read, write, and manipulate large structured and semi-structured datasets.
-
Experience using task tracking and version control technology to include JIRA and GitHub.
-
Experience using SQL and database technologies such as MySQL or SQL server.
-
Experience using and managing ElasticSearch clusters.
-
Experience tuning and optimizing ElasticSearch.
-
Experience creating and managing ElasticSearch indices.
-
Experience with Apache Spark and managing Spark clusters.
-
Experience tuning and optimizing Apache Spark.
-
Experience with the Databricks platform and Databricks API & CLI.
-
Experience understanding Hadoop Distributed File System (HDFS).
-
Experience understanding Databricks File System (DBFS).
-
Experience with infrastructure as code (IaC) technologies including AWS Cloud Formation.
-
Experience using automated build tools such as Jenkins.
-
Experience implementing machine learning models on text and multimedia data such as Spark NLP and Computer Vision models.
-
Experience with Linux operating systems and shell scripting such as Bash.
-
Experience developing machine learning models on text data and structured datasets.
-
Bachelor's degree in Engineering, Computer Science, Mathematics, Data Science, Statistics, or related field or equivalent work experience.
-
Demonstrated professional experience with cloud technology networks, digital forensics, or cybersecurity topics.
-
Demonstrated experience using artificial intelligence and machine learning technologies/packages.
-
Demonstrated experience using data pipelines and workflow technologies.