> Markdown version of [/jobs/ext/2719430-data-scientist-in-arlington](https://www.wearedevelopers.com/jobs/ext/2719430-data-scientist-in-arlington). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist in Arlington - **Company:** Energy Jobline - **Location:** Arlington, VA, United States - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Amazon Web Services, Business Analytics Applications, Microsoft Azure, Big Data, Continuous Integration, Web Scraping, Extract Transform Load (ETL), Data Security, Relational Databases, Github, R (Programming Language), Java Database Connectivity, Python (Programming Language), Machine Learning, Natural Language Processing, Open Database Connectivity, Azure Data Lake, SQL Databases, Systems Integration, Text Mining, Unstructured Data, Azure Data Factory, Data Lakes, Pyspark, Restful APIs, Software Version Control, Data Pipelines, Databricks - **Published:** September 4, 2026 - **Apply:** https://www.energyjobline.com/job/data-scientist-arlington-31447211 ## About the Role Required * US Citizenship required * 1+ year of experience working with AWS and/or Azure services, such as Databricks, Azure Data Factory, and Azure Data Lake. * Experience with Azure DevOps and/or GitHub, CI/CD pipelines, and Agile methodologies. * Experience developing and optimizing automated data pipelines and ETL/ELT processes. * Proficiency with Python and SQL; experience with R and/or JavaScript is a plus. * Experience working with structured and unstructured data and integrating multiple data sources. * Knowledge of machine learning, artificial intelligence, NLP, predictive analytics, statistical analysis, or data/text mining. * Strong analytical, problem-solving, technical writing, and communication skills. * Ability to work independently and effectively within a collaborative team environment. * Experience with contract, procurement, and accounts payable (AP) data to identify trends, anomalies, and potential risks. * Experience using PySpark/PySpark SQL for large-scale data processing. * Experience with data lakes, data lakehouses, relational databases, and business intelligence tools. * Experience supporting federal audit, investigation, oversight, fraud, waste, or abuse-related activities. * Experience developing predictive models, NLP solutions, dashboards, and investigative or analytical visualizations. * Experience presenting technical findings, training materials, or conference presentations to large audiences. ## Description * Develop, manage, and optimize automated data pipelines to support reliable analytics and data quality. * Extract, clean, transform, normalize, and validate structured and unstructured data. * Integrate data from flat files, relational databases, APIs, external systems, and other sources using JDBC/ODBC, REST APIs, and web scraping. * Develop data and analytics solutions using AWS and/or Azure, including Databricks, Azure Data Factory, and Azure Data Lake. * Use Azure DevOps and/or GitHub to support source control, CI/CD pipelines, and Agile development. * Apply Python, SQL, R, and/or JavaScript to develop datasets, models, dashboards, visualizations, and reports. * Apply advanced analytics including machine learning, AI, NLP, predictive analytics, statistical modeling, and data/text mining. * Develop, train, evaluate, deploy, and maintain machine learning and AI models. * Use PySpark/PySpark SQL to process and analyze large-scale datasets. * Identify trends, anomalies, patterns, and risks to support audit, investigative, oversight, and fraud, waste, and abuse activities. * Translate technical findings into clear narratives, recommendations, visualizations, and presentations. * Develop technical documentation, requirements, test plans, methodologies, and training materials. * Collaborate with technical teams, stakeholders, and management and provide guidance on data access, quality, storage, and analytics. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)