> Markdown version of [/jobs/ext/1975731-data-engineer-data-architecture-for-data-science-machine-learning](https://www.wearedevelopers.com/jobs/ext/1975731-data-engineer-data-architecture-for-data-science-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - Data Architecture for Data Science & Machine Learning - **Company:** Pennsylvania State University - **Location:** Reston, VA, United States (Remote available) - **Experience:** Expert - **Salary:** $146,040.0 - $318,840.0 - **Contract:** Permanent contract - **Skills:** Airflow, Algorithm Design, BigQuery, Cloud Database, Cyber Security, Databases, Data Architecture, Information Engineering, Data Integrity, Extract Transform Load (ETL), Data Security, Amazon DynamoDB, Information Lifecycle Management, Python (Programming Language), PostgreSQL, SQL Azure, MongoDB, Neo4j, NoSQL, Performance Tuning, Query Optimization, Redis, SQL Databases, Unstructured Data, Scripting, Data Storage Technologies, Snowflake, Indexer, Amazon Relational Database Service, Data Lakes, Information Technology, Cassandra, Non-relational Database, Machine Learning Operations, Data Pipelines, Amazon Redshift - **Published:** August 7, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87968768/1 ## About the Role * 5+ years of experience in data engineering, database architecture, or related technical roles * Expert-level proficiency in PostgreSQL (query tuning, schema design, indexing, partitioning, replication) * Strong understanding of data modeling, normalization vs. denormalization tradeoffs, and query optimization * Experience with non-relational databases (e.g., MongoDB, Cassandra, Neo4j, Redis, or DynamoDB) * Familiarity with machine learning workflows and how data is consumed for training, evaluation, and deployment * Experience with cloud database services (AWS RDS/Aurora, GCP Cloud SQL, Azure Database) * Proficiency in SQL and one or more scripting languages (Python preferred) * Excellent communication and collaboration skills-comfortable working closely with data scientists, ML engineers, and software developers Preferred skills/experience includes: * Experience architecting hybrid data ecosystems spanning relational, NoSQL, and analytical databases. * Knowledge of data lake, warehouse, and feature store architectures (e.g., Snowflake, Redshift, BigQuery, Feast) * Familiarity with ETL/ELT frameworks and data orchestration tools (e.g., Airflow, dbt) * Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field Work location can be fully on-site located in State College, PA or Reston , VA. Questions related to flexible work should be directed to the hiring manager during the interview process. MINIMUM EDUCATION, WORK EXPERIENCE & REQUIRED CERTIFICATIONS If filled as R&D Engineer - Data Science (ARL) - Principal Professional, this position requires: Bachelor's Degree - Engineering or Science 19+ years of relevant experience Required Certifications: None If filled as R&D Engineer - Data Science (ARL) - Advanced Professional, this position requires: Bachelor's Degree - Engineering or Science 5+ years of relevant experience Required Certifications: None If filled as R&D Engineer - Data Science (ARL) - Senior Professional, this position requires: Bachelor's Degree - Engineering or Science 14+ years of relevant experience Required Certifications: None ## Description The ideal candidate is a PostgreSQL expert who also brings hands-on experience with other modern data storage technologies-such as NoSQL, graph, and time-series databases-and can guide the organization in choosing the right tools and structures for each data use case., * Design and maintain scalable, high-performance database solutions to support data science workflows and ML experimentation * Partner with data scientists to understand data access patterns and develop storage strategies that accelerate analysis and model training * Serve as the internal subject matter expert on PostgreSQL-including schema design, indexing, partitioning, and query optimization * Evaluate and integrate alternative database technologies (e.g., MongoDB, Neo4j, Redis, Cassandra) where they provide clear advantages * Lead efforts to optimize data pipelines for both structured and unstructured data used in algorithm development * Ensure data integrity, security, and governance across storage systems * Implement monitoring, automation, and performance-tuning tools for all database environments * Advise on data lifecycle management-balancing accessibility for R&D with efficiency and compliance requirements, ARL's purpose is to research and develop innovative solutions to challenging scientific, engineering, and technology problems in support of the Navy, the Intel Community (IC), and other federal government customers. FOR FURTHER INFORMATION on ARL, visit our website atwww.arl.psu.edu. BACKGROUND CHECKS/CLEARANCES Employment with the University will require successful completion of background check(s) in accordance with University policies. Notice regarding employment at the Applied Research Laboratory (ARL): Employees must be eligible to obtain a government security clearance, participate in the ARL drug testing program, and comply with electronic and physical monitoring requirements applicable to federal contractors. ARL operates in a secure information environment involving Unclassified, Controlled Unclassified Information (CUI), and Classified information. Personal electronic devices brought onsite must be registered and may be restricted from certain areas. You must be a U.S. citizen to apply. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)