> Markdown version of [/jobs/ext/198167-senior-data-engineer-design-architecture](https://www.wearedevelopers.com/jobs/ext/198167-senior-data-engineer-design-architecture). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer: Design & Architecture - **Company:** K2Share LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Unity 3d, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Big Data, Code Review, Encodings, Cyber Security, Information Engineering, Data Integration, Extract Transform Load (ETL), Data Sharing, Data Systems, Relational Databases, Identity and Access Management, JSON, Python (Programming Language), PostgreSQL, Operational Databases, Power BI, Migration Manager, Data Processing, Cloud Platform System, Microsoft Power Automate, Apache Spark, Mitre Att&ck, Change Data Capture, Indexer, Git, Powerquery, Pandas, Build Management, Data Lakes, Pyspark, Semi-structured Data, Git Flow, Information Technology, AWS Fargate, AWS Data Analytics, Functional Programming, Software Version Control, Serverless Computing, Powerapps, Databricks - **Published:** May 31, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=53c496dc134a28f2 ## About the Role Do you have experience in Version control systems?, Do you have a Bachelor's degree?, You are a senior data engineer who moves comfortably between designing a relational schema for audit-grade compliance data, building the analytical pipelines that turn it into insight, and using AI assistants as a daily part of how you build. You think in terms of system design - not just code - and you bring the discipline to weigh tradeoffs, document design choices, and defend them in front of technical and non-technical reviewers alike. You have direct experience with OSCAL or other structured compliance schemas and you understand why round-trip fidelity matters for data that must stand up to audit. You think about schema decisions and what they will need to support five years out, not just the next sprint. And you treat AI tooling as an accelerator that has already changed the engineering job, not a future trend you are watching from the sidelines., * 5+ years of production data engineering experience, with a track record of designing and owning data systems end-to-end * Strong relational database expertise with PostgreSQL or equivalent - schema design at scale, indexing and partitioning strategy, access control, and patterns for handling semi-structured data including JSON. * Strong system design and architecture instincts - able to translate business and compliance requirements into data system designs, document tradeoffs, and lead design reviews with technical and non-technical stakeholders. * 3+ years of big-data processing on Databricks, Spark, or equivalent - PySpark, Delta Lake, Unity Catalog, and medallion (bronze/silver/gold) architecture patterns * Strong AWS experience including S3, Bedrock, Lambda, Fargate, EC2, relational database services, change-data-capture services, serverless compute, IAM, and KMS - ideally in GovCloud or other regulated-cloud environments * Strong Python development skills, including data manipulation libraries such as pandas, for ETL, transformation, and analytical workflows * Proficient with Git and modern version-control practices - branching strategies, code review discipline, and collaborative workflows in a team setting * Experience working with structured external schemas - OSCAL or similar standards-based data - including the discipline of preserving fidelity through transformation * Demonstrated focus on data system optimization - identifying I/O bottlenecks, remediating performance issues, balancing cost and performance, and scaling compute responsibly * Schema evolution discipline - migration strategy, backward compatibility, change-data-capture-friendly design, and the operational rigor of running production schemas under change control * Experience with async data processing patterns - task queues, message-based pipelines, idempotent task design * Active use of AI development tools as a routine part of the engineering workflow, with informed views on where they accelerate and where they need supervision * Familiarity with responsible AI deployment patterns - RAG architectures, vector databases, embedding management, prompt and output guardrails, and evaluation methods * Working knowledge of federal cybersecurity frameworks: FISMA, NIST RMF, NIST SP 800-53, NIST CSF * Demonstrated ability to interpret regulatory and policy guidance and translate it into product or data-product requirements * Comfort working across ambiguous, fast-moving federal programs with minimal supervision and strong collaborative instincts Preferred Knowledge, Skills, and Abilities * Direct OSCAL experience - schema implementation, document handling, version migration, or extension namespace design * Direct experience supporting U.S. Government federal civilian agency clients * Familiarity with FedRAMP authorization processes, ATO lifecycle management, and 3PAO assessment data needs * Familiarity with the Microsoft Power stack (Power BI, Power Query, Power Apps, Power Automate) for federal client reporting and workflow automation * Experience with Delta Sharing, cross-environment Databricks integration, or modern data-sharing patterns * Familiarity with MITRE ATT&CK, STIX/TAXII, or threat-intelligence data integration * Experience building or operating data products for compliance, audit, or risk-management use cases Education Bachelor's Degree in Computer Science, Data Engineering, Management Information Systems, Cybersecurity, or equivalent from an accredited institution Security Clearance Must hold, or be eligible to obtain and maintain, a Federal Security Clearance or Public Trust ## Description The Senior Data Engineer will own the data engineering function on K2Share's Federal Team, partnering with technical and product leadership to deliver data products that support mission-critical decision-making for federal agency clients. * Design and build relational data layers that handle OSCAL and other structured compliance data - including ingestion, validation, transformation, and export workflows that preserve fidelity to source schemas across the full data lifecycle * Design and maintain data models that support governance, risk, compliance, scoring, and reporting workflows for federal cybersecurity programs, with OSCAL as the connective layer across them - including long-term retention and archival policies that align with federal recordkeeping and audit requirements * Design and build big-data processing pipelines on Databricks (PySpark, Delta Lake, Unity Catalog) that normalize cybersecurity data from across federal agency environments and produce analytical layers for trend analysis, executive reporting, and cross-program insights * Optimize data systems for performance and cost - identifying I/O and compute bottlenecks, scaling compute responsibly, and balancing throughput against the cost discipline federal engagements require * Architect, build and maintain AWS data infrastructure that meets federal security and operational requirements - working across services such as S3, Bedrock, Lambda, Fargate, and EC2 in support of compliance and analytical workloads * Design and implement audit-ready data primitives - change capture, access controls, validation, and lineage - that support agency reporting and continuous monitoring needs * Lead AI-first development and responsible AI deployment on the data team - using AI development tools as a standard part of the engineering loop, prototyping AI-assisted compliance workflows, and designing the production AI systems behind them (RAG architectures, vector store management, conversational agents, prompt and output guardrails, and evaluation pipelines), in alignment with federal AI governance guidance (OMB, NIST AI RMF) * Engage with federal agency stakeholders and internal teams during requirement discovery, delivery, and ongoing support - translating compliance needs into data products and customer feedback into improvements ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)