> Markdown version of [/jobs/ext/3548732-data-platform-engineer-cybersecurity-operations](https://www.wearedevelopers.com/jobs/ext/3548732-data-platform-engineer-cybersecurity-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Platform Engineer, Cybersecurity Operations - **Company:** Citi - **Location:** Irving, TX, United States - **Experience:** Expert - **Salary:** $156,160.0 - $234,240.0 - **Contract:** Permanent contract - **Skills:** Training Data, Query Performance, Java (Programming Language), Artificial Intelligence, Airflow, Amazon Web Services, Business Analytics Applications, Automation of Tests, Microsoft Azure, BigQuery, Cyber Security, Databases, Continuous Integration, Data Dictionary, Data Infrastructure, Data Warehousing, Database Schema, Intrusion Detection and Prevention, Python (Programming Language), Meta-Data Management, Metadata Standards, NoSQL, Cloud Services, Security Information and Event Management, Data Streaming, Feature Store, Snowflake, Apache Spark, Event Driven Architecture, Data Lakes, Apache Flink, Apache Kafka, Data Management, Software Version Control, Data Pipelines, Amazon Redshift, Golang, Programming Languages - **Published:** October 1, 2026 - **Apply:** https://citi.wd5.myworkdayjobs.com/2/job/Irving-Texas-United-States/Data-Platform-Engineer--Cybersecurity--Operations_26997079/apply ## About the Role * Deep, demonstrable expertise in data platform architecture and engineering - a proven track record of architecting and operating large-scale, mission-critical data platforms from the ground up, with strong judgment on trade-offs across scalability, cost, performance, and reliability. * Extensive, hands-on experience building and operating large-scale data pipelines (batch and streaming) using technologies such as Cribl, Kafka, Spark, Flink, Airflow, or equivalent. * Deep expertise in data lake and data warehouse architecture (e.g., Delta Lake, Iceberg, Snowflake, BigQuery, Redshift) at enterprise scale, including experience making foundational design decisions on storage formats, partitioning strategies, and query engines. * Proven, in-depth experience with cloud-native data infrastructure (AWS/Azure/GCP), including storage tiering, cost optimization, and event-driven architectures. * Strong understanding of security telemetry types - logs, network flow, endpoint/EDR data, identity events, cloud audit logs - and the architectural challenges of ingesting and normalizing them at scale. * Advanced proficiency in a major programming language (Python, Go, Java, or Scala) for pipeline development and automation. * Deep experience with schema design, data modeling, and metadata management for large, heterogeneous, high-velocity datasets. * Solid grounding in cybersecurity fundamentals, particularly SOC/OSVM operations, log management, and detection/analytics use cases. * Extensive experience with database technologies spanning relational, NoSQL, time-series, and columnar/analytical stores. * Strong software engineering fundamentals: CI/CD, infrastructure-as-code, version control, automated testing, and observability/monitoring practices. * Demonstrated ability to architect for scale and reliability in mission-critical, 24x7 operational environments. * Excellent cross-functional collaboration skills, able to work with security operations, detection engineering, and data science teams., * 10+ years of experience in architecting data platforms specifically for SOC, SIEM, or XDR environments. * Familiarity with AI/ML pipeline requirements (feature stores, training data pipelines) to support advanced security analytics. * Deep experience with open table formats (Iceberg, Delta Lake, Hudi) and modern lakehouse architectures. * Relevant certifications (e.g., cloud data engineering certifications, GIAC) are a plus but not required in lieu of hands-on expertise., A platform architect and builder at heart - someone who has spent their career deep in data platform engineering and wants to apply that expertise to one of the highest-stakes data domains in the enterprise: security telemetry at scale. You should be equally comfortable leading architectural / system design discussions and rolling up your sleeves to build and operate production systems., * Bachelor's degree/University degree or equivalent experience * Master's degree preferred ## Description We are seeking a Data Platform Engineer with deep, specialized expertise in data platform architecture and engineering to build and operate the scalable data pipelines, storage backbones, and data lakes that power our next-generation cybersecurity operations capabilities. This role sits at the foundation of our security analytics capability, ingesting and normalizing massive volumes of SOC telemetry into modern, high-performance analytics platforms., * Pipeline Engineering: Design, build, and operate scalable, resilient, high-throughput data pipelines (batch and streaming) that ingest security relevant telemetry - logs, alerts, network flow data, endpoint events, identity signals, and cloud events - from across the enterprise. * Data Normalization: Develop robust normalization and enrichment logic to transform heterogeneous SOC telemetry formats into standardized schemas consumable by downstream analytics and detection platforms. * Storage Backbone Architecture: Architect and maintain the storage backbone (data lakes, warehouses, streaming stores) that serves as the durable, queryable source of truth for SOC data at enterprise scale. * Data Lake Operations: Build and operate data lake infrastructure optimized for high-volume security telemetry, balancing cost, performance, retention, and query flexibility. * Integration with Analytics Platforms: Ensure seamless, low-latency delivery of normalized SOC data into modern analytics platforms, SIEM/XDR tooling, and other downstream systems. * Scalability & Performance Engineering: Continuously tune pipeline throughput, storage partitioning, and query performance to keep pace with growing telemetry volume and evolving analytics demands. * Reliability & Observability: Implement monitoring, alerting, and self-healing capabilities to ensure pipeline uptime, data completeness, and data quality at production scale. * Schema & Taxonomy Governance: Define and enforce data schemas, taxonomies, and metadata standards across ingested telemetry sources to enable consistent downstream consumption. * Automation of Onboarding: Build reusable frameworks and automation to rapidly onboard new telemetry sources as the tool and sensor ecosystem expands. * Security & Compliance: Ensure secure handling, encryption, access control, and regulatory compliance for sensitive security data throughout its lifecycle (ingestion, storage, processing, access). * Cross-Functional Collaboration: Partner closely with cybersecurity operations analysts, detection engineers, threat and exposure management teams, and data science teams to ensure the platform meets operational and analytical needs. * Documentation & Knowledge Transfer: Maintain clear architecture documentation, data dictionaries, and operational runbooks to support platform sustainability and team scalability. * ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Tomorrow's cloud data platforms - fully managed database-as-a-service (DBaaS)](https://www.wearedevelopers.com/videos/254-tomorrow-s-cloud-data-platforms-fully-managed-database-as-a-service-dbaas) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)