> Markdown version of [/jobs/ext/1423000-data-platform-engineer-vp-ii-state-street-investment-management](https://www.wearedevelopers.com/jobs/ext/1423000-data-platform-engineer-vp-ii-state-street-investment-management). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Platform Engineer, VP II - State Street Investment Management - **Company:** State Street - **Location:** Boston, MA, United States - **Salary:** $120,000.0 - $202,500.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Analysis, Apache HTTP Server, Automation of Tests, Microsoft Azure, Continuous Integration, Information Engineering, Data Infrastructure, Data Security, Data Sharing, Data Virtualization, Software Debugging, DevOps, Distributed Computing Environment, Information Lifecycle Management, Python (Programming Language), Online Analytical Processing, Online Transaction Processing, Performance Tuning, Systems Development Life Cycle, Reliability Engineering, Cloud Services, Software Engineering, SQL Databases, Data Streaming, Unstructured Data, Management of Software Versions, Google Cloud, GitHub Copilot, Snowflake, Apache Spark, Caching, Data Strategy, Cloudformation, Information Technology, Collibra, Data Management, Physical Data Models, Terraform, Databricks - **Published:** July 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f6e1e7f8abbcb41e ## About the Role * Master's degree in Computer Science or a related engineering discipline. * 15+ years of technology experience in the financial services industry. * Strong hands-on programming in Python and SQL (Java/Scala a plus), with the ability to design, debug, and optimize distributed data processing; familiarity with AI-assisted development tools (e.g., GitHub Copilot, Codex, Anthropic tools). * Deep experience with Spark (batch and streaming) including performance tuning, troubleshooting, and cost-aware optimization. * Hands-on experience building and operating modern Lakehouse platforms using Snowflake and Databricks on AWS (Azure acceptable), supporting high concurrency analytical workloads. * Working knowledge of open table formats and metadata ecosystems, including Apache Iceberg and table catalogs/governance services (e.g., Unity Catalog or equivalent). * Demonstrated hands-on experience designing and implementing modern, open architecture data platforms supporting OLAP, OLTP, and near real-time workloads; familiarity with catalog services that enable multiple compute engines, high concurrency, and strong performance. * Extensive hands-on experience architecting and implementing enterprise DevOps and CI/CD pipelines with cloudagnostic solutions across public cloud providers (e.g., AWS, GCP, Azure). * Expertise in zero-copy data sharing, data virtualization, catalog federation, and governed data access patterns, enabling scalable and trusted data consumption across platforms. * Experience working with both unstructured data sources and processing patterns * Strong operational and engineering excellence mindset, with experience establishing SDLC standards, CI/CD, observability, reliability engineering, incident management, and continuous improvement practices for mission-critical platforms., * Previous experience in data/platform development within a financial institution. * Experience with both structured and unstructured data management tools and associated patterns * Familiarity with modern governance frameworks (Unity Catalog, Collibra, Alation). * Expertise in AI-assisted coding using GitHub Copilot or Claude Code Assistant. ## Description * Execute the long-term vision, target architecture, and roadmap for a unified, cloud-native data platform spanning ingestion, transformation, storage, governance, and secure access at scale. * Collaborate with enterprise and application architects to define data strategies and deliver logical/physical data models aligned to analytical workloads. * Collaborate on the on the core platform capabilities (Lakehouse, batch, events, streaming, metadata/catalog, observability, and security), ensuring reliability, scalability, and operability. * Design, implement, and optimize Iceberg-based tables (partitioning, compaction, metadata) for consistent, performant analytical access across multiple compute engines. * Establish data engineering excellence: CI/CD, IaC (Terraform/CloudFormation), automated testing, schema/versioning practices, data quality controls, and end-to-end observability. * Partner with Product, Analytics, ML, and domain SMEs to define data semantics and data product contracts (schemas, SLAs, documentation, versioning/backward compatibility), sand to enforce governance, stewardship, and compliance standards. * Build and publish curated semantic layer models (serving models/marts), exposing governed BI endpoints and/or consumption APIs to ensure consistent metrics and business definitions. * Optimize pipeline and query performance through effective partitioning/clustering, caching, archiving, and data lifecycle (purge/retention) patterns. * Evaluate and advance the use of agentic AI frameworks to improve software delivery across the full software development life cycle (SDLC). ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)