> Markdown version of [/jobs/ext/2283299-senior-principal-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2283299-senior-principal-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Principal Performance Engineer - **Company:** The Depository Trust & Clearing Corporation - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Query Performance, Cloud Computing, Profiling, Databases, Disaster Recovery, Distributed Systems, Java Virtual Machine (JVM), Apache JMeter, Performance Tuning, Regression Testing, Reliability Engineering, Prometheus, Test Data, Datadog, Delivery Pipeline, Snowflake, Grafana, Gatling, Apache Kafka, Splunk, Dynatrace - **Published:** August 28, 2026 - **Apply:** https://ebxr.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1/requisitions/preview/214540 ## About the Role * Minimum of 10 years of related experience * Bachelor's degree preferred or equivalent experience Talents Needed for Success: * 10+ years of hands-on performance engineering and reliability experience in latency-sensitive or high-volume transaction systems, with continued depth in profiling, dump analysis, and test harness development. * Strong performance engineering discipline across load, stress, soak, and spike testing, using workload models grounded in real production traffic and analyzing results through p99 and p99.9 behavior rather than averages. * Deep diagnostic capability across JVM tuning, heap and thread dump analysis, flame graphs, async profiling, and bottleneck isolation across application, database, network, and storage layers-even when standard dashboards appear healthy. * Resiliency and chaos engineering expertise including fault injection, failure-mode cataloging, game day design, and DR exercises supported by measured RTO/RPO evidence. * Distributed systems and data platform performance depth spanning Kafka throughput, consumer lag, rebalancing, partition skew, backpressure, exactly-once tradeoffs, object storage, lakehouse query performance, and Snowflake warehouse tuning. * Observability architecture experience covering OpenTelemetry, distributed tracing, metrics taxonomy, SLI/SLO and error budget design, actionable alerting, and disciplined control of cardinality and telemetry cost. * Automation embedded into delivery pipelines through CI performance gates, automated baselining, regression detection, and practical test data and environment strategies for constrained end-to-end environments. * Broad tooling proficiency with engineering depth across JMeter, Gatling or k6; Prometheus and Grafana; Dynatrace, Datadog or Splunk; and the ability to build custom harnesses when standard tools do not fit the workload. * Regulated environment fluency including DR obligations, production readiness reviews, audit evidence, and disciplined incident and postmortem practices. * Influence without authority with the credibility to set and enforce NFR standards across delivery squads and support release-readiness decisions with clear, evidence-based judgment. ## Description As a Senior Principal Performance & Resiliency Engineer, you will play a critical role in ensuring the reliability, scalability, and operational excellence of DTCC's most business-critical platforms. You will lead the strategy, architecture, and execution of performance engineering, resilience testing, capacity planning, and operational readiness initiatives across complex distributed systems and cloud-native environments. In this role, you will partner closely with engineering, infrastructure, architecture, and product teams to proactively identify performance bottlenecks, eliminate systemic risks, and enhance platform resilience. Your leadership will help ensure that critical financial market infrastructure systems can meet evolving business demands, regulatory expectations, and client service commitments while maintaining the highest standards of availability and reliability. You will drive a culture of engineering excellence by championing observability, automation, chaos engineering, disaster recovery preparedness, and continuous performance optimization. Your contributions will directly strengthen DTCC's ability to deliver secure, resilient, and highly available services that support global financial markets. Your Primary Responsibilities: * Make NFRs measurable and release-gated. Define clear targets for latency, throughput, availability, recovery, and capacity, and validate them under production-like volume before release. * Embed continuous performance and resilience validation. Automate regression testing and failure-injection scenarios in CI so reliability issues are identified before production impact. * Establish estate-wide observability and readiness standards. Create a common instrumentation, SLI, and capacity-planning baseline so risk platforms can scale, recover, and be operated consistently across the estate. ## Related Videos - [Continuous testing - run automated tests for every change!](https://www.wearedevelopers.com/videos/190-continuous-testing-run-automated-tests-for-every-change) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Load Testing AI: Aiming at a Moving Target](https://www.wearedevelopers.com/videos/1914-load-testing-ai-aiming-at-a-moving-target) - [Systems Thinking for Performance: How to Diagnose and Fix Slow Systems without Guessing](https://www.wearedevelopers.com/videos/2103-systems-thinking-for-performance-how-to-diagnose-and-fix-slow-systems-without-guessing) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Benefits of Using JMeter For Performance Testing](https://www.wearedevelopers.com/magazine/96-benefits-of-using-jmeter-for-performance-testing) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)