> Markdown version of [/jobs/ext/2045060-lead-application-support-engineer](https://www.wearedevelopers.com/jobs/ext/2045060-lead-application-support-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Application Support Engineer - **Company:** The Depository Trust & Clearing Corporation - **Location:** Boston, MA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Application Integration Architecture, CA Workload Automation Ae, Command-Line Interface, Databases, Data Warehousing, Linux, DevOps, Disaster Recovery, Domain Name System (DNS), Middleware, File Transfer, Data-Flow Analysis, Monitoring of Systems, HP Systems Insight Manager, IBM WebSphere MQ, Identity and Access Management, Knowledge Management, PostgreSQL, Log Analysis, Enterprise Messaging Systems, Networking Basics, Release Management, Reliability Engineering, Runbook, Software Deployment, Software Engineering, SQL Databases, TCP/IP, Datadog, File Transfer Protocol (FTP), Load Balancing, Microsoft Power Automate, Snowflake, Grafana, Software Troubleshooting, Generative AI, Firewalls (Computer Science), AWS Glue, Data Management, Cloudwatch, Splunk, Dynatrace, Servicenow - **Published:** August 13, 2026 - **Apply:** https://ebxr.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1/requisitions/preview/214281 ## About the Role * Minimum of 6+ years of experience in Application Support, Production Support, Site Reliability Engineering (SRE), or related roles * Bachelor's degree and/or equivalent practical experience * Strong experience supporting a 24x7 production environment * Excellent analytical, troubleshooting, and problem-solving skills with the ability to lead complex incident investigations Talents Needed for Success: * Amazon Web Services (AWS) experience REQUIRED, including: + ECS + EC2 + RDS PostgreSQL + AWS Glue + Kinesis + S3 + CloudWatch Monitoring + IAM Security and Access Controls + Lambda (preferred) * Strong PostgreSQL administration and SQL experience + SQL query development and performance troubleshooting + Command-line database support + Monitoring and diagnostics * Experience supporting Snowflake data platforms * Experience with IBM MQ + Queue Manager administration + Channel troubleshooting + Message flow analysis * Experience with Autosys scheduling and batch operations * Linux/Unix administration and command-line experience * Experience troubleshooting file transfer solutions (SFTP, Managed File Transfer platforms) * Understanding of application integrations, APIs, middleware, and messaging platforms * Experience with log analysis and monitoring tools * Knowledge of networking fundamentals including DNS, TCP/IP, firewalls, and load balancing Automation & AI * Strong willingness and aptitude to quickly adopt new technologies, including: + Generative AI tools + Microsoft Copilot + AI-assisted troubleshooting solutions * Experience leveraging AI to improve operational efficiency and incident response is highly desirable Operational Excellence * Experience supporting Production, Disaster Recovery (DR), and Operational Resiliency testing * Experience performing root cause analysis and driving permanent corrective actions * Ability to lead or participate in Major Incident Management (MIM) calls * Experience with release management, change management, and production deployments ITSM & Governance * Experience using ServiceNow or similar ITSM platforms for: + Incident Management + Change Management + Problem Management + Knowledge Management * Experience creating and maintaining technical documentation, runbooks, and support procedures * Understanding of audit, compliance, and operational risk requirements Soft Skills * Excellent verbal and written communication skills * Ability to communicate effectively with business, development, infrastructure, and executive stakeholders * Strong ownership mindset with the ability to drive issues to resolution * Ability to prioritize multiple competing tasks in a fast-paced environment * Strong collaboration and teamwork skills, * Experience supporting: + Financial market infrastructure applications + Trade processing platforms + Data warehousing solutions + Real-time messaging systems * Experience with observability tools such as Splunk, Datadog, Dynatrace, or similar monitoring platforms * Experience working in Agile, DevOps, or SRE organizations * Understanding of cloud modernization and application migration initiatives Key Success Factors * Rapid troubleshooting of complex production issues * Automation-first mindset * Continuous service improvement focus * Strong customer and stakeholder orientation * Ability to learn new technologies quickly and become a subject matter expert * Commitment to operational stability, resiliency, and platform modernization ## Description Being a member of IT CSS WRAFT Delivery team, in this role, you will help ensure the stability, resiliency, and continuous improvement of business-critical applications supporting DTCC's global financial market infrastructure. Leveraging deep production support expertise across AWS, PostgreSQL, Snowflake, IBM MQ, Linux, and related technologies, you will lead complex incident resolution, strengthen operational controls and disaster recovery readiness, and automate monitoring and support activities. Your work will reduce service disruption, mitigate operational risk, and enable secure, reliable, and modernized platforms for DTCC's clients and business partners. Your Primary Responsibilities: * Verify analysis performed by team members and implement changes required to prevent reoccurrence of incidents * Resolve Critical application alerts in a timely fashion including production defects, providing business impact and analysis to teams, handling minor enhancements as needed * Review and update knowledge articles and runbooks with application development teams to confirm information is up to date * Collaborate with internal teams to provide answers to application issues and escalate to as needed * Validate and submit responses to requests for information from ongoing audits * Review and Execute Disaster Recovery scripts during planned and unplanned outages, providing BCM evidence as needed * Identify and implement automation opportunities to reduce manual effort associated with application monitoring * Partner with development teams to provide input into the design and development stages of applications * Execute the pre-production/production application code deployment plans and end to end vendor application upgrade process (e.g. SNOW, SF, etc) * Aligns risk and control processes into day to day responsibilities to monitor and mitigate risk; escalates appropriately, Work Schedule Requirements * Willingness to participate in: + On-call support rotations + Weekend support activities + Late-night maintenance windows + Disaster Recovery and Resiliency testing events * Ability to respond to production incidents during off-hours when required ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)