Counterparty Credit Risk (CCR) Operations IT Cloud Engineer

PROPERTY & CASUALTY MANAGEMENT SYSTEMS, INC.
New York, NY, United States
18 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Agile Methodology Confluence JIRA Microsoft Azure Cloud Computing Cloud Engineering Software Documentation Continuous Integration Data Validation Data Integrity Extract Transform Load (ETL)
+20 more
DevOps Identity and Access Management Information Technology Operations Python (Programming Language) Reliability Engineering Power BI Software Engineering SQL Stored Procedures SQL Databases Data Streaming System Availability Snowflake Apache Spark Mttr Containerization AngularJS Pyspark Information Technology Servicenow Databricks

Job description

  • We are is seeking a Counterparty Credit Risk (CCR) Operations IT Cloud Engineer consultant to support production stability, operational execution, and continuous improvement across CCR technology platforms.
  • This consultant role will assist with maintaining availability, performance, data integrity, and control adherence for systems supporting Exposure Management (EM), Credit Risk Reporting (CRR), Risk Analytics, Front Office users, and Regulatory Reporting. The position supports reliable exposure calculation, limit monitoring, and intraday what-if analysis capabilities across OTC, Exchange-Traded Derivatives (ETD), and Securities Financing Transactions (SFT).
  • The successful candidate will work with the CCR Operations IT Lead, onsite team members, offshore resources, engineering teams, and business stakeholders to help operate and improve a cloud-enabled IT operations model while reducing manual effort and operational risk through automation and process discipline.
  • The candidate should be a hands-on cloud engineer with strong production support discipline, the ability to troubleshoot issues, communicate clearly, manage multiple tasks, and contribute effectively within a collaborative team environment.

Major Responsibilities:

Exposure Management Application Support:

  • Support key production operations for applications supporting EM, including EPE/PFE calculations, limit monitoring, what-if pre-trade intraday analysis, and reporting.
  • Ensure platforms consistently meet the availability, performance/SLAs, and data-quality expectations of EM users during daily, intraday, and regulatory cycles (monthly, quarterly, annually).
  • Coordinate with onsite and offshore IT members to translate business priorities into day-to-day operational execution.

Production Operations & Continuous Improvement:

  • Demonstrate a continuous improvement mindset through their example, with a focus on reducing incidents, manual interventions, and operational risk, and improving turnaround time for incidents.
  • Contribute to measurable operational improvements using data points and KPIs/metrics such as availability, incident trends, MTTR, SLA breaches, reruns, and batch success rates.
  • Identify and execute practical opportunities for process standardization, tooling enhancements, and Operational simplification across the CCR application stack.

Global Delivery & Offshore Leverage:

  • Support a globally distributed operating model, leveraging offshore teams for L1/L2 support, batch operations oversight, upstream data feed validation, data quality monitoring.
  • Follow and help refine clear onshore/offshore operating boundaries, escalation paths, and ownership models to support seamless production coverage.
  • Contribute knowledge transfer, documentation standards, and runbook maturity to improve offshore effectiveness and reduce dependency on key individuals, ensuring service quality, productivity, and continuous skill uplift.

Incident, Problem & Change Management:

  • Timely escalate highseverity production incidents, particularly those impacting exposure reporting, limit breaches, what-if intraday analysis or regulatory deliverables.
  • Perform root cause analysis (RCA) for impactful incidents (Calculation breaks, data validation/reconciliation issues), ensuring recurring issues are eliminated through permanent fixes rather than shortterm workarounds.
  • Support change and release governance aligned with Client CAB, DevOps, CI/CD, risk, audit, and control standards.

Platform Reliability, Automation & Resilience:

  • Implement automation and selfhealing for batch monitoring, data validation, reconciliations, and recovery processes for production as well as lower testing environments.
  • Drive with Engineering and Infrastructure teams to improve observability, alerting, capacity planning, and resilience of CCR platforms.
  • Ensure DR/BCP readiness for EMcritical systems, including regular/annual testing and documented recovery procedures.

Regulatory, Audit & Risk Controls:

  • Ensure CCR applications meet regulatory, audit, security, and internal risk management requirements.
  • Support IT Operations activities related to internal audits, regulatory exams, and model governance reviews by providing evidence, documentation, and control execution support.
  • Maintain robust IT control frameworks (Access management, change control, data integrity).

Requirements

  • 10+ years of experience in IT Operations / Production Support within financial services.
  • Experience a senior consultant role supporting risk, exposure, or comparable financial technology platforms with hands-on work capabilities.
  • Demonstrated experience operating the offshoreleveraged global support models in a regulated environment.
  • Bachelor s Degree in Computer Science, Management of Information Systems, or related business discipline(s).

Technical & Domain Expertise:

  • understanding of Exposure Management and Counterparty Credit Risk concepts, including derivative trade cycles, market data, PFE, EPE, Collateral/margin management, limits, stress testing, and whatif analysis.
  • Experience supporting batchintensive and intraday realtime risk platforms (Java, Angular, Python, PySpark, ETL, Stored Procedure, SQL, Snowflake, PowerBI).
  • Hands-on experience with Cloud computing and infrastructure (Databricks, Medallion architecture, Azure, ADF, Spark based distributed compute, other Cloud native technologies).
  • Proven track record driving operational excellence, automation, and process maturity with an Agile based application/software development.
  • Proven ability to utilize ServiceNow, JIRA, Confluence, PowerBI to manage incidents, tasks, and releases, generate a system diagnosis report, and demonstrate KPIs based operational improvement.
  • Familiarity with ITIL based service management (ITSM) frameworks, DevOps, CI/CD, and reliability engineering practices.

Leadership & Soft Skills:

  • Continuous improvement mindset, with the ability to contribute to operational maturity across onshore and offshore teams.
  • Strong experience collaborating with distributed, multicultural teams and thirdparty vendors.
  • Excellent interpersonal, written and verbal communication skills with the ability to engage Exposure Management users, and senior Technology team leads.
  • Ability to prioritize and effectively manage multiple tasks under time-critical and regulatory pressure.
  • Highly motivated, self-directed individual with the ability to work independently and in team environments.
  • Effective presenter capable of articulate strategic execution plans including system/architecture/data flow diagrams, combined with extreme attention to detail.
  • Strong analytical, problem solving (breaking ambiguolarge problems to smaller/less complex ones), and decision-making skills.
  • Collaborative team player and relationship builder.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all