Data Platform Operations Engineer

Scotiabank Group
Dallas, TX, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
0 years minimum
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Microsoft Azure Bash Shell Cloud Computing Data Infrastructure Identity and Access Management Information Technology Operations Python (Programming Language) Knowledge Management Log Analysis Networking Basics Windows PowerShell
+9 more
Reliability Engineering SQL Databases Enterprise Data Management Scripting Apache Spark AI Platforms Information Technology Data Management Databricks

Job description

The Data Platform Operations Engineer will play a critical role within the Enterprise Data & AI Technology organization - one of Scotiabank’s most significant enterprise-wide strategic initiatives. This organization drives data enabled decision making, AI innovation, and technology modernization across the Bank., The Data Platform Operations Engineer works under the guidance of senior engineers and platform leads to help maintain site reliability for our Data & AI platforms. This role focuses on monitoring alerts and dashboards, completing routine operational and maintenance tasks using predefined SOPs/runbooks, triaging and escalating incidents, and providing after-hours on-call support on a rotational basis. You will partner with teams such as IAM, Network, Cloud Ops, Security, and client delivery teams to resolve issues and keep the platform stable, secure, and available.

What You’ll Do

  • Monitoring & Alert Response: Monitor dashboards and alerts (Azure Monitor/Log Analytics, Databricks, and platform tooling), validate signal vs. noise, and take first-response actions according to SOPs.
  • Incident Triage & Escalation: Triage incidents by collecting logs/metrics, identifying likely impact, documenting findings, executing approved remediation steps, and escalating to on-call leads or SMEs with clear context.
  • SOP-Driven Maintenance: Perform routine maintenance tasks from predefined runbooks (e.g., operational checks, certificate/secret rotations as directed, housekeeping activities, basic platform validation, scheduled jobs health checks).
  • Ticket Handling & Request Fulfillment: Work intake from service queues, follow standard procedures for common requests (access requests, connectivity validation, workspace onboarding steps), and ensure requests are completed and communicated within agreed SLAs.
  • Operational Communications: Provide timely and accurate updates during incidents and maintenance activities, including status, next steps, and handoffs, using established communication channels and templates.
  • Reliability & Continuous Improvement: Identify recurring issues and operational pain points, suggest improvements to alerts/runbooks, and contribute to post-incident actions (e.g., updating SOPs, adding monitoring coverage).
  • Documentation & Knowledge Management: Maintain clear operational documentation, ensure runbooks are current, and capture lessons learned to improve onboarding and reduce time-to-resolve.
  • After-Hours Support (Rotation): Participate in after-hours/on-call rotations and perform approved response actions, escalating when required to meet service reliability targets.

Requirements

Do you have experience in Technical troubleshooting support?, * 0-2 years of experience in IT operations, production support, or a similar support role.

  • Foundational knowledge of cloud concepts (Azure preferred): identity/access basics, networking fundamentals, and how to navigate cloud portals and logs.
  • Comfort monitoring alerts/dashboards and troubleshooting using logs and metrics (Azure Monitor/Log Analytics or similar tools).
  • Exposure to data platforms (Databricks, Spark, SQL warehouses, or similar) is an asset; ability to learn quickly is essential.
  • Basic scripting ability (Python, Bash, or PowerShell) to run operational checks or automate simple repetitive tasks.
  • Familiarity with ITIL-style incident and change processes (ticketing, triage, documentation, and handoffs) is an asset.
  • Strong communication and customer support mindset: able to provide clear updates, follow SOPs precisely, and escalate effectively.
  • Willingness to participate in after-hours/on-call rotations as required.
  • Degree/college diploma in Computer Science, Engineering, Information Technology, or a related field is preferred (or equivalent practical experience).

About the company

Scotiabank is a leading bank in the Americas. Guided by our purpose: “for every future”, we help our customers, their families and their communities achieve success through a broad range of advice, products and services, including personal and commercial banking, wealth management and private banking, corporate and investment banking, and capital markets.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:38 min

Royal Bank of Canada document processing capabilities deployment

Mohak Chadha Mohak Chadha · WWC Europe 2026

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · WWC 2022

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

54 sec

Interpreting complex terminal commands safely using external explanation utilities

Dan Cranney +2 · LIVE

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

Videos

See all

Related articles

See all