> Markdown version of [/jobs/ext/1826093-systems-operations-manager-data-platforms-teradata-hadoop](https://www.wearedevelopers.com/jobs/ext/1826093-systems-operations-manager-data-platforms-teradata-hadoop). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Operations Manager - Data Platforms -Teradata & Hadoop - **Company:** Wells Fargo - **Location:** Minneapolis, MN, United States - **Experience:** Experienced - **Salary:** $119,000.0 - $206,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Cloud Engineering, Cyber Security, Continuous Integration, Data Infrastructure, Disaster Recovery, Distributed Systems, Apache Hadoop, Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Teradata SQL, Software Vulnerability Management, Cloud Platform System, Mttr, Containerization, Kubernetes, Data Management, Devsecops - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a2493f0b76756090 ## About the Role * 5+ years of Systems Engineering, and Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education * 2+ years of Leadership experience * Hands-on experience with: + Teradata and Hadoop platforms + Distributed systems and data platform operations + Incident, problem, and change management processes Desired Qualifications: * Experience supporting enterprise-scale Teradata and Hadoop platforms * Demonstrated leadership in 24x7 production support and SRE environments * Strong experience in: + Automation, AIOps, and operational transformation + DevSecOps and CI/CD practices + Observability, monitoring, and platform telemetry * Familiarity with Kubernetes, containerization, and cloud-native architectures * Strong understanding of: + Multi-tenant data platforms and workload management + Regulatory, audit, and risk-controlled environments ## Description This role is accountable for platform stability, reliability, and operational excellence across a complex, multi-tenant ecosystem supporting 100+ tenants. The manager will lead a 24x7 operations team, apply Site Reliability Engineering (SRE) principles, and drive automation-led transformation to ensure predictable, resilient service delivery at scale. This is a hands-on leadership role requiring strong execution discipline, ownership, and the ability to operate in a high-risk, regulated environment, ensuring SLA adherence, compliance, and business continuity outcomes. In this role, you will: Operational Leadership & Platform Ownership * Lead end-to-end platform operations for Teradata and Hadoop environments, ensuring availability, performance, and resilience * Provide clear ownership and accountability for production services, operational outcomes, and service stability * welDrive incident, problem, and change management, including major incident command and recovery leadership * Lead 24x7 global support operations, including on-call governance and escalation management Operational Excellence & Service Performance * Own and drive SLA/OLA adherence, uptime, and service health metrics * Lead capacity management, performance tuning, and proactive issue prevention initiatives * Establish and enforce operational standards, runbooks, and service management practices * Drive root cause analysis (RCA) and long-term remediation of systemic issues * Drive adoption of automation, observability, and AIOps practices to reduce manual toil and improve MTTR. Governance, Risk & Compliance * Ensure alignment with enterprise risk, compliance, and change management frameworks * Drive patching, vulnerability remediation, and platform security posture * Maintain audit readiness, documentation quality, and control adherence * Identify, escalate, and mitigate operational and platform risks Multi-Tenant Platform Operations * Manage operations across shared, multi-tenant platforms, ensuring workload isolation and stability * Oversee resource allocation, scheduler configuration, and workload prioritization * Execute in high-risk production environments where changes impact multiple tenants simultaneously Site Reliability Engineering (SRE) & Automation * Apply SRE principles to improve reliability, availability, and scalability of data platforms * Drive automation-first operations to eliminate manual toil and standardize service delivery * Implement and enhance observability, monitoring, and self-service capabilities * Partner with engineering teams to improve platform reliability, operability, and service maturity * Drive adoption of automation, observability, and AIOps practices to reduce manual toil and improve MTTR. Stakeholder Engagement & Execution Alignment * Partner with Engineering, CIO-aligned teams, Cybersecurity, and LOB stakeholders * Provide clear, executive-ready communication on platform health, risks, and priorities * Drive cross-functional accountability and execution discipline across teams People Leadership & Talent Development * Lead, coach, and develop a team of Systems Operations engineers and analysts * Build a culture of ownership, accountability, and operational excellence * Manage resource allocation, workforce planning, and vendor/partner support * Develop team capabilities in SRE practices, automation, and platform operations maturity Resiliency & Business Continuity * Ensure resiliency posture across Teradata and Hadoop platforms, including: + Disaster recovery (DR) readiness and execution + RTO/RPO alignment and validation + Continuous improvement of recovery capabilities * Lead BCP execution and failover coordination for critical platforms ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)