Lead Data Center Support Services Technician

Oracle
Santa Clara, CA, United States
2 months ago
Apply on oracle.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Audit Trail Computer Maintenance Data Centers Interoperability Network Configuration and Change Management Oracle (Applications) Cloud Services Performance Testing

Job description

Leads complex, data center hardware support services to ensure resilience, reliability, and availability at scale. Designs, governs, and signs off on advanced change activities and site augmentations, driving standardization of SOPs/MOPs across regions. Serves as the highest operational escalation point for incidents, leading cross-functional triage, root-cause analysis, and long-term corrective actions. Partners with engineering, capacity, networking, and security teams to anticipate risks, implement architectural improvements, and automate repeatable workflows. Mentors and develops team mates, uplifts best practices, and influences policy and process changes that improve performance across the group and organization., Datacenter Services Operations-Break/Fix and Hardware Maintenance:

  • Oversees lifecycle management for servers and components; sets standards for diagnosis, replacement, and performance validation without impacting critical workloads
  • Coaches teams on advanced power distribution practices and complex hardware interventions; validates that work meets reliability, safety, and SLA objectives Datacenter Services Operations-Network Configuration, Installation, and Augmentation:

  • Leads the planning, risk assessment, and execution governance for large-scale network builds and migrations; validates architecture, cabling, and end-to-end connectivity
  • Defines acceptance criteria and sign-off gates for network changes; ensures interoperability and error-free integration with upstream and downstream systems Technical Support and Trouble Shooting:

  • Owns the escalation queue and trend analysis; directs complex investigations, establishes fix-forward plans, and ensures high-quality documentation of resolutions within SLAs
  • Shapes the ticket taxonomy, triage playbooks, and automation triggers to improve response consistency and speed Safety and Compliance-Safety, Security, Compliance, and Documentation:

  • Governs adherence to SOPs/MOPs and site rules for all high-risk work; ensures alignment with physical security procedures, local regulations, and audit requirements; verifies functional testing of security controls
  • Ensures comprehensive documentation, audit trails, and change records; prepares for and leads compliance reviews and remediations Continuous Improvement:

  • Leads post-incident reviews and enterprise-level RCAs; socializes learnings and implements durable design/process changes that reduce risk across the fleet
  • Partners with engineering to mitigate availability risks, reduce single points of failure, and improve observability and alert fidelity Collaboration, Vendor Relations, and Leadership:

  • Mentors technicians and specialists; sets expectations for execution quality, safety, and customer focus; builds capability through training and coaching
  • Influences staffing and readiness plans for on-call and peak events; models inclusive practices and cross-team collaboration

Core Responsibilities

Planning & Execution - Drives completion of complex and/or ambiguous work areas in accordance with broad project requirements. Provides leadership to team members on shifting priorities and resources according to team needs.

Collaboration & Partnership - Collaborates across Lines of Business to lead the delivery of impactful work and provide support for critical team objectives. Leverages advanced understanding of business, stakeholder, and/or customer needs to build partnerships within and outside of team.

Problem Solving - Proactively develops and shares procedures to identify and resolve complex issues, serving as the final non-managerial escalation point for the team. Collects and reviews data and/or information to troubleshoot the most complex errors.

Continuous Learning - Models continuous learning by actively seeking to expand knowledge and learning new skills and/or tools, and staying current with trends and best practices in one’s field. Seeks out and leverages feedback and advanced training to refine improve skills.

Continuous Improvement - Recommends, implements, and shares strategies to improve the efficiency and effectiveness of complex processes, protocols, and workflows for own and other teams.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

Requirements

Are you an experienced Data Center Technician who has experience in data center design/operation, racking, stacking, hardware break-fix, and everything in between? Then we’d like to talk to you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on oracle.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:56 min

Leveraging GitOps for AI auditing and instant rollbacks

Jaroslaw Gajewski Jaroslaw Gajewski · World Congress 2026 Europe

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

2:41 min

Transitioning artificial intelligence infrastructure into scalable commodity cloud services

juarezjunior juarezjunior · World Congress 2024

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:09 min

Managing enterprise execution with the Operate runtime

Marcin Makowski Marcin Makowski · World Congress 2026 Europe

1:29 min

Tech infrastructure capacity and AI product innovations

Videos

See all

Related articles

See all