Data Center Operations Manager

Amazon.com, Inc.
Boardman, OR, United States
7 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$125,000.0 - $155,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Cloud Computing Data Centers Monitoring of Systems Networking Hardware Uptime Nagios Alwayson Information Technology AWS Data Analytics

Job description

Amazon Web Services, Inc. seeks a Data Center Night Manager to lead mission-critical night-shift operations for AWS data centers. In this field-based role, you will manage technicians, oversee incident response, and safeguard uptime, security, and safety for global cloud infrastructure. You’ll coordinate with engineering, facilities, and security teams to resolve complex issues, enforce runbooks and change processes, and deliver clear shift reporting. In AWS’s customer-obsessed, innovative culture, you’ll drive continuous improvement, develop talent, and learn cutting-edge cloud technologies while supporting large-scale, always-on services., * Lead night-shift operations across AWS data center facilities, ensuring uptime, security, and safety

  • Monitor infrastructure health, incident queues, and SLAs; coordinate rapid response to outages or performance issues
  • Manage and coach night-shift technicians, scheduling, and workload prioritization
  • Implement and enforce operational runbooks, change management, and escalation procedures
  • Partner with engineering, facilities, and security teams to resolve complex technical issues
  • Track operational metrics, create shift reports, and communicate status to daytime leadership
  • Ensure compliance with AWS security, safety, and regulatory standards
  • Drive continuous improvement initiatives to reduce incidents and improve efficiency

Requirements

  • Data center operations management
  • Incident and problem management
  • IT infrastructure monitoring tools (e.g., Cloud
  • Watch, Nagios, Solar
  • Winds)
  • Server and network hardware troubleshooting
  • Change management and runbook execution
  • People management and shift scheduling
  • Root cause analysis and reporting
  • ITIL-based service operations
  • Physical security and access control procedures
  • Environmental and power systems awareness (HVAC, UPS, generators)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

2:22 min

Leveraging unique cultural backgrounds in engineering design

Ixchel Ruiz · LIVE

6:51 min

Audience questions on cloud security and operational capacity

Steffen Heilmann · World Congress 2021

4:35 min

Defining service-level indicators based on user behavior

Maxim Schepelin Maxim Schepelin · World Congress 2026 Europe

9:18 min

Provisioning external availability and performance monitoring endpoints

Liam Hurrell +1 · World Congress 2021

2:42 min

Measuring engineering success metrics and system stability across frameworks

Thomas Fuchs +3 · LIVE

Videos

See all

Related articles

See all