> Markdown version of [/jobs/ext/2014616-data-center-operations-and-maintenance-engineer](https://www.wearedevelopers.com/jobs/ext/2014616-data-center-operations-and-maintenance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Operations and Maintenance Engineer - **Company:** Designworks NY LLC - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Data Centers, Reliability Engineering, Prometheus, AI Infrastructure, Datadog, Graphics Processing Unit (GPU), Computer Network Operations, Grafana, Information Technology, Hardware Infrastructure, Pagerduty - **Published:** August 10, 2026 - **Apply:** https://www.careerbuilder.com/job-details/data-center-operations-and-maintenance-engineer-bellevue-wa--2e0e58df-f992-402c-850d-417c38008dff ## About the Role * Experience in a data center operations, site reliability, or infrastructure operations role, ideally supporting GPU or large-scale compute environments. * Strong incident response and troubleshooting skills across hardware, networking, and systems layers. * Comfortable working in an early-stage environment where processes and tooling are still being established. * Experience in multiple technical environments is a plus. Preferred Qualifications * Experience with on-call/incident-management tooling (e.g., PagerDuty, Opsgenie) and monitoring stacks (e.g., Prometheus, Grafana, Datadog). * Background supporting GPU cluster or data center operations post-launch., America Online (AOL), Artificial Intelligence (AI), Automation, Cloud Computing, Computer Systems, Customer Experience, Customer Support/Service, GPU (Graphics Processing Unit), Incident Management, Incident Response, Intelligence Agencies, Leadership, Machine Tool, Network Operations Center, Network System Hardware, On Call, Operating Systems, Operational Support, Startup, Systems Scalability, Team Player ## Description This is an individual contributor role responsible for helping to build and scale the Operations & Maintenance function supporting mission-critical AI infrastructure. The focus of these roles will be to support post-launch monitoring, incident response, reliability and provide exceptional customer support. You will partner closely with Engineering, Infrastructure, Networking, Hardware, and Operations and Maintenance leadership to create a highly reliable, scalable operating system capable of supporting one of the industry's most advanced AI infrastructure platforms. What You'll Do * Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues. * Lead or support incident response for production issues, driving toward fast, effective resolution. * Collaborate with the engineering teams to make sure monitoring, alerting, and operational tooling are first class and make issues are caught before they impact customers. * Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence. * Contribute to runbooks, on-call practices, and operational maturity as the platform scales., * Build the Operations & Maintenance organization, processes, tooling and systems from its earliest stages. * Influence the strategy, processes, tooling, and culture supporting next-generation AI infrastructure. * Work alongside experienced engineers and engineering leaders solving some of the industry's most complex infrastructure challenges. * Join a company combining startup agility with substantial long-term investment and stability. * Play a foundational role in scaling the infrastructure powering tomorrow's AI applications. ## Related Videos - [AI in Leadership: How Technology is Reshaping Executive Roles](https://www.wearedevelopers.com/videos/1705-ai-in-leadership-how-technology-is-reshaping-executive-roles) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)