> Markdown version of [/jobs/ext/2722506-data-center-operations-systems-engineer-iii](https://www.wearedevelopers.com/jobs/ext/2722506-data-center-operations-systems-engineer-iii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Center Operations Systems Engineer III - **Company:** Lambda Inc. - **Location:** Los Angeles, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** JIRA, Booting (BIOS), Data Centers, Data Center Infrastructure Management (CIM), Network Topologies, Issue Tracking Systems, InfiniBand, Linux System Administration, Software Engineering, Computer Networking Systems, High Performance Computing, Hardware Infrastructure, Zendesk - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/data-center-operations-systems-engineer-iii-los-angeles-lambda-3-8816336 ## About the Role * Have strong past experiences with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management * Be familiar with carrier DIA circuit test and turn ups, fiber testing and troubleshooting * Basic knowledge of cable optics and the different types of use * Solid understanding of single and three phase power theories * PDU balancing and why it is important * Familiar with multiple cable media types and their uses * Knowledge of cold isle and hot isle containment * Solid understanding of server hardware and boot process * Ability to structure, collaborate and iteratively improve on complex maintenance MOPs. * Working with product management, support, and other teams to align operational capabilities with company goals. * Translating business priorities into technical and operational requirements. * Supporting cross-functional projects where infrastructure plays a critical role. * Are action-oriented and willingness to train junior staff on best practices * Are willing to travel for bring up of new data center locations as needed (25%-30%), * Have 5+ years experience with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management * Experience with/or knowledge of network topology and configurations and 400gb Infiniband architectures. * Experience with/or knowledge of DDP or SCM cluster storage systems. * Have 5+ years working with and reporting from a ticketing systems like JIRA and Zendesk * Advanced experience with Linux administration * Experience with High Performance Compute GPU systems (air or water cooled) - especially Nvidia NVL72 ## Description * Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured. * Troubleshoot hardware and software issues in some of the world's most advanced GPU and Networking systems. * Document and update data center layout and network topology in DCIM software * Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments * Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers * Partner with HW Support teams to ensure data center hardware incidents with higher level troubleshooting challenges are resolved, reported on and solutions are disseminated to the large operations organization. * Work with RMA team to ensure faulty parts are returned and replacements are ordered * Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [The AI Agent Path to Prod: Building for Reliability](https://www.wearedevelopers.com/videos/1523-the-ai-agent-path-to-prod-building-for-reliability) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [5 Years in Cloud Native: The Good, the Bad, and the Bill](https://www.wearedevelopers.com/videos/100111-5-years-in-cloud-native-the-good-the-bad-and-the-bill) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)