> Markdown version of [/jobs/ext/2073235-sre-l1-support-cloud-platform-ops-engineer](https://www.wearedevelopers.com/jobs/ext/2073235-sre-l1-support-cloud-platform-ops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE L1 Support/Cloud Platform Ops Engineer - **Company:** Bitdeer Technologies Group - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, JIRA, Command-Line Interface, Data Centers, Network Interface Controllers, Firmware, Monitoring of Systems, Issue Tracking Systems, Linux System Administration, Log Analysis, Nagios, Network Diagnostics, Prometheus, Diagnostic Tools, Graphics Processing Unit (GPU), Cloud Platform System, Grafana, Servicenow - **Published:** August 15, 2026 - **Apply:** https://www.wayup.com/i-j-SRE-L1-Support-Cloud-Platform-Ops-Engineer-Bitdeer-Technologies-Group-324719202980592/ ## About the Role + 2+ years in NOC, data center operations, or IT support role + Basic Linux system administration (command line, log analysis, service management) + Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent) + Experience with ticketing systems (ServiceNow, Jira Service Management) + Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement + Strong communication skills for shift handoffs, incident documentation, and escalation + Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation) + Curiosity about automation - you don't just execute the runbook, you notice when it's the third time this month and ask what should change. + Comfort with structured data - you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed. ## Description You are the first human in the loop - the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans - it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM-8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents. What you'll own + Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards. + Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts. + Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection. + Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery. + Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports. + Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira). + Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles). + Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team. + Maintain and update operational runbooks based on recurring issues. + Assist with hardware deployment, firmware updates, and inventory management under SME guidance. Feed the AIOps substrate + Every novel incident you resolve is data the platform team needs - you tag it, describe it, and hand it back so it becomes an automation. + Every runbook you touch should get closer to being executable by the platform, not by you. + Your handoff notes are structured signal, not free-form email. Why this role is different from a NOC job + You are not the last line of defense - the platform is. You are the training signal. + Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors. ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)