> Markdown version of [/jobs/ext/2246751-devops-infrastructure-lead](https://www.wearedevelopers.com/jobs/ext/2246751-devops-infrastructure-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps & Infrastructure Lead - **Company:** CloudFlare - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $130,000.0 - **Contract:** Permanent contract - **Skills:** Secure Shell (SSH), Microsoft Access, PHP (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, JIRA, LAMP (Software Bundle), Command-Line Interface, Cloud Computing, Databases, Cron, Cursor (Graphical User Interface Elements), DevOps, Domain Name System (DNS), Github, Identity and Access Management, Laravel, Linux System Administration, MongoDB, MySQL, Node.Js, Redis, Software Tools, Digitalocean, Zero Trust Network Access, Runbook, Backup and Restore, Datadog, AWS Cdk, Data Logging, Pulumi, Scripting, GitHub Copilot, ReactJS, Amazon Virtual Private Cloud (VPC), Cloudformation, Deployment Automation, Cloudflare, React Native, Front End Software Development, Functional Programming, Cloudwatch, Terraform, GPT, Serverless Computing - **Published:** August 26, 2026 - **Apply:** https://careers.jobscore.com/apply_flow/applications/go?job_id=aBgLCYFM9idQkn0tqCIfqR&name=BuiltIn&ref=rss&sid=69 ## About the Role * Senior-level experience operating production infrastructure. * Deep, hands-on Linux server administration (the traditional, "old-school" kind): operating, securing, and troubleshooting manually managed production servers (LAMP/LEMP, system services, cron, networking, SSH) directly at the command line, not only through a cloud console. * Experience with DigitalOcean, Linode, AWS EC2, bare VPS hosting, or comparable environments. * Senior database operations: migrating self-managed MySQL to a managed service, replication, backup validation, restore testing, and IO isolation. * Strong Cloudflare across DNS, WAF, CDN and caching behavior, page rules, Workers, Pages, and Zero Trust/Access, including traffic routing and origin protection. * PHP/Laravel application environments, and experience with a managed Laravel runtime (Laravel Cloud and/or DigitalOcean App Platform). * Datadog or a comparable observability platform for monitoring, alerting, dashboards, logs, and incident investigation. * Infrastructure-as-code such as Terraform, Pulumi, AWS CDK, Serverless Framework, or CloudFormation. * CI/CD pipelines and deployment automation. * Practical AWS experience (Lambda, IAM, VPC, CloudWatch, S3, SSM/Secrets Manager, queues). * Good judgment around production safety, access control, secrets, backups, and incident response. * Willingness to carry real on-call responsibility and respond to production incidents outside normal business hours; this is not a strict 9-to-5 role. * A habit of documenting what you learn and creating runbooks others can follow. * Practical experience using AI tools (ChatGPT, Claude, Cursor, GitHub Copilot, or similar), with strong judgment about where human verification is required. * Ability to work independently in a small, remote engineering organization where practical ownership matters more than bureaucracy. Nice to Have * Experience migrating manually managed services onto managed platforms or IaC. * Experience moving static frontends onto Cloudflare Pages. * Managed migrations for MongoDB, OpenSearch, or Valkey/Redis. * Experience supporting Node.js, React, and React Native alongside PHP. * Experience helping organizations reduce infrastructure bus-factor risk. * Experience working with external DevOps/security partners or auditors. ## Description This is not a standard 9-to-5 role. Production issues do not keep business hours, so it carries real on-call responsibility: you need to be reachable and able to respond when unforeseen incidents arise. What You'll Do * Administer and improve existing DigitalOcean infrastructure. * Support and improve Linux-based production server environments. * Migrate self-managed databases onto managed database services, with validated failover, backups, and recovery. * Move applications onto managed runtimes (including Laravel Cloud where it fits), replacing manual deploy processes with automated, repeatable pipelines. * Expand and harden our use of Cloudflare for edge, static hosting, caching, and security. * Build a clear inventory of servers, services, databases, domains, access paths, backups, monitoring, and operational risks. * Create and maintain practical runbooks for common and emergency infrastructure workflows. * Improve incident response, escalation paths, monitoring, logging, and alerting. * Review and improve backup, restore, and disaster-recovery procedures. * Identify recurring manual work and convert it into safer procedures, scripts, automation, or infrastructure-as-code. * Help define infrastructure-as-code standards and move appropriate infrastructure into repeatable, version-controlled workflows. * Work with AWS services where needed (Lambda, VPC, IAM, CloudWatch, S3, SSM/Secrets Manager, queues). * Use AI tools to accelerate discovery, documentation, scripting, troubleshooting, and automation, with strong production-safety judgment. * Partner with engineering leadership to prioritize infrastructure risk and modernization; track work clearly in Jira/GitHub and communicate proactively about risks, tradeoffs, and blockers. What Success Looks Like In the first 30-60 days, you'll take ownership of how we see and operate our infrastructure, building on what we already track and closing the gaps. You'll validate and take ownership of what already exists: * Our infrastructure inventory and server map * Our monitoring and alerting * Our DNS / Cloudflare configuration * Our prioritized infrastructure risk register You'll create what we're missing: * An access and credential map * Verified backup and restore status for critical systems (tested, not assumed) * Runbooks for the highest-risk operational workflows In the first 90 days, you'll move us toward a durable, consolidated model. Success means: * The first core database migrated to a managed service, with a tested restore, plus a clear, sequenced plan for the rest. * The first application running on a managed runtime (App Platform or Laravel Cloud). * The first static frontend served from Cloudflare Pages. * A measurably stronger edge security posture. * Critical systems no longer understood by only one person; common tasks have documented procedures; manual processes are being converted to automation; AI is used safely to reduce toil. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [From clicks to cribs - How to find your dream home with web scraping](https://www.wearedevelopers.com/videos/767-from-clicks-to-cribs-how-to-find-your-dream-home-with-web-scraping) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 103 - Superb Owl Trafficking](https://www.wearedevelopers.com/magazine/388-dev-digest-103-superb-owl-trafficking) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)