> Markdown version of [/jobs/ext/2227867-lead-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2227867-lead-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead DevOps Engineer - **Company:** Exadel, Inc. - **Location:** Boulder, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $187,200.0 - $208,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Macintosh Application Environment, Microsoft Azure, Bash Shell, Information Systems, Continuous Integration, DevOps, Disaster Recovery, Domain Name System (DNS), Fault Tolerance, Python (Programming Language), Key Management, Performance Tuning, Windows PowerShell, Prometheus, Strategies of Testing, Software Vulnerability Management, Datadog, Policy as Code, Scripting, Transport Layer Security, Load Balancing, Cloud Platform System, Grafana, Mttr, Infrastructure as Code (IaC), Cloudformation, Information Technology, Deployment Automation, Bicep, Terraform - **Published:** August 25, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/lead-devops-engineer-application-it-operations/10834453 ## About the Role * Extensive, lead-level experience in DevOps engineering and AppOps, with a focus on operating critical applications in enterprise cloud environments (Azure and/or AWS). * Deep technical knowledge of cloud infrastructure services and components, including networking, load balancers, DNS, SSL/TLS certificates, storage, and messaging services. * Hands-on expertise with enterprise observability stacks (such as Datadog, Grafana, Prometheus, ELK/OpenSearch, and OpenTelemetry), alert engineering, and log/metric/trace analysis. * Solid practical understanding of continuous integration and continuous deployment (CI/CD) pipelines, multi-environment application lifecycles (Dev, QA, UAT, Prod), and validation strategies in lower/production environments. * Proficiency in deployment strategies (including blue/green, rolling, canary) and traffic management. * Advanced scripting and automation skills using PowerShell, Bash, or Python to develop runbooks, health checks, self-healing, and remediation workflows. * Strong command of Infrastructure as Code (IaC), specifically with Terraform (modules, workspaces), Azure ARM/Bicep, or AWS CloudFormation, including environment drift detection and policy-as-code. * Practical understanding of security and compliance protocols in operations, including secrets and key management, vulnerability remediation, least-privilege access, and audit readiness. * Proven track record of engineering leadership, stakeholder management, technical mentorship, and leading incident response / RCA processes in global, fast-paced environments. * Excellent communication and advisory skills, with the ability to translate technical risks into clear business metrics for stakeholder decision-making. * Strong ownership mindset and the ability to ensure 24x7 application reliability and operational excellence. * Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience. Nice to have * Preferred Certifications: Azure Administrator/Architect, AWS SysOps/DevOps Professional, ITIL Foundation (or higher), SRE Foundation, Terraform Associate/Professional, or other industry-recognized DevOps/SRE credentials. ## Description * Mentor AppOps engineers, providing technical guidance, conducting code/review for automation scripting, and developing on-call operational excellence. * Own production reliability for critical applications by defining, tracking, and enforcing SLOs, SLAs, error budgets, and capacity/performance baselines. * Lead major incident response and production triage, driving clear business and technical communications, and ensuring data-driven root cause analysis (RCA) with long-term preventative actions. * Direct release, deployment, and change operations by coordinating application deployments, assessing operational risks, enforcing readiness gates, ensuring compliance with client change processes, and validating post-deployment health to improve change success rates. * Architect and maintain operational observability by designing and implementing enterprise dashboards, alert strategies, log/trace pipelines, and runbook automation for rapid system diagnosis and recovery. * Establish and continuously improve operational standards, guardrails, and runbooks, while automating repeatable workflows and repetitive tasks to systematically reduce manual toil and improve operational efficiency. * Partner cross-functionally with Engineering, CloudOps, Security, and Compliance teams to resolve issues, improve service quality, and consult on resiliency patterns (including circuit breakers, bulkheads, graceful degradation, and retries) and performance tuning. * Plan and execute capacity management, scaling strategies, and Disaster Recovery (DR)/Business Continuity Planning (BCP) readiness, including failover testing and simulated scenario exercises. * Champion security-by-default and compliance alignment in operations by enforcing secrets hygiene, patch/vulnerability remediation, certificate/DNS management, least-privilege access, and general security standard adherence. * Drive service reviews with stakeholders, publishing key operational KPIs (such as MTTR, change success rate, and incident rate) to lead continuous improvement roadmaps. * Monitor application health, availability, and performance across all environments, proactively identifying anomalies, resolving application environment issues, and optimizing runtime behavior. * Participate in on-call rotation responsibilities with the Service Delivery and Operations Team. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [The journey from developer to devops - what i've learnt along the way](https://www.wearedevelopers.com/videos/238-the-journey-from-developer-to-devops-what-i-ve-learnt-along-the-way) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)