AI DevOps Engineer 1

Barton Malow
Southfield, MI, United States
3 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Automation of Tests Cloud Computing Continuous Integration Software Debugging DevOps Github Prometheus Datadog Data Logging Cloud Platform System
+11 more
Large Language Models Grafana Multi-Agent Systems Gitlab-ci Kubernetes Cloudwatch Terraform Dynatrace Serverless Computing Docker Databricks

Job description

We are hiring an AI DevOps Engineer to own how our AI systems get deployed, stay up, and stay observable. You will own the infrastructure and operational substrate for Barton Malow’s agentic platform - the CI/CD pipelines, the cloud environments, the release machinery, the reliability practice, and the telemetry layer that turns opaque agent behavior into something you can measure, alert on, and debug.

Our engineers build the AI agents; you make shipping them safe, repeatable, and observable. Right now most of our apps have no CI/CD, the infrastructure is fragmented, and the agentic systems being stood up have little more than print statements for observability. Building that operational foundation is the job.

Agentic systems fail differently from ordinary services - nondeterministic output, silent quality drift, runaway tool-call loops, and cost that spikes without warning. Standard DevOps is necessary but not sufficient. This role exists because someone has to own reliability and telemetry for systems that don’t fail the way the runbooks assume.

What you’ll do

Own CI/CD for the platform. Build the pipelines that take AI systems from commit to production - automated testing, evaluation gates, security and dependency checks, controlled and canary releases, and one-command rollback. Most of our apps have no pipeline today; you will establish the pattern and make the safe path the default path.

Manage cloud infrastructure as code. Own the cloud environments (AWS primarily) the platform runs on - provisioning, networking, secrets, environment parity, and cost controls - as versioned, reviewable infrastructure-as-code, not hand-tuned consoles. You are accountable for environments that are reproducible, least-privilege by default, and cheap to stand up and tear down.

Run the reliability practice. Own production reliability: SLOs, on-call and incident response, capacity and cost management, self-healing loops that detect and recover from failures, and blameless post-incident review. You will help define what “up” and “healthy” even mean for a nondeterministic system.

Build the agent telemetry and observability layer. Instrument the platform so agent behavior is legible: structured traces of agent runs and tool calls, token and cost accounting, latency and success metrics, output-quality tracking over time, and the dashboards and alerts that surface a regression before a user does. When an agent misbehaves in production, the telemetry you built is how the team finds out and figures out why.

Set the operational standard by example. On a small, high-leverage team, your pipelines and dashboards are the template. You establish the deployment patterns others adopt, the observability every new system gets wired into by default, and the operational discipline that lets a lean team run production systems well.

Requirements

  • 6-8 years in DevOps, SRE, platform, or infrastructure engineering, with a track record of running production systems you were accountable for
  • Deep CI/CD experience - you have built and owned pipelines (GitHub Actions, GitLab CI, or similar) that gate, test, and safely release real production software
  • Strong cloud operations, ideally AWS - provisioning, networking, secrets, and cost management as infrastructure-as-code (Terraform, CDK, or similar)
  • Hands-on observability experience - metrics, logging, distributed tracing, dashboards, and alerting (OpenTelemetry, Prometheus/Grafana, Datadog, CloudWatch, or similar) - and the instinct to instrument first
  • SRE fundamentals: SLOs, incident response, on-call, capacity planning, and blameless postmortems
  • Enough software fluency to read application code, wire telemetry into it, and debug a failing deploy without waiting for someone else

Nice to have

  • Experience operating AI or LLM systems in production - token and cost accounting, prompt and evaluation-score tracking, or LLM observability tooling (LangSmith, Langfuse, Arize, or similar)
  • Familiarity with the failure modes of nondeterministic systems: quality drift, runaway loops, cost spikes, non-reproducible output
  • Experience with Databricks or a similar lakehouse platform, and with tool-integration layers such as MCP
  • Container and orchestration experience (Docker, Kubernetes, or serverless equivalents)
  • Experience in a non-software-company engineering organization - internal tools, corporate IT transformation, or similar
  • Experience standing up an internal platform or golden-path deployment pattern that other teams adopted

About the company

Barton Malow is a builder. For over 100 years we have delivered some of the most complex construction projects in North America - schools, hospitals, stadiums, manufacturing plants, and industrial facilities. Today we are doing something most construction companies are not: rebuilding how we build software, with AI at the center. Our APEX team owns the AI and engineering platform that the rest of the company builds on, and we are investing seriously in it.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all