> Markdown version of [/jobs/ext/3029835-principal-ai-operations-infrastructure-automation](https://www.wearedevelopers.com/jobs/ext/3029835-principal-ai-operations-infrastructure-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal, AI Operations & Infrastructure Automation - **Company:** Idc Inc - **Location:** Boston, MA, United States (Remote available) - **Experience:** Experienced - **Salary:** $128,320.0 - $160,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, JIRA, Audit Trail, Microsoft Azure, System Configuration, Continuous Delivery, Continuous Integration, Noise Reduction, Information Technology Operations, Python (Programming Language), Reliability Engineering, Prometheus, User Provisioning Software, Datadog, Large Language Models, Grafana, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Virtual Agents, Restful APIs, Terraform, Splunk, New Relic (SaaS), Dynatrace, Human in the Loop, Servicenow - **Published:** September 22, 2026 - **Apply:** https://www.jofdav.com/jobs/59820423-principal-ai-operations-infrastructure-automation ## About the Role * 8+ years in infrastructure, platform engineering, or IT operations, including 3+ years leading automation or AIOps initiatives. * Demonstrated track record of redesigning an operations function around automation - not just introducing point tools but changing how a team works. * Deep hands-on experience with cloud infrastructure automation (AWS strongly preferred), infrastructure-as-code (Terraform, CloudFormation, or similar), and CI/CD pipelines. * Working knowledge of AIOps and agentic automation platforms, and a clear point of view on where AI genuinely improves operational outcomes versus adds noise. * Practical automation skill in Python or an equivalent language, including building integrations against REST APIs and operational tooling. * Hands-on experience with observability and monitoring platforms - Datadog, Splunk, Dynatrace, New Relic, Prometheus/Grafana, Elastic, or similar - including instrumenting services and building useful alerting. * Strong program leadership skills: able to define a roadmap, sequence dependencies, and hit aggressive milestones. * Excellent executive communication skills - comfortable distilling complex technical transformation into clear, concise updates for senior leadership. Preferred * Experience operating in a regulated or compliance-sensitive environment (SOC 2, ISO 27001, or similar). * Prior experience partnering directly with a hyperscaler (AWS, Azure, or GCP) on joint automation or AI initiatives. * Background in site reliability engineering (SRE) practices and observability tooling. * Experience with agentic AI frameworks, RAG pipelines, or building internal AI copilots for technical teams. * Familiarity with ITSM platforms and process - ServiceNow, Jira Service Management, or similar - and comfort automating against them. * Familiarity with FinOps practices and cloud cost automation. * Relevant certifications (cloud architect or engineer, Kubernetes, ITIL v4, Terraform). ## Description You'll personally write and review automation, get into the tooling directly, and lead by example for the team you manage. You'll partner closely with the CIO, Cyber & Infrastructure leadership, and AWS as a strategic cloud and co-development partner to re-architect how infrastructure operations get done. This role builds and leads the team responsible for automation - setting technical direction, reviewing their work, and growing their skills as the operating model evolves. Additional responsibilities include: * Own the end-to-end redesign of the infrastructure operating model, moving core workflows (provisioning, monitoring, incident response, patching, capacity management) from manual execution to automated, self-healing pipelines. * Define a phased transformation roadmap with clear milestones, automation-coverage targets, and go-live gates - and drive execution against it at pace. * Lead change management for the infrastructure team through the transition - defining new roles, skills, and ways of working as automation is adopted. * Build the AIOps layer; evaluate, select, and deploy AIOps platforms and agentic tooling - anomaly detection, predictive alerting, automated remediation - suited to IDC's AWS-centric environment. Implement event correlation and noise reduction so a hundred alerts resolve into one actionable incident with a probable root cause attached. * Stand up LLM (large language model)-assisted triage that classifies, enriches, and routes incidents, pulling relevant runbooks, past resolutions, and change history into the ticket automatically. * Develop predictive capacity and failure models for critical services, and act on them before thresholds are breached. * Automate end to end, build and champion an infrastructure-as-code (IaC) standard across the organization, converting legacy manual processes into version-controlled, repeatable automation. * Deliver self-healing automation and closed-loop remediation for the highest-volume recurring incidents - restart, rollback, scale, reroute, and verify without a human in the loop. Automate service request fulfillment: access provisioning, environment setup, onboarding and offboarding, and routine change execution. * Own CI/CD (continuous integration/continuous delivery-deployment) pipelines for operational tooling, including testing and safe rollback for automations that touch production. * Lead the team and prove the results Directly manage and develop a team of infrastructure and automation engineers - hiring, coaching, and setting technical direction - while staying hands-on in the tooling and code yourself. * Partner with AWS and other strategic vendors to pilot and scale automation and AI-driven operations capabilities, including Bedrock/AgentCore-based tooling where relevant. * Establish new operational KPIs - automation coverage, mean time to detect and resolve, deployment frequency, toil eliminated - and report progress in concise, executive-ready formats. * Set guardrails for AI in operations: human-in-the-loop thresholds, blast-radius limits, audit logging, evaluation of model output quality, and clear rollback paths. * Ensure automated operations meet IDC's security, compliance, and resiliency standards, working closely with Cyber & Compliance., Measurable improvement in mean time to detect and resolve incidents. A growing share of routine operational workflows runs without manual intervention, and infrastructure changes are consistently deployed through version-controlled, automated pipelines. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)