> Markdown version of [/jobs/ext/2273724-ai-observability-engineer](https://www.wearedevelopers.com/jobs/ext/2273724-ai-observability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Observability Engineer - **Company:** NCR Voyix Corporation - **Location:** Atlanta, GA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Application Performance Management, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, DevOps, Python (Programming Language), Knowledge Management, Knowledge-Based Systems, Machine Learning, Windows PowerShell, Reliability Engineering, Software Tools, Cloud Services, Azure Machine Learning, Data Logging, Google Cloud, Chatbots, Microsoft Power Automate, Delivery Pipeline, Prompt Engineering, Mttr, Generative AI, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Performance Monitor, Bicep, Virtual Agents, ArcSight Event Correlation, Terraform, Splunk, Appdynamics, Dynatrace, Api Management - **Published:** August 27, 2026 - **Apply:** https://dejobs.org/x/x/5640804F599047A99D6710C5FF0AFA54/job/ ## About the Role * Bachelor's degree in Computer Science, Engineering, Information Technology, or related field. * 3+ years of experience in Observability, Site Reliability Engineering (SRE), Infrastructure Engineering, or Operations Engineering. * Strong experience with monitoring, logging, tracing, and performance management platforms. * Proficiency in scripting and automation using Python, PowerShell, Bash, or similar languages. * Experience working with cloud platforms such as Microsoft Azure, AWS, or Google Cloud. * Knowledge of CI/CD pipelines, Infrastructure as Code (Terraform, Bicep, ARM, etc.), and DevOps practices. * Experience with Python, PowerShell, or similar languages for automation and API integrations. * Experience developing automation workflows and integrating observability platforms with cloud services and operational tools. * Familiarity with AI-assisted engineering practices and Copilot-enabled development workflows. * Strong analytical and troubleshooting skills., * Experience with Microsoft Copilot, Azure OpenAI, Copilot Studio, or other AI-powered engineering tools. * Knowledge of AIOps platforms and machine learning concepts related to observability. * Experience implementing OpenTelemetry standards and distributed tracing solutions. * Familiarity with Kubernetes, containers, and cloud-native monitoring architectures. * Experience building automated remediation and self-healing workflows. * Hands-on experience with Microsoft Copilot, Azure AI Services, Azure OpenAI, Copilot Studio, LangChain, Semantic Kernel, Agentic AI frameworks, or similar technologies. * Experience building AI agents, retrieval-based knowledge systems, AI-powered chatbots, or autonomous operational workflows. * Knowledge of RAG architectures, vector databases, prompt engineering, and AI governance best practices. * Familiarity with AIOps platforms and event intelligence solutions. ## Description * Design, deploy, and maintain enterprise observability platforms for monitoring, logging, tracing, and alerting. * Develop dashboards, KPIs, and service health metrics to provide actionable operational insights. * Implement and optimize observability solutions using tools such as Splunk, AppDynamics, or Splunk Observability Cloud platforms. * Automate operational processes, alert management, health checks, and incident response workflows using scripting and orchestration tools. * Collaborate with engineering and operations teams to improve application performance, reliability, and scalability. * Analyze incidents, identify root causes, and implement preventive measures through proactive monitoring and automation. * Drive adoption of AI-powered observability capabilities, including anomaly detection, predictive analytics, event correlation, and intelligent alerting. * Leverage Microsoft Copilot, Generative AI, and automation technologies to enhance troubleshooting, operational efficiency, and engineering productivity. * Develop AI-assisted runbooks, knowledge bases, and self-healing solutions to reduce manual intervention and Mean Time to Resolution (MTTR). * Participate in on-call support and major incident management activities as needed. * Design and implement AI-driven observability solutions using telemetry, monitoring, logging, and distributed tracing platforms. * Develop automated remediation, self-healing workflows, and operational runbooks using scripting, orchestration, and infrastructure-as-code tools. * Build and integrate Agentic AI solutions that can autonomously analyze alerts, retrieve operational context, recommend actions, and execute approved remediation workflows. * Leverage Microsoft Copilot and Generative AI tools to improve incident investigation, root-cause analysis, knowledge management, and engineering productivity. * Implement AI-powered anomaly detection, event correlation, capacity forecasting, and predictive monitoring capabilities. * Develop integrations between observability platforms and AI agents to automate repetitive operational tasks and improve MTTR. * Collaborate with application, SRE, cloud, and platform teams to identify opportunities for AI-assisted operations and process automation. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) - [Agentic DevOps: How AI-Powered Automation Transforms Software Delivery on GitHub and Azure](https://www.wearedevelopers.com/videos/1539-agentic-devops-how-ai-powered-automation-transforms-software-delivery-on-github-and-azure) - [The AI-Ready Stack: Rethinking the Engineering Org of the Future](https://www.wearedevelopers.com/videos/1706-the-ai-ready-stack-rethinking-the-engineering-org-of-the-future) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [GitHub Copilot: Beyond the Basics – 10 Ways to Elevate Your Coding](https://www.wearedevelopers.com/magazine/524-github-copilot-beyond-the-basics-10-ways-to-elevate-your-coding)