AI Observability Engineer

NCR Voyix Corporation
Atlanta, GA, United States
1 day ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Application Performance Management Microsoft Azure Bash Shell Cloud Computing Cloud Engineering DevOps Python (Programming Language) Knowledge Management Knowledge-Based Systems Machine Learning
+25 more
Windows PowerShell Reliability Engineering Software Tools Cloud Services Azure Machine Learning Data Logging Google Cloud Chatbots Microsoft Power Automate Delivery Pipeline Prompt Engineering Mttr Generative AI Kubernetes Infrastructure Automation Frameworks Information Technology Performance Monitor Bicep Virtual Agents ArcSight Event Correlation Terraform Splunk Appdynamics Dynatrace Api Management

Job description

  • Design, deploy, and maintain enterprise observability platforms for monitoring, logging, tracing, and alerting.
  • Develop dashboards, KPIs, and service health metrics to provide actionable operational insights.
  • Implement and optimize observability solutions using tools such as Splunk, AppDynamics, or Splunk Observability Cloud platforms.
  • Automate operational processes, alert management, health checks, and incident response workflows using scripting and orchestration tools.
  • Collaborate with engineering and operations teams to improve application performance, reliability, and scalability.
  • Analyze incidents, identify root causes, and implement preventive measures through proactive monitoring and automation.
  • Drive adoption of AI-powered observability capabilities, including anomaly detection, predictive analytics, event correlation, and intelligent alerting.
  • Leverage Microsoft Copilot, Generative AI, and automation technologies to enhance troubleshooting, operational efficiency, and engineering productivity.
  • Develop AI-assisted runbooks, knowledge bases, and self-healing solutions to reduce manual intervention and Mean Time to Resolution (MTTR).
  • Participate in on-call support and major incident management activities as needed.
  • Design and implement AI-driven observability solutions using telemetry, monitoring, logging, and distributed tracing platforms.
  • Develop automated remediation, self-healing workflows, and operational runbooks using scripting, orchestration, and infrastructure-as-code tools.
  • Build and integrate Agentic AI solutions that can autonomously analyze alerts, retrieve operational context, recommend actions, and execute approved remediation workflows.
  • Leverage Microsoft Copilot and Generative AI tools to improve incident investigation, root-cause analysis, knowledge management, and engineering productivity.
  • Implement AI-powered anomaly detection, event correlation, capacity forecasting, and predictive monitoring capabilities.
  • Develop integrations between observability platforms and AI agents to automate repetitive operational tasks and improve MTTR.
  • Collaborate with application, SRE, cloud, and platform teams to identify opportunities for AI-assisted operations and process automation.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field.
  • 3+ years of experience in Observability, Site Reliability Engineering (SRE), Infrastructure Engineering, or Operations Engineering.
  • Strong experience with monitoring, logging, tracing, and performance management platforms.
  • Proficiency in scripting and automation using Python, PowerShell, Bash, or similar languages.
  • Experience working with cloud platforms such as Microsoft Azure, AWS, or Google Cloud.
  • Knowledge of CI/CD pipelines, Infrastructure as Code (Terraform, Bicep, ARM, etc.), and DevOps practices.
  • Experience with Python, PowerShell, or similar languages for automation and API integrations.
  • Experience developing automation workflows and integrating observability platforms with cloud services and operational tools.
  • Familiarity with AI-assisted engineering practices and Copilot-enabled development workflows.
  • Strong analytical and troubleshooting skills., * Experience with Microsoft Copilot, Azure OpenAI, Copilot Studio, or other AI-powered engineering tools.
  • Knowledge of AIOps platforms and machine learning concepts related to observability.
  • Experience implementing OpenTelemetry standards and distributed tracing solutions.
  • Familiarity with Kubernetes, containers, and cloud-native monitoring architectures.
  • Experience building automated remediation and self-healing workflows.
  • Hands-on experience with Microsoft Copilot, Azure AI Services, Azure OpenAI, Copilot Studio, LangChain, Semantic Kernel, Agentic AI frameworks, or similar technologies.
  • Experience building AI agents, retrieval-based knowledge systems, AI-powered chatbots, or autonomous operational workflows.
  • Knowledge of RAG architectures, vector databases, prompt engineering, and AI governance best practices.
  • Familiarity with AIOps platforms and event intelligence solutions.

About the company

NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide., Integrated into our shared values is NCR Voyix’s commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law. NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party AgenciesTo ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

β€œWhen applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.com email domain.”

Help us run the world’s top brands.

At NCR Voyix (http://www.ncr.com/) , we specialize in turning routine transactions into meaningful connections. With a rich history (http://www.ncr.com/about/history) of innovation, we’ve been at the forefront of problem-solving through technology. Operating globally in over 30 countries, we lead in Retail, Restaurant, Digital banking, and Payments. Our solutions optimize banking operations, streamline restaurant services, enhance retail interactions, and foster trust through secure payment systems.

We take pride in our strong culture (http://www.ncr.com/about) and a history of providing robust career paths. Come work for a leading technology company where you can grow your career. Join us and be part of revolutionizing transactions across these pivotal industries.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou Β· Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 Β· LIVE

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 Β· World Congress 2022

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley Β· World Congress 2021

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell Β· Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 Β· LIVE

Videos

See all

Related articles

See all