> Markdown version of [/jobs/ext/2003425-aws-sre-observability-engineer-focused-on-ai](https://www.wearedevelopers.com/jobs/ext/2003425-aws-sre-observability-engineer-focused-on-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS SRE / Observability Engineer focused on AI... - **Company:** Insight Global - **Location:** Charlotte, NC, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Software as a Service, Monitoring of Systems, Log Analysis, Reliability Engineering, Software Engineering, Datadog, Data Logging, Prompt Engineering, Software Troubleshooting, Technical Debt, Infrastructure as Code (IaC), Gitlab, Git, Front End Software Development, Virtual Agents, Api Design, Api Gateway, Terraform, Splunk, Software Version Control, Dynatrace, Serverless Computing, Microservices - **Published:** August 9, 2026 - **Apply:** https://www.juju.com/job/00000000gm953h ## About the Role Terraform expertise for Infrastructure as Code (IaC) Strong AWS experience, particularly with EC2 and cloud-native services Experience supporting AWS environments, including API Gateway technologies Strong experience with Datadog monitoring and observability Experience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk GitLab and source control experience using Git Strong troubleshooting, incident response, and production support experience Strong log analysis experience within enterprise monitoring environments Ability to quickly triage production issues and determine root cause using observability tools Observability engineering experience, including identifying performance bottlenecks and system issues Experience leading incident response efforts and supporting application teams Experience creating and facilitating blameless postmortems Ability to improve logging strategies, operational runbooks, and reduce technical debt Experience supporting high-availability production applications and infrastructure Experience supporting AI-powered operational workflows with human oversight and validation Strong analytical mindset with the ability to rapidly understand system behavior and operational impact Strong collaboration and communication skills when working across development, operations, and platform teams Experience with Agentic AI systems or AI-enabled operational platforms Prompt engineering experience, including creating and optimizing AI prompts Experience working with Anthropic Claude or comparable enterprise AI models Experience evaluating and validating AI-generated outputs Experience supporting AI observability and AI Operations initiatives API development experience Experience building and supporting microservices architectures Front-end development experience Automation and platform engineering experience Capacity planning and system design consulting experience Experience promoting observability best practices across engineering organizations Previous experience in large-scale cloud environments ## Description Insight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows. This position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability. In addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes. This is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities. -Monitor and support AWS application and infrastructure environments -Plan and execute application and infrastructure configuration changes -Respond to production incidents, critical outages, and operational emergencies -Triage application and infrastructure issues using monitoring and observability platforms -Analyze logs, metrics, traces, and alerts to identify root cause and operational impact -Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related -Lead incident response efforts and coordinate cross-functional resolution activities -Develop and maintain postmortems, operational documentation, and technical runbooks -Partner with software engineering teams to improve reliability, resiliency, and performance -Enhance monitoring, logging, alerting, and observability capabilities across the environment -Drive faster issue detection and resolution through monitoring best practices -Support emerging agentic AI solutions being deployed across operational workflows -Perform human-in-the-loop validation of AI-generated recommendations and outputs -Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness -Collaborate with AI, development, and operations teams to improve operational intelligence -Develop automation solutions that reduce manual effort and improve operational efficiency -Participate in system design reviews, capacity planning initiatives, and architectural discussions -Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Unlocking the AI Black Box: Building Trust in the Era of Agentic Production](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)