AWS SRE / Observability Engineer focused on AI...

Insight Global
Charlotte, NC, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Software as a Service Monitoring of Systems Log Analysis Reliability Engineering Software Engineering Datadog Data Logging Prompt Engineering Software Troubleshooting
+14 more
Technical Debt Infrastructure as Code (IaC) Gitlab Git Front End Software Development Virtual Agents Api Design Api Gateway Terraform Splunk Software Version Control Dynatrace Serverless Computing Microservices

Job description

Insight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows.

This position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability.

In addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes.

This is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities.

-Monitor and support AWS application and infrastructure environments

-Plan and execute application and infrastructure configuration changes

-Respond to production incidents, critical outages, and operational emergencies

-Triage application and infrastructure issues using monitoring and observability platforms

-Analyze logs, metrics, traces, and alerts to identify root cause and operational impact

-Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related

-Lead incident response efforts and coordinate cross-functional resolution activities

-Develop and maintain postmortems, operational documentation, and technical runbooks

-Partner with software engineering teams to improve reliability, resiliency, and performance

-Enhance monitoring, logging, alerting, and observability capabilities across the environment

-Drive faster issue detection and resolution through monitoring best practices

-Support emerging agentic AI solutions being deployed across operational workflows

-Perform human-in-the-loop validation of AI-generated recommendations and outputs

-Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness

-Collaborate with AI, development, and operations teams to improve operational intelligence

-Develop automation solutions that reduce manual effort and improve operational efficiency

-Participate in system design reviews, capacity planning initiatives, and architectural discussions

-Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence

Requirements

Terraform expertise for Infrastructure as Code (IaC)

Strong AWS experience, particularly with EC2 and cloud-native services

Experience supporting AWS environments, including API Gateway technologies

Strong experience with Datadog monitoring and observability

Experience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk

GitLab and source control experience using Git

Strong troubleshooting, incident response, and production support experience

Strong log analysis experience within enterprise monitoring environments

Ability to quickly triage production issues and determine root cause using observability tools

Observability engineering experience, including identifying performance bottlenecks and system issues

Experience leading incident response efforts and supporting application teams

Experience creating and facilitating blameless postmortems

Ability to improve logging strategies, operational runbooks, and reduce technical debt

Experience supporting high-availability production applications and infrastructure

Experience supporting AI-powered operational workflows with human oversight and validation

Strong analytical mindset with the ability to rapidly understand system behavior and operational impact

Strong collaboration and communication skills when working across development, operations, and platform teams Experience with Agentic AI systems or AI-enabled operational platforms

Prompt engineering experience, including creating and optimizing AI prompts

Experience working with Anthropic Claude or comparable enterprise AI models

Experience evaluating and validating AI-generated outputs

Experience supporting AI observability and AI Operations initiatives

API development experience

Experience building and supporting microservices architectures

Front-end development experience

Automation and platform engineering experience

Capacity planning and system design consulting experience

Experience promoting observability best practices across engineering organizations

Previous experience in large-scale cloud environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Videos

See all

Related articles

See all