AWS SRE / Observability Engineer focused on AI...
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
Insight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows.
This position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability.
In addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes.
This is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities.
-Monitor and support AWS application and infrastructure environments
-Plan and execute application and infrastructure configuration changes
-Respond to production incidents, critical outages, and operational emergencies
-Triage application and infrastructure issues using monitoring and observability platforms
-Analyze logs, metrics, traces, and alerts to identify root cause and operational impact
-Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related
-Lead incident response efforts and coordinate cross-functional resolution activities
-Develop and maintain postmortems, operational documentation, and technical runbooks
-Partner with software engineering teams to improve reliability, resiliency, and performance
-Enhance monitoring, logging, alerting, and observability capabilities across the environment
-Drive faster issue detection and resolution through monitoring best practices
-Support emerging agentic AI solutions being deployed across operational workflows
-Perform human-in-the-loop validation of AI-generated recommendations and outputs
-Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness
-Collaborate with AI, development, and operations teams to improve operational intelligence
-Develop automation solutions that reduce manual effort and improve operational efficiency
-Participate in system design reviews, capacity planning initiatives, and architectural discussions
-Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence
Requirements
Terraform expertise for Infrastructure as Code (IaC)
Strong AWS experience, particularly with EC2 and cloud-native services
Experience supporting AWS environments, including API Gateway technologies
Strong experience with Datadog monitoring and observability
Experience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk
GitLab and source control experience using Git
Strong troubleshooting, incident response, and production support experience
Strong log analysis experience within enterprise monitoring environments
Ability to quickly triage production issues and determine root cause using observability tools
Observability engineering experience, including identifying performance bottlenecks and system issues
Experience leading incident response efforts and supporting application teams
Experience creating and facilitating blameless postmortems
Ability to improve logging strategies, operational runbooks, and reduce technical debt
Experience supporting high-availability production applications and infrastructure
Experience supporting AI-powered operational workflows with human oversight and validation
Strong analytical mindset with the ability to rapidly understand system behavior and operational impact
Strong collaboration and communication skills when working across development, operations, and platform teams Experience with Agentic AI systems or AI-enabled operational platforms
Prompt engineering experience, including creating and optimizing AI prompts
Experience working with Anthropic Claude or comparable enterprise AI models
Experience evaluating and validating AI-generated outputs
Experience supporting AI observability and AI Operations initiatives
API development experience
Experience building and supporting microservices architectures
Front-end development experience
Automation and platform engineering experience
Capacity planning and system design consulting experience
Experience promoting observability best practices across engineering organizations
Previous experience in large-scale cloud environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.juju.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
MLOps And AI Driven Development
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
What is Software Engineering in the Age of AI?