SRE AI Engineer

SWIFT
Charlotte, NC, United States
17 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$188,200.0 - $227,700.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Artificial Intelligence Amazon Web Services Application Performance Management Microsoft Azure Cloud Computing Databases Continuous Integration Cursor (Graphical User Interface Elements) DevOps Middleware
+30 more
Monitoring of Systems Information Technology Operations Nagios Network Control Open Source Technology Ping (Networking Utility) Redis Reliability Engineering Site Reliability Engineering Practices Prometheus Runbook Session Management Systems Integration Google Cloud Retrieval-Augmented Generation Large Language Models Grafana Multi-Agent Systems Siteminder Kubernetes Low Latency SolarWinds (Software) Performance Monitor Virtual Agents BIG-IP Access Policy Manager (APM) Kibana Splunk Cisco Switches GPT Dynatrace

Job description

We are currently seeking a highly skilled SRE hands-on AI Engineer with solid experience in AI Observability and instrumentation approaches for AI systems, AI Agents development to perform detection, diagnosis and autonomous self-healing operates on AI Control Plane. Observability data collection and automation to help lead transformational initiatives within IT operations, encompassing development as well. As a crucial figure in this role, you will participate/help with various technology domain groups and cross functional teams on unified observability gap analysis and solutioning (automation and manual fixes) Responsibilities:

  • Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and DevOps/deployment processes.
  • Experience building agentic workflows using LLMs, tool-calling, function-calling, multi-agent orchestration, and event-driven automation.
  • Experience with Agent-to-Agent communication, AI agent federation, and enterprise AI control-plane concepts.
  • Experience implementing AI control-plane governance, including policy-based execution, approval workflows, audit trails, guardrails, and risk-based remediation controls.
  • Expertise in Observability as a service, Dashboard as a services, monitoring as a services and alert as a service in all technology domains (application, infrastructure, database, security, middleware, network etc.,) Telemetry data collection using Dynatrace APM, SolarWinds, CISCO Switches, F5, Databases, Open-Source tools (Prometheus and Grafana), Log Aggregations (Kibana or Splunk) and AIOPS Tools.
  • Practical experience implementing Golden Signals (latency, traffic, errors, saturation) using related telemetry sources.
  • Configure application performance monitoring (APM), infrastructure monitoring, synthetic monitoring, RUM, and log monitoring.
  • Integrate Dynatrace with CI/CD pipelines, alerting tools, ITSM systems, and incident automation frameworks.
  • Tune alert thresholds, baselines, and AI-driven anomaly detection to reduce noise and improve actionable insights.
  • Deeper understanding of Login authentication mechanisms using Ping, ForgeRock and SiteMinder technologies (session management and cookie management)
  • Define best practices and principles for SRE, including monitoring, alerting, and automation.
  • Collaborate with development teams on resiliency to ensure that services and applications are designed with operational reliability in mind.
  • Implement monitoring systems to assess the performance of applications and infrastructure and proactively identifying areas for optimization.
  • Ability to develop close relationship with other operational teams to integrate SRE practices and drive overall operational improvements across enterprise.
  • Stay up to date on industry trends, new technologies, and best practices in SRE and applying relevant advancements to the organization.
  • Ability to build strong working relationships across different levels, client focus mindset.

Requirements

  • Around 7-10 years of SRE hands on experience with AI OPS, cloud technologies, development, SRE toolsets and automation
  • Hands-on experience implementing Retrieval-Augmented Generation using approved enterprise knowledge sources such as runbooks, SOPs, RCA documents, incident history, architecture documents, and knowledge articles.
  • Expertise SPEC driven and Prompting using Ai IDE tools Cursor, Kiro and Good Antigravity
  • Hands-on experience AI LLM’s GPT - Open AI, Claude, Gemini etc.,
  • Experience with LangChain, LangGraph, Bedrock Agents, Azure AI Foundry, or Vertex AI Agent Builder
  • Experience with vector databases such as OpenSearch, Pinecone, Chroma, Redis Vector, or pgvector.
  • Experience performing Observability current-state assessments, gap analysis and solutioning (automation and manual fixes) in all technology domains (application, infrastructure, database, security, middleware, network etc.,),
  • Strong hands-on automation experience in Observability as a code, dashboard as a code, monitoring as a code, alert as a code (Instrumentation, templates, automatic deployment, visualization and alerting)
  • Strong hands-on experience with any Cloud Technology (AWS): Control Tower, Project Setup, Creating Accounts, RDS, SSO
  • Monitoring & alerting setup experience with Splunk, Prometheus, Grafana, Kibana, ELK, with pref. for APM (Dynatrace).
  • Strong skills in APM, distributed tracing, synthetic & real user monitoring, log monitoring, and Davis AI configuration
  • Own the design, configuration, CICD deployment, and optimization for enterprise-wide observability tools.
  • Experience integrating, automation, and cloud platforms (AWS, Azure, GCP).
  • Extended experience instrumenting OTEL Framework.
  • Hands on experience with Dynatrace Plug-and-play observability modules (OKit) development for Observability Developers Java and .Net applications.
  • Define monitoring standards, best practices, and governance to ensure consistency and scalability.
  • Experience to deploy and tune OneAgent, build end-to-end PurePath tracing, and leverage Smartscape topology for proactive performance monitoring and root-cause analysis.
  • Collaborate with application and infrastructure teams to troubleshoot performance issues and implement permanent fixes.

Good to have:

  • Any of the relevant professional certifications AIOPS related certifications, Certified Site Reliability Engineer (CSRE), Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer Professional, , Google Cloud Professional; DevOps Engineer

Benefits & conditions

  • $130,000-160,000 per year

About the company

Semrush Inc.

  • Austin, TX
  • $188,200-227,700 per year Semrush is a brand visibility platform, empowering marketers to command their online presence and create measurable impact. We unify SEO authority and AI visibility, so brands are …

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all