> Markdown version of [/jobs/ext/2303694-senior-ai-observability-engineer](https://www.wearedevelopers.com/jobs/ext/2303694-senior-ai-observability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Observability engineer - **Company:** Lam Research International Holding Company - **Location:** Fremont, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $92,000.0 - $211,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Border Gateway Protocol, Cloud Computing, Configuration Management Databases, Configuration Management, Encodings, Databases, Continuous Integration, Data Centers, Data Deduplication, Noise Reduction, DevOps, Domain Name System (DNS), Multi-protocol Systems, Machine Learning, Network Virtualization, Open Shortest Path First (OSPF), Reliability Engineering, Software Safety, TCP/IP, Virtual Local Area Networks, Wide Area Networks, Private Cloud Environment, Data Logging, JavaScript Pagination Plugin, Google Cloud, Load Balancing, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), AI Platforms, Kubernetes, Information Technology, Machine Learning Operations, ArcSight Event Correlation, Software Version Control, Serverless Computing, Servicenow - **Published:** August 30, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/86123874/1 ## About the Role * BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent practical experience. * Eight or more years in SRE, DevOps, infrastructure, observability, network engineering, or platform engineering, with a track record of shipping production systems yourself. * Production experience with LLM applications: prompt engineering, retrieval-augmented generation, embeddings and vector databases, function and tool calling, and agent orchestration. * Practical use of machine learning for anomaly detection, forecasting, event correlation, and alert noise reduction on real operational telemetry. * Working knowledge of at least one major AI platform and the ability to remain portable across them, including private-network deployment, quota, and cost management. * Hands-on production experience with at least one major public cloud and the ability to design portable, cloud-agnostic patterns across the others, covering identity and access boundaries, compute, storage, managed databases, serverless functions, and event services. * Solid on-premises infrastructure background across virtualization, storage, data-center operations, and private cloud platforms. * Strong networking fundamentals spanning both worlds: TCP/IP, BGP and OSPF, VLANs and overlays such as VXLAN and EVPN, MPLS and SD-WAN, firewalls, load balancers, DNS, and cloud virtual network, VPC, and transit routing constructs. You can read a flow log or a routing table and reason about a failure. ## Description We are seeking a hands-on Senior AIOps Reliability Engineer to build the AI-native operations layer for our hybrid enterprise estate. You will design and ship LLM-based agents, retrieval-augmented knowledge pipelines, and machine-learning anomaly detection that find, triage, and remediate incidents across public cloud, private cloud, and on-premises data centers, all resting on strong SRE and network engineering fundamentals. Our environment spans Azure, AWS, and Google Cloud alongside on-premises data centers, colocation sites, and manufacturing and HPC facilities, so cloud-agnostic design and hybrid network fluency matter more than depth in any single provider. This is a deep individual-contributor role: you will write the code, instrument the telemetry, tune the models, and own the reliability of both the infrastructure and the AI systemsoperatingon it. The impact you'll make Join Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure. As a crucial member of our IT team, you'll contribute your technical assistance and guidance to projects for various systems and infrastructures. Acting as a technical liaison, you'll address complex business problems with automated systems solutions. Your expertise will be instrumental in driving Lam's commitment to innovation and efficiency. What you'll do AI and Agentic Operations * Build agentic AI workflows using LLM agents, tool and function calling, and orchestration frameworks such as LangGraph, Semantic Kernel, AutoGen, or the Model Context Protocol, applied to autonomous fault detection, triage, and remediation. * Develop the AIOps intelligence layer: time-series anomaly detection, dynamic baselining, alert deduplication and correlation, event clustering, and predictive failure and capacity forecasting across infrastructure, network, and application telemetry from both cloud and on-premises sources. * Engineer the retrieval knowledge fabric by chunking, embedding, and indexing runbooks, post-mortems, architecture documents, CMDB and ServiceNow records into a vector store, then tuning retrieval quality against measurable evaluations. * Ship AI-assisted incident response: automated summarization, root-cause hypothesis generation, blast-radius analysis, and telemetry-grounded draft post-mortems wired into the paging and ITSM toolchain. * Automate remediation safely through event-driven pipelines and configuration-management runbooks invoked by agents, with human-in-the-loop approval gates, scoped least-privilege boundaries, rollback paths, and complete audit trails for every autonomous action. * Own AI safety and governance in production: guardrails, prompt-injection defense, hallucination and drift monitoring, PII redaction, and evaluation harnesses that gate every model or prompt change. * Run LLMOps and MLOps, covering prompt and model versioning, offline and online evaluation, shadow and A/B testing, inference logging, token cost and latency observability, and CI/CD for every AI component. * Instrument AI systems as first-class services with OpenTelemetry GenAI tracing, model SLOs, and quality, cost, and latency dashboards for every agent in production. ## Related Videos - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Mastering AI-Driven Problem Solving in Engineering with Observability](https://www.wearedevelopers.com/videos/994-mastering-ai-driven-problem-solving-in-engineering-with-observability) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)