> Markdown version of [/jobs/ext/1505975-senior-engineer-ml](https://www.wearedevelopers.com/jobs/ext/1505975-senior-engineer-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Engineer - ML - **Company:** Albertsons Companies - **Location:** Pleasanton, CA, United States - **Experience:** Expert - **Salary:** $109,700.0 - $156,838.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Cloud Engineering, Configuration Management Databases, Information Systems, Continuous Integration, Data Cleansing, Data Deduplication, Information Engineering, Noise Reduction, Graph Database, Monitoring of Systems, Information Technology Operations, Python (Programming Language), Machine Learning, Automation of Marketing, Neo4j, NumPy, Tensorflow, Prometheus, Azure Machine Learning, Software Engineering, Systems Integration, Management of Software Versions, Enterprise Search, Feature Engineering, Pytorch, Large Language Models, Grafana, Multi-Agent Systems, Prompt Engineering, Backend, Pandas, AI Platforms, Kubernetes, Information Technology, Operational Systems, Machine Learning Operations, Restful APIs, Splunk, Appdynamics, GPT, Software Version Control, Docker, Servicenow, Microservices - **Published:** July 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1401373563ae5662 ## About the Role * Strong experience designing and building AI/ML systems for real-world production use cases. * Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, and NumPy. * Experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, or recommendation problems. * Strong understanding of time-series modeling techniques for forecasting and operational prediction use cases.Experience with alert classification, incident prediction, event deduplication, prioritization, or * signal correlation in observability or IT operations contexts. * Knowledge of causal ML, causal inference, graph-based reasoning, and dependency-aware analysis techniques for RCA and impact analysis. * Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph. * Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, or assistant workflows. * Strong understanding of prompt engineering, RAG architecture, embeddings, vector stores, tool calling, memory handling, and agent evaluation techniques. * Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, and operational systems. * Experience building backend services and APIs for AI/ML model inference and agent orchestration. * Good understanding of observability data such as logs, metrics, traces, topology, incidents, and alerts. * Experience with data engineering concepts including feature engineering, data preprocessing, model pipelines, and batch or streaming inference. * Familiarity with graph databases such as Neo4j and their use in dependency mapping, causal analysis, and knowledge-driven AI systems. * Experience with REST APIs, microservices architecture, Docker, Kubernetes, and cloud-native deployment patterns. * Familiarity with CI/CD, MLOps, model lifecycle management, experiment tracking, and model versioning practices. * Knowledge of OpenTelemetry, monitoring systems, and observability platforms is highly desirable. * Strong understanding of software engineering fundamentals, system design, and scalable architecture patterns. * Strong analytical and problem-solving skills, with the ability to convert ambiguous operational problems into measurable AI/ML solutions. * Excellent communication and collaboration skills to work with SREs, platform engineers, product owners, and business stakeholders. * Self-driven mindset with strong curiosity, innovation, and the ability to learn and apply emerging AI techniques effectively., * Bachelor's degree in computer science, Information Systems, Engineering, Data Science, Artificial Intelligence, or a related field, or equivalent practical experience. * 6 to 10 plus years of overall experience in software engineering, machine learning, or AI system development. * 3 plus years of hands-on experience building and deploying machine learning systems in production. * Strong experience in Python-based AI/ML development is required. * Experience working on observability, monitoring, or AIOps-related platforms is strongly preferred. * Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred. * Experience in AIOps, Observability, SRE, IT operations, or incident management domains. * Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, or remediation recommendations. * Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence. * Experience using vector databases and retrieval frameworks for enterprise search and agentic applications. * Experience integrating AI services with tools such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, or similar platforms. Familiarity with MCP-based client or agent integrations is a plus. ## Description * Design, develop, and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making. * Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation. * Develop causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services. * Create and fine-tune models for incident intelligence use cases such as forecasting service degradation, capacity risk prediction, alert prioritization, and anomaly explanation. * Design feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data. * Build intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph. * Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval. * Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents. * Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms. * Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures. * Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness. * Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices. * Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration. * Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions. * Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through experimentation and monitoring. * Establish engineering best practices for ML development, prompt engineering, evaluation, model deployment, testing, versioning, and documentation. * Support production incident analysis for AI/ML services and drive root cause identification and remediation for model or agent failures. * Create technical documentation covering model design, feature logic, training pipelines, evaluation metrics, deployment architecture, and agent workflows. Drive innovation in AI-enabled observability, causal intelligence, and agentic SRE workflows to enhance the value of the AIOps platform. ## Related Videos - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Navigating the AI Revolution in Software Development](https://www.wearedevelopers.com/videos/1266-navigating-the-ai-revolution-in-software-development) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)