> Markdown version of [/jobs/ext/1401276-ai-ml-observability-engineer](https://www.wearedevelopers.com/jobs/ext/1401276-ai-ml-observability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Observability Engineer - **Company:** STRADIT LLC - **Location:** Dallas, TX, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, ARM Architecture, Baselining, Cloud Computing, Continuous Integration, Distributed Systems, Python (Programming Language), Machine Learning, NumPy, OpenFlow, Tensorflow, DataOps, Web Services, Datadog, Pytorch, Delivery Pipeline, Large Language Models, Snowflake, Grafana, Prompt Engineering, Multi-Cloud, Generative AI, Pandas, Matplotlib, NetScout, Scikit Learn, Kubernetes, Infrastructure Automation Frameworks, SolarWinds (Software), Machine Learning Operations, ArcSight Event Correlation, Data Pipelines, Dynatrace, Microservices - **Published:** July 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=999731d2d66ddf42 ## About the Role Core Technical Skills * Strong proficiency in Python and data science/ML libraries: NumPy, Pandas, scikit learn, TensorFlow, PyTorch, Matplotlib, Seaborn. * Experience with Generative AI, LLM fine tuning, prompt engineering, RAG pipelines, and LLM evaluation frameworks. * Expertise in developing and deploying ML models in production (batch & streaming). * Strong understanding of statistics, time series modeling, and anomaly detection. Observability & Telemetry * Experience with OpenTelemetry for logs, metrics, traces, spans. * Familiarity with Observability concepts: Golden Signals, SLO/SLI design, APM, RUM, Synthetics, event correlation, baselining. * Experience with Observability tools such as: Grafana (Alloy agents, dashboards, ML capabilities), Dynatrace, Monte Carlo (Data Observability), Netscout, ThousandEyes, SolarWinds, NetBrain. Cloud, Data & Platform * Hands on with AWS (SageMaker, Bedrock), Snowflake ML, Snowflake/Openflow, Snowflake AI Observability tooling. * Experience building Snowflake data pipelines (streams, tasks, UDFs) - plus for Cortex features. * Strong understanding of distributed systems and microservices telemetry requirements. Automation & Engineering Quality * Experience with automation pipelines, CI/CD, and infrastructure as code patterns supporting Observability adoption. * Ability to build asynchronous Python APIs or services for model inference and operational integration., * Experience developing agentic AI systems that analyze telemetry, generate action recommendations, or execute automated operational responses. * Experience building self-healing patterns, including automated rollback, service restarts, configuration corrections, and predictive maintenance. * Experience in Snowflake ML workflows, Snowflake Cortex Agents, and data pipeline automation. * Exposure to AI-enabled alerting, RCA automation, and operational self-healing concepts. * Experience with large-scale operational telemetry and multi-cloud ecosystems. Soft Skills * Strong analytical thinking and problem solving. * Excellent communication skills for cross functional collaboration with infrastructure, SRE, engineering, business, and leadership teams. * Curiosity, continuous learning mindset, and passion for applied AI and Observability. ## Description We are seeking a passionate and hands-on AI/ML Engineer to accelerate our Enterprise Observability strategy. This role will design, build, and operationalize AI/ML capabilities that enhance end to end telemetry pipelines, anomaly detection, intelligent alerting, and proactive system resiliency. You will work at the intersection of AI/ML engineering, Observability platforms, and automation, developing solutions that improve detection, diagnosis, and prevention of operational issues across distributed systems., * Design and deploy AI/ML models supporting anomaly detection, baselining, event correlation, and predictive operational analytics. * Build and integrate AI-enabled capabilities into enterprise Observability platforms, including Grafana, APM/RUM tools, network telemetry systems, and data observability tools. * Develop AI Agents that can autonomously triage issues, recommend corrective actions, and initiate automated remediation workflows to reduce recovery time and improve system resilience. * Implement self-healing automation using AI-driven decisioning, integrating with orchestration frameworks, service APIs, and infrastructure automation pipelines. * Engineer and maintain real-time and batch data pipelines using Snowflake ML Jobs, Snowflake Cortex, streams, tasks, and UDFs. * Implement and manage OpenTelemetry-based telemetry ingestion for logs, metrics, traces, and spans across distributed systems. * Build asynchronous Python APIs and services for model inferencing and operational integration. * Enhance observability intelligence with AI-powered capabilities such as root-cause acceleration, chatbot/search enablement, and automated insights. * Contribute to SLO/SLI modeling, Golden Signals instrumentation, and Observability NFR adoption. * Collaborate across engineering, SRE, platform and business teams to embed proactive intelligence and Observability standards throughout the ecosystem. ## Related Videos - [Unlocking the AI Black Box: Building Trust in the Era of Agentic Production](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Mastering AI-Driven Problem Solving in Engineering with Observability](https://www.wearedevelopers.com/videos/994-mastering-ai-driven-problem-solving-in-engineering-with-observability) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)