> Markdown version of [/jobs/ext/243256-senior-ai-tools-engineer-sre-operations-geforce-now](https://www.wearedevelopers.com/jobs/ext/243256-senior-ai-tools-engineer-sre-operations-geforce-now). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Tools Engineer, SRE Operations - GeForce NOW - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $144,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Data Visualization, Python (Programming Language), Operational Databases, Reliability Engineering, Application Enhancement Tool, Cloud Platform System, Large Language Models, Grafana, Kubernetes, Information Technology, Data Management, Machine Learning Operations, Data Pipelines - **Published:** May 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3ef64e0a4e3cb979 ## About the Role Do you have experience in Tooling?, Do you have a Bachelor's degree?, We are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. Applicants with SRE or equivalent experience are encouraged., * B.S. in Computer Science, Statistics, or Engineering (or equivalent experience), and 5+ years of experience. * Strong proficiency in Python; familiarity with Go or other systems languages is a plus. * Practical experience building, optimizing, and deploying AI tools. * Strong knowledge of the AI space and current developments, including understanding how LLM-based platforms are built, optimized, and which platforms work best. * Hands-on experience with container orchestration (Kubernetes) and cloud environments (AWS cloud). * Active engagement with developments in the AI field and the ability to distinguish meaningful advances from noise when making technical decisions. * Expertise in automation and handling large-scale data pipelines. * Experience applying monitoring and visualization tools, such as Grafana, to interact with data. * Excellent ability to handle data sources and pipelines to transform and manage data. Ways to stand out in a crowd: * Understanding of SRE principles and experience managing production environments. * Strong in LLM improvement pipelines as well as a strong grasp of recent developments in LLM training. * Someone with excellent knowledge of LLMs and AI Models who can reason and recommend an approach that sustains the team and product long term. This person helps prevent grave mistakes by avoiding the wrong platform choice. * Understanding of SRE concepts and managing production environments as well as experience with Kubernetes, AWS, and other cloud technologies. * Proficiency in automation. ## Description You will build and deploy sophisticated AI-powered tools and products. These tools support the operation and optimization of a critical production global Geforce Now service. This role is critical for transforming extensive production data streams-such as signals, metrics, and logs-into actionable intelligence. The intelligence automates root cause analysis for incidents and predicts future service trends and patterns. * Build and implement robust AI/ML tools capable of analyzing production data to identify root causes for complex incidents and identify future operational trends. * Lead the development of brand-new LLM- and Agent-based systems to improve operational efficiency. * Establish and maintain excellent data management practices, including building pipelines to transform and handle large-scale data sources vital for model development. * Take charge of and enhance LLM-based pipelines while integrating a strong grasp of LLM progress into product development. * Act as a resident authority on AI Frameworks, recommending the best platforms, toolsets, and architectural approaches to ensure the long-term technical sustainability of the product. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)