> Markdown version of [/jobs/ext/2278964-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/2278964-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** F5 Networks, Inc. - **Location:** San Jose, CA, United States - **Salary:** $176,600.0 - $265,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Business Software, C++ (Programming Language), Cloud Computing, Nvidia CUDA, Data Centers, Monitoring of Systems, Python (Programming Language), Linux Kernel, Load Testing, Open Source Technology, Software Systems, Rust (Programming Language), Graphics Processing Unit (GPU), Google Cloud, Large Language Models, Kubernetes, Low Latency, Performance Monitor, Hardware Acceleration, Machine Learning Operations, TensorRT, Decoding, Automation Anywhere, Docker, Golang, Programming Languages - **Published:** August 28, 2026 - **Apply:** https://ffive.wd5.myworkdayjobs.com/f5jobs/job/San-Jose/AI-Inference-Engineer_RP1038660 ## About the Role * Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows. * Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization. * Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure. * Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs. Preferred Experience: * Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention. * Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels). * Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges. * Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts. ## Description The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments-from GPU-rich data centers to resource-constrained edge devices-with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy. This role is pivotal in advancing F5's AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance., High-Performance AI Serving * Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale. * Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications. Hardware Acceleration and Optimization * Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs. * Collaborate with hardware teams to maximize utilization and performance across various computational environments. Inference Orchestration and Scalability * Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration. * Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability. Performance Monitoring and Observability * Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs). * Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale., Success Metrics (KPIs): * Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second. * Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization. * Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments. * Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput. Why Join F5? F5 empowers you to push boundaries in AI optimization and high-performance engineering. Joining our team means: * Collaborating with cutting-edge technologies and hardware solutions to support real-time AI applications. * Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving. * Driving transformative projects that deliver real-time AI reliability to global customers while maintaining cost and efficiency standards. * Working on advanced MLOps solutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment. What Success Looks Like: As an AI Inference Engineer at F5, success is measured by your ability to: * Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems. * Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement. * Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals. * Translate complex AI and inference workflows into practical solutions that align with F5's strategic objectives. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)