> Markdown version of [/jobs/ext/283928-ai-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/283928-ai-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Infrastructure Engineer - **Company:** Zensors Inc. - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Computer Vision, C++ (Programming Language), Software as a Service, Compilers, Program Optimization, Nvidia CUDA, Computer Programming, Databases, Distributed Computing Environment, FFmpeg, Python (Programming Language), Machine Learning, Neo4j, Performance Tuning, SQL Databases, Data Streaming, Video Editing, AI Infrastructure, Pytorch, Deep Learning, Parallel Computation, Information Technology, ONNX (Open Neural Network Exchange) Format, Data Management, Machine Learning Operations, TensorRT - **Published:** May 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a772cba89893f848 ## About the Role Do you have experience in Performance tuning?, * BS/MS or Ph.D. in Computer Science, Electrical Engineering, or a related discipline. * Strong programming skills in C/C++ and Python. * Experience with model optimization, quantization, and efficient deep learning techniques (e.g., knowledge distillation, pruning). * Deep understanding of GPU hardware performance, including execution models, thread hierarchy, memory/cache management, and the cost/performance trade-offs of video processing. * Experience with profiling and benchmarking tools (e.g., Nsight Systems, Nsight Compute) to validate performance on complex architectures. * Experience identifying and resolving compute and data flow bottlenecks, particularly in high-bandwidth video processing pipelines. * Strong communication skills and the ability to work cross-functionally between research and infrastructure teams., * Familiarity with database systems (e.g., SQL, Neo4j). * Work in Computer Vision, Deep Learning, and Vision Transformers. * Experience with video processing frameworks such as NVIDIA DeepStream, DALI, or FFmpeg. * Familiarity with ML compilers (e.g., TVM, MLIR) or inference engines like TensorRT or ONNX Runtime. * Knowledge of distributed training systems or cloud-scale inference serving (e.g., Triton Inference Server). ## Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video streams. As a Machine Learning Engineer in ML Runtime & Optimization, you will develop technologies to accelerate the training and inference of computer vision models that power smart spaces and cities. Your responsibilities will include: * Optimizing Core ML Pipelines: Identifying key bottlenecks in our current video analytics pipeline and performing in-depth analysis to ensure the best possible performance on current server and edge compute architectures. * Cross-Stack Collaboration: Collaborating closely with AI research and platform engineering teams to optimize core parallel algorithms and influence the design of our next-generation inference infrastructure. * Model Acceleration: Applying advanced model optimization techniques-such as quantization (Int8/FP16), pruning, and layer fusion-to our Vision Transformers (ViTs) and CNNs to maximize throughput and minimize latency. * Building Efficient Operators: Working across the entire ML framework/compiler stack (e.g., PyTorch, CUDA, TensorRT, and NVIDIA DeepStream) to write custom optimized ML operator libraries. * Resource Efficiency: Reducing the compute cost per video stream to enable massive scalability of our SaaS product. * Data Management: Building, improving, maintaining, and operating systems to facilitate the collection, labeling, and use of visual data for ML training. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [AI vs Recruiters and Applicants, Turmoil in the Games Industry, What to Put on a CV](https://www.wearedevelopers.com/videos/1363-ai-vs-recruiters-and-applicants-turmoil-in-the-games-industry-what-to-put-on-a-cv) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Making neural networks portable with ONNX](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)