AI Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
OPSWAT, a global leader in IT, OT, and ICS critical infrastructure cybersecurity, delivers an end-to-end platform that gives public and private sector organizations and enterprises the critical advantage needed to protect their complex networks, secure their devices, and ensure compliance. Over the last 20 years our commitment to innovative technology has earned the trust of more than 1,700 organizations, governments, and institutions globally, solidifying our role in protecting the world's critical infrastructure and securing our way of life.
, <p>OPSWAT is committed to harnessing the power of AI to drive meaningful advancements across all areas of our business. We are currently seeking a highly skilled Senior AI Engineer to lead the deployment and optimization of Large Language Models for efficient, high performance inference on GPU hardware.</p>
In this role, you will own the infrastructure layer that turns our AI models into fast, cost effective, production grade services. You will work closely with the AI/ML Team, the Enterprise & Data Team, and Infrastructure to design local LLM serving environments, maximize GPU utilization, and apply model compression techniques such as quantization from FP32 down to FP8/FP4. This is a chance to work at the deep intersection of models, serving software, and hardware within a global cybersecurity company, delivering the performance that powers OPSWAT's AI products at scale.
What You Will Do
- Local LLM Deployment & Hardware Setup: Design and build local LLM serving environments on GPU hardware, selecting GPU configurations based on VRAM, memory bandwidth, and workload. Install and maintain the full inference stack (NVIDIA drivers, CUDA, cuDNN, inference engines) across single and multi GPU deployments, including tensor and pipeline parallelism.
- Inference Optimization & Model Compression: Optimize LLM models for efficient GPU serving, balancing latency, throughput, memory, and cost. Apply quantization from FP32 to FP16/BF16, FP8, and FP4 using PTQ methods (GPTQ, AWQ, SmoothQuant) and QAT. Design mixed precision strategies and run calibration to minimize quality loss. Apply complementary techniques including pruning, distillation, and sparsity.
- Serving Engine & Performance Engineering: Deploy and tune high throughput serving engines (vLLM, TensorRT-LLM, TGI), optimizing KV cache (PagedAttention), continuous batching, and speculative decoding. Leverage optimized kernels (FlashAttention) and FP8 tensor cores. Build benchmarking for TTFT, inter token latency, throughput, GPU utilization, and cost per million tokens.
- Production Serving Infrastructure: Deploy quantized models with autoscaling, load balancing, and observability. Establish quality regression gates and run A/B tests of quantized vs full precision models on real traffic.
- Research & Propose Innovative Solutions: exploring and implementing novel inference optimization and model compression techniques.
What We Need From You
Requirements
</ul>
It Would Be Nice if You Have
- Experience writing or tuning custom CUDA / Triton kernels.
- Experience with multi GPU and distributed inference.
- Certifications in AWS or Azure architecture.
- Experience with CI/CD and cloud production deployment (Azure, AWS, or GCP), including GPU backed instances.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLOps And AI Driven Development
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence