ML Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
- Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
- Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing.
- Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints.
- Build autoscaling and capacity management systems that balance latency, throughput, and cost.
- Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
- Integrate model serving with API gateways, identity systems, and observability platforms.
- Implement caching, prompt deduplication, and response reuse strategies where appropriate.
- Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking.
- Develop deployment workflows including canary releases, shadow testing, and automated rollback.
- Operate incident response for high-availability AI services and drive durable reliability improvements.
- Collaborate with ML and product teams to support new model releases and capability rollouts.
- Implement security controls including request signing, content filtering, and abuse detection at the serving layer.
- Document operational procedures, performance characteristics, and tuning guidance for internal teams.
- Stay current with AI serving research and translate advances into production capabilities.
Requirements
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Bachelor’s or Master’s degree in Computer Science or a related field.
- 10 or more years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Deep experience operating high-throughput, low-latency services in production.
- Hands-on experience with LLM or large model inference frameworks such as vcLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
- Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
- Experience with observability stacks including metrics, tracing, and structured logging.
- Solid grounding in performance engineering and capacity planning.
- Strong communication and incident response skills., * Open-source contributions to model serving infrastructure.
- Experience with multi-region or globally distributed AI serving.
- Familiarity with model quantization, distillation, and compression techniques.
- Exposure to FinOps for AI workloads and cost-efficient serving design.
- Experience supporting external-facing AI APIs at scale.
Benefits & conditions
We offer a wide range of career opportunities across various domains, from engineering and research to marketing and operations. We provide comprehensive benefits, competitive compensation packages, and a supportive work-life balance to ensure your well-being and success.
Explore our current job openings, and discover how you can contribute to our mission of driving innovation and shaping a brighter tomorrow. We are excited to learn about your skills, experiences, and aspirations, and how they align with our company’s vision.
Thank you for considering Bright Vision Technologies as your potential employer. We look forward to welcoming you to our talented team and embarking on an exciting journey together.
About the company
At Bright Vision Technologies, we believe that innovation and talent go hand in hand. We are delighted that you are considering a career with us and taking the first step towards joining our dynamic team.
As a leading technology company, we are driven by a shared vision to create a brighter future through groundbreaking solutions. We embrace creativity, collaboration, and a passion for excellence in everything we do.
Our work environment is built on trust, respect, and inclusivity. We value diversity and understand the power of different perspectives coming together to drive innovation. We foster a culture of continuous learning and growth, where you can expand your skills, explore new horizons, and reach your full potential.
Joining Bright Vision Technologies means being part of a team that pushes boundaries and pioneers cutting-edge technologies. We tackle complex challenges and deliver transformative solutions that make a real impact in the world. Your contributions will be valued, and your ideas will shape the future of our company., Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Highest Paying Tech Companies for Developers
Dev Digest 121 - AI goes offline
7 Cloud Computing Trends Coming in 2025 for Developers