Foundation Model Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
We are looking for an Foundation Model Engineer to design, execute, and operationalize fine-tuning workflows for large language models across supervised, preference-based, and reinforcement learning approaches. The role requires deep practical experience with modern training stacks, careful dataset construction, rigorous evaluation methodology, and the engineering discipline to operate complex training pipelines reliably. The ideal candidate combines strong ML intuition with production-grade engineering practices, and is comfortable navigating the trade-offs between data quality, compute budget, evaluation rigor, and shipping velocity. In this role you will work closely with cross-functional partners - product, design, engineering, operations, and business stakeholders - to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong
Requirements
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production., * 10 or more years of combined ML research and engineering experience, with significant LLM exposure.
- Strong proficiency in Python and modern deep learning frameworks, especially PyTorch.
- Hands-on experience fine-tuning transformer-based language models at non-trivial scale.
- Familiarity with distributed training strategies including FSDP, ZeRO, and pipeline parallelism.
- Experience with RLHF, DPO, or other preference optimization techniques.
- Strong understanding of evaluation methodology, benchmarks, and human evaluation design.
- Experience operating training jobs on GPU clusters and recovering from failures.
- Strong written and verbal communication skills.
- Track record of shipping or publishing impactful LLM work., * Publications at top-tier ML venues.
- Experience with multimodal model fine-tuning.
- Familiarity with synthetic data generation and dataset distillation.
- Open-source contributions to LLM training libraries.
- Exposure to responsible AI evaluation and red-teaming practices.
Benefits & conditions
4.24.2 out of 5 stars Remote $200,000 - $230,000 a year - Full-time
About the company
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
The Best Large Language Models on The Market
How to Become an AI Engineer
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Dev Digest 121 - AI goes offline