Senior/Principal Visual ML Engineer

Clarvos LLC
Charleston, SC, United States
25 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Languages
Sign Languages
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Computer Vision Microsoft Azure Cloud Engineering Nvidia CUDA Distributed Systems Python (Programming Language) Machine Learning Language Modeling Open Source Technology
+20 more
Systems Architecture Video Editing Data Ingestion Pytorch Large Language Models Multi-Agent Systems Deep Learning Model Validation Generative AI Kubernetes Information Technology ONNX (Open Neural Network Exchange) Format HuggingFace Production Code Machine Learning Operations TensorRT Stable Diffusion GPT Docker Microservices

Job description

Travel Requirement: 0-10%, Quarterly for meetings, We are seeking a Senior/Principal Machine Learning Engineer with deep expertise in Visual Language Models (VLMs), Large Vision Models (LVMs), Generative AI, and multimodal foundation models to build the next generation of AI-powered creative technologies.

This is a hands-on technical role responsible for architecting, developing, and deploying state-of-the-art AI systems for image, video, and creative generation. You will work closely with Product, Engineering, Data Science, and Design teams to build production-scale GenAI capabilities that power creative automation, digital advertising, content personalization, and campaign optimization.

You will drive innovation across the entire lifecycle, from research and experimentation through large-scale production deployment, while helping establish the company’s long-term Visual AI strategy.

ESSENTIAL FUNCTIONS AND RESPONSIBILITIES:

Visual AI & Generative AI Development

  • Design and develop production-grade AI systems for:
  • Image generation
  • Video generation
  • Image editing and enhancement
  • Creative optimization
  • Style transfer
  • Multimodal content understanding
  • Brand-aware content generation
  • AI-assisted creative workflows
  • Build scalable pipelines for automated creative generation across multiple marketing channels.
  • Research and implement state-of-the-art diffusion, transformer, autoregressive, and multimodal architectures.
  • Fine-tune and optimize foundation models for enterprise production use cases

Vision & Multimodal ML

Develop and optimize systems using:

  • Vision Language Models (VLMs)
  • Large Vision Models (LVMs)
  • Multimodal LLMs
  • Diffusion models
  • Transformer-based image/video generation
  • Contrastive vision-language models
  • Image-text alignment models
  • Visual reasoning models

Model Development & Optimization

  • Train, fine-tune, optimize, and deploy large-scale generative AI models using advanced techniques including LoRA, QLoRA, PEFT, distillation, quantization, and prompt optimization.
  • Build robust model evaluation frameworks to measure creative quality, visual fidelity, consistency, brand alignment, safety, hallucination risk, and human preference alignment.
  • Improve model performance across quality, latency, scalability, and cost through continuous experimentation, benchmarking, and production optimization

Lead development and building of AI systems for:

  • Text-to-video generation
  • Image-to-video generation
  • Video editing
  • AI avatars
  • Motion transfer
  • Storyboarding
  • Creative sequencing
  • Marketing video generation
  • Dynamic creative optimization

System Architecture & Technical Requirements

  • Lead decisions around foundation models, fine-tuning strategies, RAG pipelines, embeddings, and ranking systems.
  • Deep expertise in Generative AI, multimodal foundation models, Vision Language Models (VLMs), Large Vision Models (LVMs), diffusion models, transformers, and autoregressive architectures, with hands-on experience building image and video generation systems using leading models such as FLUX, Stable Diffusion, Imagen, Veo, Runway, Kling, and open-source video diffusion models.
  • Strong experience with computer vision and multimodal AI frameworks (CLIP, Florence, Qwen-VL, LLaVA, SAM, YOLO, Grounding DINO) and applying them to visual understanding, generation, editing, and creative optimization.
  • Proven ability to productionize large-scale AI models using modern ML infrastructure including Hugging Face, Diffusers, DeepSpeed, FSDP, TensorRT, ONNX, CUDA/GPU optimization, and cloud-native MLOps platforms (AWS/GCP/Azure, Kubernetes, Kubeflow, MLflow, distributed inference).
  • Architect and oversee scalable LLM/GenAI systems for MarTech/AdTech use cases
  • Design and deploy multi-agent systems using frameworks such as LangGraph, AutoGen, CrewAI, MCP, or equivalent.
  • Own end-to-end ML system design: data ingestion, feature pipelines, training, inference, evaluation, and monitoring.

Cross-Functional Collaboration

  • Work closely with Product, Data, and Platform teams to translate business needs into scalable ML capabilities.
  • Communicate complex ML concepts clearly to executive leadership, stakeholders, and the Board.
  • Contribute to technical narratives used for fundraising, company valuation, and strategic planning.

Requirements

  • MS or PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, Robotics, or a related field.
  • 8-12+ years of experience developing production ML systems.
  • 5+ years of experience in Deep Learning and Computer Vision.
  • 3+ years of hands-on experience with Generative AI for images and video.
  • Expert-level proficiency in Python and PyTorch.
  • Strong software engineering fundamentals with production-quality code.
  • Solid understanding of distributed systems, GPU optimization, batching, and cost-aware inference.
  • Excellent software engineering fundamentals (Python, APIs, microservices, Docker, Kubernetes).

PHYSICAL REQUIREMENTS/WORKING CONDITIONS:

Standing/Walking/Mobility: Must have mobility to attend meetings remotely and in person.

Climbing/Stooping/Kneeling: 0% - 10% of the time.

Lifting/Pulling/Pushing: 0% - 10% of the time.

Fingering/Grasping/Feeling: Must be able to write, type and use a telephone system 100% of the time.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all