Lead Data Scientist (Nvidia)

SoftServe, Inc.
United States
1 day ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Profiling Nvidia CUDA Distributed Systems Python (Programming Language) Machine Learning Language Modeling NumPy Tensorflow Software Engineering
+16 more
AI Infrastructure Pytorch Large Language Models Deep Learning Generative AI Pandas Kubernetes Information Technology Low Latency ONNX (Open Neural Network Exchange) Format HuggingFace Machine Learning Operations TensorRT Hardware Infrastructure Virtual Agents Nim (Programming Language)

Job description

In this role, you will combine deep hands-on expertise in GPU-accelerated AI, Generative AI, and modern AI infrastructure with the opportunity to shape enterprise AI strategies for global clients. You will lead technical engagements from discovery through production, influence architectural decisions, drive AI adoption, and contribute to NVIDIA-focused go-to-market initiatives while collaborating with client stakeholders and multidisciplinary engineering teams., * Lead end-to-end AI engagements, from discovery and solution strategy through architecture design, implementation planning, and production delivery

  • Translate complex business challenges into AI use cases, solution roadmaps, and scalable enterprise architectures
  • Design and validate production-ready AI solutions leveraging NVIDIA technologies across cloud and on-premises environments
  • Define reference architectures for Generative AI and Agentic AI solutions, including GPU infrastructure, Kubernetes orchestration, and inference optimization strategies
  • Evaluate and optimize AI inference performance by analyzing GPU utilization, latency, throughput, scalability, and infrastructure efficiency
  • Benchmark and optimize LLM serving frameworks and deployment configurations
  • Lead pre-sales activities, including technical discovery, workshops, solution positioning, proposals, and proof-of-concept initiatives
  • Collaborate with NVIDIA stakeholders and internal teams to develop reusable accelerators, solution blueprints, and industry offerings
  • Drive technical thought leadership through whitepapers, technical content, conference presentations, and mentorship

Requirements

  • 6+ years of experience in AI consulting, Generative AI, Agentic AI, Machine Learning, or Deep Learning, including ownership of client-facing engagements
  • Bachelor’s or Master’s degree in Computer Science, Applied Mathematics, Physics, Engineering, or related technical field preferred
  • Advanced expertise in Generative AI, Agentic AI, multimodal AI, transformers, Large Language Models (LLMs), and Vision Language Models (VLMs)
  • Hands-on experience with Python and modern AI/ML frameworks, including PyTorch, TensorFlow, Hugging Face, Pandas, and NumPy
  • Strong experience with NVIDIA AI technologies, including at least three of the following: NeMo, NIM, Triton, TensorRT-LLM, Riva, DeepStream, Metropolis, or Omniverse
  • Practical experience in deploying AI workloads on Kubernetes using Helm, NVIDIA GPU Operator, GPU device plugins, MIG/vGPU partitioning, and modern inference platforms such as vLLM or Ollama
  • Working knowledge of model quantization, inference optimization, and GPU profiling tools, including NVIDIA Nsight Systems, Nsight Compute, DCGM, PyTorch Profiler, and Triton or vLLM monitoring
  • Proven skill in analyzing GPU performance, identifying compute, memory, or I/O bottlenecks, and optimizing AI infrastructure for performance and cost efficiency
  • Experience in designing and deploying enterprise AI solutions on AWS, Azure, or GCP using CUDA, TensorRT, Triton Inference Server, DeepStream, and ONNX
  • Solid understanding of enterprise architecture, distributed systems, MLOps, AI governance, and modern software engineering practices
  • Strong advisory and stakeholder management skills
  • English proficiency for leading technical discussions with global clients and stakeholders

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all