Staff Software Engineer

Cloudradiant Corp.
San Jose, CA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$184,000.0 - $230,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Nvidia CUDA Computer Programming Python (Programming Language) Node.Js Open Source Technology Software Engineering Graphics Processing Unit (GPU) Large Language Models Multi-Agent Systems Prompt Engineering
+9 more
Generative AI AI Platforms Kubernetes ONNX (Open Neural Network Exchange) Format HuggingFace Machine Learning Operations GPT Serverless Computing Docker

Job description

  • Enterprise AI Services: Design and implement elegant, scalable application services (Go/Node.js) that wrap AI capabilities for enterprise use.
  • K8s-Native AI Orchestration: Lead the deployment of inference servers (vLLM, Triton) using KServe, KubeRay, or Knative to ensure serverless-style scaling for AI workloads.
  • Developer Velocity: Build internal tooling, SDKs, and “AI Gateways” that enhance team agility and simplify the integration of Foundation Models (Llama, GPT) into product features.
  • RAG & Prompt Engineering: Architect robust Retrieval-Augmented Generation (RAG) pipelines and prompt management services that integrate seamlessly with vector databases and enterprise data sources.
  • Cross-Functional Collaboration: Partner with UI engineers, UX designers, and Product Management to ensure the AI platform is not just powerful, but highly usable for internal developers.
  • Infrastructure & Security: Ensure AI workloads are secure, multi-tenant, and optimized for GPU resource scheduling (MIG, fractional GPUs) within Kubernetes.

Requirements

  • Bachelor’s degree with 6+ years of software engineering experience (or equivalent Masters/PhD tenure), with at least 2+ years focused on AI/ML systems.
  • Expert proficiency in Python (for AI ecosystem) and strong competence in a systems language like Go or Rust/C++ (for high-performance serving layers).
  • Deep understanding of LLM deployment challenges and runtimes (e.g., vLLM, ONNX, TorchServe, Triton). Familiarity with quantization techniques (AWQ, GPTQ) to optimize model size/speed.
  • Experience building complex workflows using tools like LangChain or LlamaIndex, and deploying them on containerized infrastructure (Docker/Kubernetes).
  • Ability to navigate the rapidly changing AI landscape, filtering hype from practical engineering solutions, and driving technical alignment across teams.

You May Also Have:

  • Model Fine-Tuning: Experience with efficient fine-tuning techniques (PEFT, LoRA/QLoRA) on custom datasets.
  • GPU Optimization: Familiarity with CUDA programming or profiling GPU performance (Nsight systems).
  • Open Source: Contributions to open-source AI projects (HuggingFace transformers, vLLM, etc.).

Benefits & conditions

This is more than cloud management, it’s about building the foundation for a consistent, secure, and compliant cloud experience that gives organizations 100% access to 100% of their data, anywhere.

With the recent acquisition of Taikun, we are simplifying Kubernetes and cloud management even further, creating a platform that is unified, scalable, and future-ready.

If you are passionate about Kubernetes, not just using it but building it at the core managing workloads across hybrid clouds and datacenters and obsessed with performance, devops, etc. this is where you belong.

This role is not eligible for immigration sponsorship.

The anticipated annual base salary range for this position is:

  • California: $184,000- $230,000

Individual compensation within the published range is determined by the candidate’s skills, experience, qualifications, and primary work location. In addition to base pay, sales roles are eligible for Cloudera’s commission plan, while non-sales roles are eligible for the corporate incentive plan. All employees receive a comprehensive benefits package.

What you can expect from us:

  • Generous PTO Policy
  • Support work life balance with Unplugged Days
  • Flexible WFH Policy
  • Mental & Physical Wellness programs
  • Phone and Internet Reimbursement program
  • Access to Continued Career Development
  • Comprehensive Benefits and Competitive Packages
  • Paid Volunteer Time
  • Employee Resource Groups

About the company

At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we’re the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.

Ready to take cloud innovation to the next level? Join Cloudera’s Anywhere Cloud team and help deliver a true “build your own pipeline, bring your own engine” experience - enabling data and AI workloads to run anywhere, without friction or vendor lock-in.

We bring the best of public cloud - cost efficiency, scalability, elasticity, and agility - to wherever data lives: public clouds, private data centers, and the edge. Powered by Kubernetes, our hybrid architecture separates compute and storage to maximize flexibility and optimize infrastructure usage.

This isn’t just cloud management - it’s about building a consistent, secure, and compliant cloud experience that gives organizations full access to all their data, anywhere.

With the acquisition of Taikun, we’re simplifying Kubernetes and cloud management even further, creating a unified, scalable, future-ready platform. If you’re passionate about Kubernetes - not just using it, but building it at the core, managing workloads across hybrid clouds and data centers, and obsessing over performance and DevOps - this is where you belong.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 ¡ World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz ¡ World Congress 2025

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset ¡ World Congress 2023

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 ¡ World Congress 2026 Europe

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky ¡ World Congress 2025

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice ¡ World Congress 2025

Videos

See all

Related articles

See all