> Markdown version of [/jobs/ext/1810255-lead-software-engineer-data-cloud](https://www.wearedevelopers.com/jobs/ext/1810255-lead-software-engineer-data-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Software Engineer, Data Cloud - **Company:** Salesforce.com, Inc. - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $172,500.0 - $260,100.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Batch Processing, Big Data, Software as a Service, Cloud Computing, Computer Clusters, Code Review, Encodings, Computer Programming, Customer Data Management, Cursor (Graphical User Interface Elements), Distributed Systems, Job Scheduling, Python (Programming Language), Network Control, Open Source Technology, Software Tools, Prometheus, Software Engineering, Data Streaming, Google Cloud, GitHub Copilot, Grafana, Apache Spark, Build Management, Kubernetes, Apache Flink, Integration Frameworks, Machine Learning Operations, Presto, Hardware Infrastructure, Stream Analytics, Dynatrace, Serverless Computing, Golang - **Published:** July 15, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17597304?backUrl=%2Fcareer%2F17597304%2FLead-Software-Engineer-Data-Cloud-Washington-Bellevue ## About the Role * 10+ years of software engineering experience with deep expertise in distributed systems * Expert-level Kubernetes experience: operators, controllers, scheduling, resource management * Production experience running large-scale infrastructure serving mission-critical workloads * Strong programming skills in Go, Java, Python, or Rust * Deep understanding of containers, orchestration, networking, storage, and compute isolation * Experience designing and building control planes for distributed infrastructure * Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows * Proven technical leadership: driving architecture, mentoring engineers, influencing technical direction, * Experience with GPU infrastructure and ML/AI workload orchestration * Background in building serverless or function-as-a-service platforms * Expertise in observability tools: Prometheus, Grafana, OpenTelemetry, distributed tracing * Experience with big data processing frameworks (Spark, Flink, Presto, etc.) * Knowledge of vector databases and embedding infrastructure * Contributions to open-source infrastructure projects * Experience in multi-tenant SaaS environments with strong isolation requirements * Familiarity with major cloud providers: AWS, GCP, Azure ## Description Salesforce Data 360 is transforming how enterprises unify, analyze, and activate their data at scale. We're seeking a Lead Member of Technical Staff to architect and build the next generation of our control plane services, powering workloads that serve millions of customers worldwide. In this role, you'll lead the evolution of our compute platform, designing systems that orchestrate diverse compute workloads including GPU clusters, batch processing, real-time analytics, and AI/ML systems. You'll work at the intersection of infrastructure, distributed systems, and emerging AI capabilities. You architect and deliver highly scalable, high-quality software while collaborating with geographically distributed engineers and product owners. You bring deep technical expertise and a commitment to engineering excellence, helping define the standards and culture that drive the team forward. What You'll Actually Be Doing: * Design and build our next gen Control Plane services * Drive architectural decisions on resource isolation, multi-tenancy, and distributed system design * Evolve compute infrastructure to support diverse workload patterns: batch processing, streaming, interactive queries, and AI/ML workloads at true petabyte scale * Design intelligent job scheduling and orchestration systems to maximize resource utilization * Build platform primitives for capacity management, priority-based allocation, and workload isolation * Implement custom Kubernetes operators and controllers for specialized workload requirements * Develop infrastructure strategies to sustain AI/ML workloads, focusing on GPU management and model deployment * Architect sophisticated resource management protocols optimized for varied and complex compute ecosystems * Partner with infrastructure engineering teams to establish integration frameworks for advancing AI technologies * Lead code reviews, define coding standards, and translate functional requirements into technical requirements - ensuring all developed features meet design and code quality standards while fulfilling customer requirements. * Collaborate with product teams and business partners on product and feature design, coordinating the evaluation, scope, and completion of new development requests. * Assess the impact of new requirements across all upstream and downstream applications, systems, and processes, and work with geographically distributed engineers to ensure successful product delivery. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Super scaling for the Super Bowl: How to survive 30 million users hitting your backend in 30 minutes](https://www.wearedevelopers.com/videos/100356-super-scaling-for-the-super-bowl-how-to-survive-30-million-users-hitting-your-backend-in-30-minutes) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)