Staff Software Engineer

Anyscale, Inc.
San Francisco, CA, United States
5 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence C++ (Programming Language) Databases Software Debugging Distributed Systems Fault Tolerance Open Source Technology Software Architecture Software Engineering System Programming Rust (Programming Language)
+6 more
Multithreading Concurrency Apache Spark Information Technology Apache Kafka Stream Processing

Job description

  • Own a major technical area of Ray Core end to end, defining the roadmap, identifying the most important technical problems, and driving execution through production.
  • Lead large, technically complex projects spanning multiple engineers, teams, and/or organizations.
  • Set technical direction and make architectural decisions for distributed computing infrastructure used by demanding production workloads.
  • Design, build, and evolve core distributed-systems primitives rather than simply integrating existing platforms or frameworks.
  • Work on problems involving areas such as distributed execution, scheduling, resource management, fault tolerance, concurrency, networking, storage, or system performance.
  • Stay hands-on with implementation and debugging in a systems-oriented codebase, while raising the technical bar for the engineers around you.
  • Help shape the longer-term architecture and evolution of Ray Core as our workloads and scale continue to grow.

Requirements

  • 6+ years of software engineering experience, with a track record of increasing technical ownership.
  • Experience leading substantial projects end to end, including defining the problem, creating a roadmap, making architectural decisions, driving implementation, and owning the outcome in production.
  • Experience leading projects that are larger than a single-engineer effort, typically spanning multiple engineers and lasting multiple quarters.
  • Deep experience with distributed systems and computer systems.
  • Strong systems programming experience in languages such as C++, Rust, Java, or similar lower-level languages.
  • Strong understanding of systems concepts such as multithreading/concurrency, distributed coordination, resource management, fault tolerance, performance, or networking.
  • Experience building foundational systems such as databases, streaming systems, distributed runtimes, operating systems, schedulers, storage systems, Spark, Kafka, or similar infrastructure is highly relevant.
  • A track record of mentoring engineers and raising the technical bar of the teams around you.
  • Nice to have: Contributions to open-source infrastructure projects.

Benefits & conditions

  • We’re on a mission to make scalable computing effortless. Ray is the AI Compute Engine at the center of some of the world’s most powerful AI platforms
  • Our tech is in production at companies like OpenAI, Uber, Spotify, Instacart, and Cruise
  • We’re backed by Andreessen Horowitz, NEA, and Addition, with $250M+ raised to date
  • Recent partnerships with Azure, CoreWeave, and Google Cloud are putting AI-native compute directly into enterprise environments
  • Competitive salary and equity, plus health/dental/vision coverage (many plans up to 99% employer-covered)
  • We offer flexible time off, paid parental leave, and mental health support

About the company

At Anyscale, we’re on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that’s creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.

With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.

Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:41 min

Scale and diversity of software development teams

Bastian Heilemann Bastian Heilemann +1 · World Congress 2025

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

5:59 min

Analyzing concurrency bottlenecks in standard serverless architectures

Marco Plaul Marco Plaul +1 · World Congress 2023

1:37 min

Introduction to Apache Kafka benchmarking and performance analysis

Kirill Kulikov · LIVE

1:39 min

Speaker introduction and software engineering background

Daniel Raniz Raneland · LIVE

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

Videos

See all

Related articles

See all