Backend Software Engineer - Core Backend, Cloud Customer Experience

Crusoe Cloud
Sunnyvale, CA, United States
1 day ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$170,000.0 - $205,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence User Authentication C++ (Programming Language) Cloud Computing Computer Programming Continuous Delivery Continuous Integration Data Centers Data Validation Data Deduplication
+36 more
Data Structures Relational Databases Database Design Software Debugging Distributed Systems Middleware Fault Tolerance Protocol Buffers Design of User Interfaces Human-Computer Interaction PostgreSQL Open Source Technology Systems Development Life Cycle Queueing Systems Cloud Services Service Design Software Engineering Management of Software Versions Graphics Processing Unit (GPU) Computer Network Operations Cloud Platform System State Machines Backend Rate Limiting Event Driven Architecture Kubernetes Infrastructure Automation Frameworks Low Latency Apache Kafka Data Management Api Gateway Terraform Stream Processing Docker Programming Languages Microservices

Job description

We are seeking a Senior Backend Software Engineer to join the Core Backend team within Cloud Customer Experience (CCX). The CCX organization is at the forefront of delivering a best-in-class user experience for our AI-focused cloud platform, and Core Backend is the engine underneath it. Our mission is to provide an intuitive, seamless user flow while ensuring the backend reliability and scalability that set us apart from the competition.

The Core Backend team owns the foundational services behind every interaction a customer has with Crusoe Cloud. Concretely, the team builds and operates:

  • API Gateway: the front door to Crusoe Cloud. Every request to create, read, or manage a cloud resource flows through the gateway this team owns. That means request routing, authentication and authorization enforcement, rate limiting and throttling, API versioning, and the tracing that makes a single call debuggable end to end.
  • Resource Management Layer: the abstraction between customer intent and physical infrastructure. Resource models and lifecycle state machines, reconciliation loops that drive declared state toward reality, capacity and placement coordination, and one consistent API surface behind every client.
  • Quota Management: the system of record for what every organization and project is entitled to consume. Our team built and owns the quota model for all Crusoe products and its enforcement path, along with the request and approval workflows that gate GPU, compute, storage, and networking capacity - a low-latency, strongly consistent, correctness-critical service sitting directly in the path of every resource creation.
  • Notifications: the platform-wide notification and eventing system that tells customers what is happening to their infrastructure - quota changes, capacity events, maintenance windows, billing thresholds, security alerts - across email, in-console, and webhook delivery. Event pipelines, templating, fan-out, delivery guarantees, deduplication, and per-user preference management at scale.
  • User Onboarding Flows: the path from sign-up to first successful workload. Identity and organization provisioning, invitations and role assignment, verification checks, and the self-serve flows that let a new customer go from account creation to running GPUs without a human in the loop.

Your work will focus on making these systems fast, reliable, and coherent - one gateway, one resource model, one entitlement story, one notification surface - as Crusoe Cloud’s customer base and hardware footprint grow.

What You’ll Be Working On:

  • Design, develop, and maintain scalable and reliable services that power our cloud platform’s user-facing experiences.
  • Own and evolve the API gateway that fronts all Crusoe Cloud resource management traffic - routing, authentication and authorization enforcement, rate limiting, versioning, and tenant isolation - without becoming a bottleneck for the teams shipping behind it.
  • Build the notification and eventing backbone that keeps customers informed about the state of their infrastructure, with delivery guarantees and preference controls they can trust.
  • Evolve the quota and entitlement system into a first-class product surface: self-serve quota requests, transparent limits, and enforcement that is both fast and correct under contention.
  • Reduce time-to-first-GPU for new customers by hardening and automating onboarding, provisioning, and identity flows.
  • Extend the resource management layer so new infrastructure offerings can be modeled, exposed, and reconciled without bespoke one-off plumbing each time.
  • Collaborate with cross-functional teams, like product and design, to evaluate tools, frameworks, and customer needs, creating innovative solutions that differentiate Crusoe Cloud.
  • Contribute to architectural decisions that support reliability and maintainability across the company.
  • Mentor engineers, enhance hiring practices, and contribute to building a strong, inclusive engineering culture.

At Crusoe, we are redefining cloud infrastructure by integrating data center operations with seamless user experiences. From developing turn-key AI cloud infrastructure to managing advanced data center operations, Crusoe is at the forefront of innovation in AI-first cloud computing. Join us to help shape a platform that reduces carbon emissions while delivering best-in-class cloud services.

Requirements

  • Customer-Centric Mindset: A passion for creating intuitive, high-quality solutions that directly impact customer success and satisfaction. Any experience building out infrastructure tooling is a plus.
  • Professional Experience: 3+ years of software development experience, including programming with modern compiled languages such as Go, Rust, Java, or C++. Our services are primarily written in Go.
  • API and Service Design: Experience designing versioned public APIs (gRPC/protobuf, REST) and the SDK, CLI, or Terraform provider surfaces built on top of them, with an eye toward backward compatibility.
  • Gateway and Edge Concerns: A clear sense of what belongs at the API edge versus in a backing service - request routing, authn/authz enforcement, rate limiting and throttling, quota checks, input validation, and graceful degradation under load.
  • Cloud Expertise: Proven ability to design and scale fault-tolerant distributed systems and develop managed cloud services.
  • Data and State Management: Strong relational database skills (PostgreSQL), including schema design, transactions, isolation levels, locking, and safe online migrations for correctness-critical data such as quotas and entitlements.
  • Event-Driven Systems: Practical experience with message queues and streaming systems (Kafka, NATS, Pub/Sub, or similar) and the patterns that make them reliable - at-least-once delivery, idempotent consumers, retries, backoff, and dead-letter handling. Directly relevant to notifications and resource lifecycle events.
  • Asynchronous Workflows: Comfort building long-running, resumable workflows and reconciliation loops - state machines, control loops, or workflow engines such as Temporal - for provisioning, onboarding, and resource lifecycle management.
  • Technical Proficiency: Strong fundamentals in data structures, algorithms, microservices, and infrastructure tools like Docker, Kubernetes, Terraform, and CI/CD systems.
  • Observability and Reliability: Experience defining SLOs, instrumenting services with metrics, logs, and traces, and using that signal to debug production issues in systems you are on-call for.
  • Collaboration Skills: Ability to work with cross-functional teams to align priorities and deliver customer-first solutions.
  • Mentorship Abilities: Experience guiding engineers, improving hiring and onboarding processes, and driving team growth.
  • Communication Skills: Exceptional ability to articulate complex ideas and align technical solutions with customer needs.

Bonus Points:

  • Direct experience building or scaling a public cloud console or developer platform.
  • Familiarity with the unique requirements of AI/ML workloads and their impact on cloud resource management.
  • Experience with financial or billing systems at scale (usage-based billing, credits, etc.).
  • Prior experience in a high-growth startup environment or a “Neocloud” infrastructure provider.
  • Contributions to open-source projects related to cloud-native infrastructure or developer experience., Algorithms, Apache Kafka, Application Programming Interface (API), Architectural Services, Artificial Intelligence (AI), Authentication, Billing, C++ Programming Language, Cloud Computing, Communication Skills, Compiled Programming Languages, Construction, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Customer Experience, Customer Satisfaction, Customer Support/Service, Data Management, Data Structures, Database Design, Debugging Skills, Distributed Computing, Docker, Financial Systems, GPU (Graphics Processing Unit), Java, Machine Tool, Manufacturing, Mentoring, Messaging Middleware, Metrics, Microservices, Network Operations Center, On Call, Onboarding, Open Source, Plumbing, PostgreSQL, Problem Solving Skills, Product Design, Production Systems, Psychiatry and Mental Health, Public Cloud, Reconciliation, Relational Databases (RDBMS), Resource Management, Rust Programming Language, Scalable System Development, Software Development, Software Engineering, Startup, Systems Administration/Management, Team Player, Technical Recruiting, Time Management, Traffic Shaping, User Interface/Experience (UI/UX), Validation Testing

Benefits & conditions

  • Competitive compensation
  • Restricted Stock Units
  • Paid time off & paid holidays
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off

Compensation Range

Compensation will be paid in the range of $170,000 - 205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

About the company

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack - from electrons to tokens - to power the world’s most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We’re in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We’re solving that - with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We’re looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:38 min

Transitioning into backend engineering from web development

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all