Staff Software Engineer, Platform Metering

NSCALE, LLC
New York, NY, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$220,000.0 - $266,667.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing Data Systems Cursor (Graphical User Interface Elements) File Systems Distributed Systems Intrusion Detection Systems Python (Programming Language) System Availability Backend Usage Tracking Event Driven Architecture
+5 more
Kubernetes Bare Metal Data Management Slurm Webhooks

Job description

Nscale is hiring a Staff Software Engineer to build the source of truth for how customers consume Nscale’s GPU cloud. You’ll turn usage of provisioned compute across Kubernetes, Slurm, bare metal, and storage into reliable, auditable records that power usage-based billing, credit enforcement, customer visibility, and internal cost intelligence.

This role sits at the intersection of cloud infrastructure, distributed systems, billing, and platform reliability. You’ll own the domain-level architecture for metering provisioned resources such as GPU compute, Kubernetes nodes, bare-metal capacity, storage, and future platform services. Your work will ensure that every billable resource has a trustworthy usage trail: accurate enough for billing, timely enough for credit enforcement, and explainable enough for customers, finance, and engineering teams. This is an opportunity to define a foundational platform domain early, setting the metering architecture and standards that future Nscale services will build on.

Cloud metering powers Nscale’s usage-based billing platform by producing accurate, deduplicated, auditable usage records for rating, credit burn-down, entitlements, cost attribution, margin analysis, and customer-facing usage dashboards.

What you’ll work on

  • GPU and compute metering: measuring provisioned GPU-hours across instances, Kubernetes, Slurm, bare metal, and future compute products.
  • Storage metering: measuring GiB-hours for file storage, object storage, and future storage products.
  • Usage event production: building controllers and resource watchers that emit durable service consumption events into the metering pipeline.
  • Usage attribution: mapping consumption to organizations, projects, regions, clusters, flavors, allocations, and workloads.
  • Credit and customer safeguards: providing the usage signals needed for prepaid credit burn-down, headroom checks, grace periods, customer notifications, and fair enforcement when credit is exhausted.
  • Reconciliation and visibility: ensuring usage can be audited, replayed, explained, and surfaced to customers, finance, support, and engineering., * Domain-level technical direction. Set the technical direction for platform metering across a defined domain, influencing engineering squads across Nscale that produce or consume usage data.
  • Design accurate metering models. Define how Nscale measures resources over time, including instance-hours, GPU-hours, GiB-hours, readiness states, allocation burn-down, and future shared-pool or pod-level usage models.
  • Build reliable event producers. Design and implement controllers and resource watchers that emit self-contained usage quanta for compute, Kubernetes, bare metal, storage, and other platform resources.
  • Engineer for idempotency and auditability. Ensure metering events can survive retries, redelivery, backfills, partial outages, and customer disputes through deterministic transaction IDs, provenance, traceability, and reconciliation.
  • Integrate with billing systems. Partner with billing, product, finance, and platform teams to ensure usage events can be rated, aggregated, credited, and surfaced to customers and internal stakeholders.
  • Create leverage through standards. Establish shared schemas, libraries, conventions, dashboards, and operational runbooks so that new platform services can add metering consistently.
  • Own production outcomes. Operate the metering platform with strong observability, alerting, incident response, data-quality checks, and reconciliation against billing records and infrastructure state.
  • Mentor and influence. Guide other engineers through architecture reviews, implementation choices, and operational best practices for high-integrity usage systems.

Requirements

  • Extensive experience designing, building, and operating distributed systems in production, ideally in cloud infrastructure, data platforms, billing, control planes, or platform engineering.
  • Experience with resource lifecycle, capacity, or usage tracking systems such as compute instances, Kubernetes nodes, storage volumes, jobs, workloads, quotas, or entitlements.
  • Strong understanding of event-driven architecture, including reliable delivery, idempotency, aggregation, replay, and failure handling.
  • Strong operational discipline, including monitoring, alerting, incident response, reconciliation, data-quality checks, and post-incident improvement.
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • Strong software engineering fundamentals, with proficiency in typed backend or systems languages. Our primary stack is Go, with some services in Rust and Python.
  • Comfortable working in a fast-paced, ambiguous environment with high ownership, pragmatic judgement, and a bias toward measurable business impact.
  • You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.

Preferred

  • Experience with cloud billing, chargeback/showback, prepaid credit systems, entitlements, quota enforcement, or customer-facing usage dashboards.
  • Experience with GPU cloud infrastructure, Kubernetes, bare-metal provisioning, workload scheduling, storage platforms, or AI/ML inference and training workloads.
  • Experience with billing, usage, ledger-style, event-sourced, or replayable data systems that support auditability, reconciliation, backfills, and dispute investigation.
  • Experience with Kubernetes controllers/operators, controller-runtime, CRDs, admission webhooks, or multi-cluster resource watchers.
  • Experience with infrastructure-as-code, cloud providers, regional control planes, and service catalogs or flavor/rate-card models.
  • Strong product sense for making usage data understandable to customers, finance, and support teams, especially when investigating billing disputes or consumption anomalies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

1:28 min

Synchronizing database webhook triggers with pipeline webhooks

Bobur Umurzokov · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:30 min

Consumer challenges in ingesting and verifying webhooks

Phil Leggetter Phil Leggetter · World Congress 2026 Europe

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all