Software Engineer - Systems

Boson AI USA Inc.
Santa Clara, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Airflow Amazon Web Services C++ (Programming Language) Cloud Computing Continuous Integration Extract Transform Load (ETL) Distributed Systems Fault Tolerance Python (Programming Language) Routing
+18 more
Rust (Programming Language) Data Logging ReactJS Large Language Models Apache Spark Caching Backend Rate Limiting Containerization Core Data Kubernetes Low Latency Apache Flink Apache Kafka Machine Learning Operations Video Streaming Data Pipelines Api Management

Job description

About the Role: Build and operate the core platform behind Boson’s model APIs and agentic products. You’ll own the infrastructure that every Boson agent runs on - API serving, state management, data pipelines, context retrieval, and execution runtime - and make it fast, reliable, and easy for product teams to build on., * Own and evolve the core platform infrastructure: API serving layer, state management, policy enforcement engine, and execution runtime for agentic workflows.

  • Design and operate high-throughput, low-latency distributed services that back our model API products - including request routing, load management, rate limiting, and multi-tenant isolation.
  • Build and maintain downstream data pipelines (ETL/ELT) for API logs, usage analytics, and billing - ensuring data correctness, freshness, and queryability at scale.
  • Develop production-grade internal SDKs and libraries with clean APIs, strong type safety, and clear contracts that product teams can build on confidently.
  • Architect context and memory systems for conversational workloads - low-latency retrieval, caching, and integration with vector stores and retrieval pipelines.
  • Instrument end-to-end observability: define SLIs/SLOs, build structured logging and tracing, and drive reliability improvements across the platform.
  • Collaborate closely with ML and product teams to integrate model serving, voice runtime, and tooling infrastructure under tight latency and quality constraints.

Requirements

  • 3+ years building and operating backend systems at scale - you’ve owned services that other teams depend on in production.
  • Strong distributed systems fundamentals: concurrency, fault tolerance, consistency tradeoffs, capacity planning.
  • Hands-on experience with data pipeline infrastructure (Kafka/Kinesis, Spark/Flink, Airflow, or similar) for log processing, analytics, or ETL workloads.
  • Track record of designing APIs and frameworks adopted by other engineering teams - you care about developer experience and long-term maintainability.
  • Proficiency in at least one systems language (Go, Rust, Java, C++) or Python in a performance-sensitive context.
  • Comfortable working across the stack: cloud infrastructure (AWS/GCP), containerized deployments (K8s), CI/CD, and production oncall., * Experience with LLM serving, agentic orchestration patterns (ReAct, planner-executor), or RAG pipelines.
  • Familiarity with emerging agent integration protocols (MCP, A2A) or orchestration frameworks (LangChain, LlamaIndex).
  • Background in real-time media systems (audio/video streaming, low-latency signaling).
  • Experience building high-stakes platform services (payments, identity, core data) where correctness and auditability are non-negotiable.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

1:38 min

Transitioning into backend engineering from web development

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all