Staff Software Engineer, Distributed Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
Orbis is seeking a Staff Software Engineer, Distributed Systems to define architecture across the messaging, state, and reconciliation layers of Catalyst, our secure distributed data platform that moves and transforms data reliably. We are looking for someone who has shaped platform architecture and wants to help set technical direction for a team building critical infrastructure., * Define, maintain, and expand the Golang and Python frameworks and libraries that support a large set of batch and/or streaming plugins to allow movement and integration of data
- Solving business needs at scale by expanding the patterns of data movement across disparate data sources
- Manage the lifecycle of critical components and dependencies, utilizing automated tooling to proactively identify vulnerabilities and facilitate safe, tested upgrades across the ecosystem
- Drive the design of high-performance, low-latency architectures optimized for specialized workloads, such as TAK protocols and RTSP media streams
- Define architecture across multiple feature areas or subsystems of Catalyst’s distributed data plane, including event streaming, queuing, and state machines, and evolve it as the platform scales
- Act as a domain expert in one or more critical areas across the data lifecycle (e.g., event streaming, state management, or storage engines)
- Define, monitor, and improve on reliability targets and SLOs using Datadog, and lead incident response and postmortems for the core distributed-systems paths
- Participate in on-call and incident response for our data platform including leading post-incident reviews, triage, prevention, and mitigation
- Document and socialize procedures, naming conventions, and other standards via Notion and Github to maintain a high-quality, searchable engineering knowledge base
- Review pull requests to ensure security, testing, and operational readiness, integrating automated scanning tools like Trivy and Dependabot into the development lifecycle
- Lead team rituals such as design discussions to improve culture, surface new ideas, risks, and strategies across the team
- Participate in sprint planning and refinement sessions to estimate work and negotiate priorities across stakeholders
- Architect and maintain robust CI/CD pipelines using GitHub Actions (Workflows) to automate testing, deployment, and quality assurance across the distributed system
- Define how AI-assisted tooling fits into the team’s workflow, where it accelerates development and increases knowledge
- Encourage others to think critically and systematically by fostering a culture of continuous learning and problem-solving
- Develop and evolve Infrastructure-as-Code (IaC) (e.g., Pulumi, Ansible) to automate the reliable provisioning of the management plane and the secure, automated enrollment of edge nodes (CDPUs) into the fleet
- Containerize and operate distributed-system services with Docker, and drive design and operational maturity of Kubernetes deployments (workload scheduling, autoscaling, networking, and cluster upgrades) across the fleet
- Architect and maintain the distributed artifact registry and the scoped-sync mechanism for content-addressed distribution of platform components across the entire fleet
- Drive the maturity of multi-platform build and release pipelines, ensuring high-integrity, verifiable artifact delivery across diverse architectures (Linux/Darwin)
- Travel up to 25% for customer engagement, architecture reviews, or team collaboration (Middle East)
Requirements
- 7+ years building distributed systems in production, with a track record of shaping whole platform or infrastructure areas rather than individual features
- Deep hands-on experience with Go in high-throughput, concurrent systems, including goroutines, bounded channels for backpressure, context-based cancellation, and graceful shutdown under load
- Working knowledge of message queues and streaming systems (for example NATS/JetStream, Kafka, or similar), including durable delivery, retention and eviction strategies, and the tradeoffs between at-least-once and exactly-once semantics
- Demonstrated experience designing state machines and idempotent operations for systems that must recover deterministically after crashes or partial failures, including write-then-reconcile patterns, rollback-safe writes, and dual durable-versus-running state
- Experience with gRPC/protobuf-based service boundaries and cross-system transfer protocols, including schema versioning and bidirectional session management between distributed components
- Fluency in the failure semantics of distributed systems: supervised restarts with backoff, deterministic reconciliation loops, and designs where durable records and running state never diverge
- Exceptional communication skills across technical and non-technical audiences; you routinely interface with customers, product leaders, and engineering teams to gather requirements, articulate trade-offs, and drive alignment from initial design to delivery
- Has set technical direction others followed and can point to a subsystem and explain what they built and what they’d do differently
- Travel up to 25% for customer engagement, architecture reviews, or team collaboration (Middle East), * Experience operating streaming or data infrastructure embedded within a larger product rather than as a standalone hosted service, with no external broker and no open control ports
- Background contributing to or reviewing API and schema evolution in multi-module, independently versioned protobuf codebases
- Experience building or operating systems for regulated or high-reliability customers
- Background in national security, intelligence, or defense environments with direct understanding of mission-critical operational requirements
- Active security clearance (Secret or above); Top Secret preferred
Physical Requirements
- Prolonged periods of sitting at a desk and working on a computer
- Participation in virtual and in-person meetings
- Ability to attend planned meetings and/or work in classified spaces for extended periods within the specified work regions
- TRAVEL: International travel to Middle East and geopolitical hot zones, up to 25% per calendar year is required. Please apply only if you can meet this requirement
Benefits & conditions
Beyond the opportunity to work alongside practitioners who have operated at the mission edge, joining Orbis means real ownership over outcomes that matter - the chance to build technology that gives the United States Government and its allies a genuine decision advantage, not just another dashboard. From competitive retirement matching to a PTO policy that encourages real time off, we’re committed to creating an environment where our team can do their best work and still have a life outside of it.
Our comprehensive benefits package is designed to meet the diverse needs of our employees and their families. A full list of benefits is shared with candidates following an initial conversation with our HR team, so you’ll have a clear picture of the support and resources available to you as part of the Orbis team.
About the company
Orbis Operations is a technology company that delivers sovereign intelligence and decision capabilities to the United States Government and its allies. We convert complex information and operational systems into decisive outcomes, giving public institutions and critical enterprises the tools to move faster, see clearer, and act first in an increasingly contested world.
We are not a consulting firm and not a systems integrator. Every Orbis team is built around practitioners who have operated at the mission edge, and every engagement is judged on one thing: whether it produces working capability, not slide decks. Solutions forged from experience; that is how we build, and it is how we hire., Orbis works in whatever arrangement the mission requires. Depending on the role, that might mean five days a week on a customer site, full-time in one of our offices, a hybrid schedule, or fully remote work - our team is spread across 25 states and four countries to match. We make sure every team member has the tools and access needed to do the work, wherever that work happens.
Travel is common across many Orbis roles, particularly those supporting customer missions in the field.
Orbis is headquartered in McLean, VA, with additional office locations in Taipei, Taiwan, and Canberra, Australia.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries