> Markdown version of [/jobs/ext/2502133-staff-software-engineer-distributed-systems](https://www.wearedevelopers.com/jobs/ext/2502133-staff-software-engineer-distributed-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Distributed Systems - **Company:** Orbis Operations - **Location:** McLean, VA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Automation of Tests, Catalyst (Software), Data Infrastructure, Data Integration, Extract Transform Load (ETL), Linux, Distributed Data Store, Distributed Systems, Github, Protocol Buffers, Scrum Methodology, Queueing Systems, Ansible, Session Management, Data Streaming, Systems Integration, Tripwire, Management of Software Versions, Web Application Frameworks, Datadog, Pulumi, Delivery Pipeline, State Machines, RTSP, Containerization, Kubernetes, Infrastructure Automation Frameworks, Transport Protocols, Apache Kafka, Codebase, Stream Processing, Multiplatform, Docker, Golang - **Published:** August 25, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18062751?backUrl=%2Fcareer%2F18062751%2FStaff-Software-Engineer-Distributed-Systems-Virginia-Mclean ## About the Role * 7+ years building distributed systems in production, with a track record of shaping whole platform or infrastructure areas rather than individual features * Deep hands-on experience with Go in high-throughput, concurrent systems, including goroutines, bounded channels for backpressure, context-based cancellation, and graceful shutdown under load * Working knowledge of message queues and streaming systems (for example NATS/JetStream, Kafka, or similar), including durable delivery, retention and eviction strategies, and the tradeoffs between at-least-once and exactly-once semantics * Demonstrated experience designing state machines and idempotent operations for systems that must recover deterministically after crashes or partial failures, including write-then-reconcile patterns, rollback-safe writes, and dual durable-versus-running state * Experience with gRPC/protobuf-based service boundaries and cross-system transfer protocols, including schema versioning and bidirectional session management between distributed components * Fluency in the failure semantics of distributed systems: supervised restarts with backoff, deterministic reconciliation loops, and designs where durable records and running state never diverge * Exceptional communication skills across technical and non-technical audiences; you routinely interface with customers, product leaders, and engineering teams to gather requirements, articulate trade-offs, and drive alignment from initial design to delivery * Has set technical direction others followed and can point to a subsystem and explain what they built and what they'd do differently * Travel up to 25% for customer engagement, architecture reviews, or team collaboration (Middle East), * Experience operating streaming or data infrastructure embedded within a larger product rather than as a standalone hosted service, with no external broker and no open control ports * Background contributing to or reviewing API and schema evolution in multi-module, independently versioned protobuf codebases * Experience building or operating systems for regulated or high-reliability customers * Background in national security, intelligence, or defense environments with direct understanding of mission-critical operational requirements * Active security clearance (Secret or above); Top Secret preferred Physical Requirements * Prolonged periods of sitting at a desk and working on a computer * Participation in virtual and in-person meetings * Ability to attend planned meetings and/or work in classified spaces for extended periods within the specified work regions * TRAVEL: International travel to Middle East and geopolitical hot zones, up to 25% per calendar year is required. Please apply only if you can meet this requirement ## Description Orbis is seeking a Staff Software Engineer, Distributed Systems to define architecture across the messaging, state, and reconciliation layers of Catalyst, our secure distributed data platform that moves and transforms data reliably. We are looking for someone who has shaped platform architecture and wants to help set technical direction for a team building critical infrastructure., * Define, maintain, and expand the Golang and Python frameworks and libraries that support a large set of batch and/or streaming plugins to allow movement and integration of data * Solving business needs at scale by expanding the patterns of data movement across disparate data sources * Manage the lifecycle of critical components and dependencies, utilizing automated tooling to proactively identify vulnerabilities and facilitate safe, tested upgrades across the ecosystem * Drive the design of high-performance, low-latency architectures optimized for specialized workloads, such as TAK protocols and RTSP media streams * Define architecture across multiple feature areas or subsystems of Catalyst's distributed data plane, including event streaming, queuing, and state machines, and evolve it as the platform scales * Act as a domain expert in one or more critical areas across the data lifecycle (e.g., event streaming, state management, or storage engines) * Define, monitor, and improve on reliability targets and SLOs using Datadog, and lead incident response and postmortems for the core distributed-systems paths * Participate in on-call and incident response for our data platform including leading post-incident reviews, triage, prevention, and mitigation * Document and socialize procedures, naming conventions, and other standards via Notion and Github to maintain a high-quality, searchable engineering knowledge base * Review pull requests to ensure security, testing, and operational readiness, integrating automated scanning tools like Trivy and Dependabot into the development lifecycle * Lead team rituals such as design discussions to improve culture, surface new ideas, risks, and strategies across the team * Participate in sprint planning and refinement sessions to estimate work and negotiate priorities across stakeholders * Architect and maintain robust CI/CD pipelines using GitHub Actions (Workflows) to automate testing, deployment, and quality assurance across the distributed system * Define how AI-assisted tooling fits into the team's workflow, where it accelerates development and increases knowledge * Encourage others to think critically and systematically by fostering a culture of continuous learning and problem-solving * Develop and evolve Infrastructure-as-Code (IaC) (e.g., Pulumi, Ansible) to automate the reliable provisioning of the management plane and the secure, automated enrollment of edge nodes (CDPUs) into the fleet * Containerize and operate distributed-system services with Docker, and drive design and operational maturity of Kubernetes deployments (workload scheduling, autoscaling, networking, and cluster upgrades) across the fleet * Architect and maintain the distributed artifact registry and the scoped-sync mechanism for content-addressed distribution of platform components across the entire fleet * Drive the maturity of multi-platform build and release pipelines, ensuring high-integrity, verifiable artifact delivery across diverse architectures (Linux/Darwin) * Travel up to 25% for customer engagement, architecture reviews, or team collaboration (Middle East) ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025)