> Markdown version of [/jobs/ext/1373785-staff-software-engineer](https://www.wearedevelopers.com/jobs/ext/1373785-staff-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - **Company:** SnapCommerce Inc. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Data Stores, Software Debugging, Distributed Systems, Prometheus, Software Engineering, Data Logging, Grafana, Kubernetes, Dynatrace - **Published:** July 22, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/p6l8xrpu5p ## About the Role * Extensive experience in SRE, software, and platform engineering, including operating, debugging, and reasoning about large-scale distributed systems in production * Proven track record in incident leadership and reliability work, including running complex incidents, writing postmortems, and delivering systemic changes that stick * Hands-on experience with observability stacks (metrics, logs, traces), including tools such as Prometheus/VictoriaMetrics, Loki, Grafana, and OpenTelemetry, alongside distributed tracing * Sound judgement to identify when an architecture is unsuitable for a given workload and to define what should replace it * Strong software engineering skills, ideally in Go, with deep experience in AWS and Kubernetes * A mindset that treats repeated manual work as a problem to solve, with automation as second nature * Ability to act as a technical reference for senior engineers and drive change across teams without formal authority, through design rigour and credibility * Comfortable operating in ambiguous, high-impact environments, with the ability to plan a year or two ahead for platform needs Nice to have * Experience building or operating a large-scale time-series, logging, or continuous-profiling platform (e.g. VictoriaMetrics, Loki, Pyroscope, or similar) * Strong understanding of the cost and performance trade-offs of data stores, and how these shape architectural decisions ## Description _As a Staff Software Engineer / Staff Platform Engineer, you'll join our platform area, where we build the foundations every product team relies on to ship quickly and run safely in production. You'll work alongside a team of strong senior engineers, setting technical direction, raising the reliability bar, and helping the whole team sharpen its skills - when they hit a wall, they'll escalate to you. _What the role involves _Reliability & Operational Excellence _ * Make operational processes - deployments, upgrades, migrations - routine, safe, and reversible * Lead incident response end-to-end, owning postmortems and turning findings into lasting systemic improvements * Automate repeated manual work, treating recurring tasks as issues to be resolved rather than accepted _Observability Depth _ * Build monitoring and alerting that reflects customer impact, so any engineer can move from "something broke" to "here's why" in minutes * Own the reliability and efficiency roadmap, identifying the biggest gaps and coordinating fixes across teams _Technical Leadership & Culture _ * Set technical direction through design reviews, RFCs, and pairing, raising the bar for strong senior engineers * Champion rigour and consistency across the platform, ensuring teams understand trade-offs, not just rules * Run blameless postmortems that surface causes rather than blame, helping the team learn and improve each cycle ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)