Senior Software Engineer II, Developer Experience / Operational Excellence

United States Digital Space LLC
Greater London, UK
26 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Clean Code Principles Artificial Intelligence Amazon Web Services Data Analysis Cloud Computing Python (Programming Language) Systems Development Life Cycle Software Engineering Datadog Grafana Infrastructure as Code (IaC) Information Technology
+4 more
Low Latency Terraform New Relic (SaaS) Pagerduty

Job description

the company is hiring a Senior Software Engineer II to join our Operational Excellence (OPX) team within the Developer Experience organization., * Own and improve incident management tooling and on-call health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden

  • Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency
  • Contribute to AI-driven operational tooling that goes beyond triage, building toward autonomous remediation where AI detects issues, takes corrective action, and self-recovers with minimal human involvement
  • Drive incident prevention by identifying systemic patterns and ruthlessly eliminating operational toil. You have deep empathy for on-call engineers and a bias toward making their lives better
  • Partner directly with product engineering teams to diagnose reliability gaps, reduce their operational burden, and help them adopt best practices for running their services
  • Define and champion operational excellence best practices across engineering through guardrails, scorecards, and standards that help teams run their services reliably by default
  • Champion, role model, and embed the company’s cultural principles (Focus on Customer Success, Build for the Long Term, Adopt a Growth Mindset, Be Inclusive, Win as a Team) as we scale globally and across new offices

Requirements

  • 8+ years of experience designing and building products in a software engineering team
  • Bachelor’s Degree in Computer Science/Engineering or equivalent practical experience
  • 3+ years of experience on infrastructure and/or platform engineering focused teams
  • Expertise in Observability and reliability, operational metrics and data analysis
  • Proven track record architecting monitoring frameworks, SLO platforms, and automated response workflows Datadog (or equivalent observabilty tooling like New Relic, Grafana).
  • Proven experience working on large-scale enterprise software applications
  • Experience in Developer Experience (DevEx) & Internal Portals: Designing and implementing solutions/tools that centralise and simplify engineering operations.
  • Familiarity with cloud platforms (AWS, GCP or the like)
  • Experience in implementing AI-driven automation across the software development lifecycle (SDLC) to reduce developer friction, automate repetitive technical tasks, and accelerate time-to-delivery. Routinely applies AI tools across your workflow
  • Familiarity with Experienced at writing high quality code (Go, Python or equivalent) focused on infrastructure, deployment and operations challenges
  • Experience mentoring and supporting engineers and role modeling engineering practices within a technical lead capacity
  • Proactive growth mindset, always looking at ways to improve the status quo, * B.S. in Computer Science or related technical discipline
  • Strong communication skills and desire to collaborate across teams
  • Experience with incident management tooling (Incident.io PagerDuty or equivalent)
  • Experienced with Infrastructure as Code (Iac) - Terraform

Benefits & conditions

At the company, we build for the people who keep the global economy moving. We want owners, not passengers, which is why our rewards are designed to fuel high-impact builders. Our compensation program delivers above-market total compensation through a combination of base salary, performance-based bonus/variable pay, and equity (for eligible roles) in a high-growth public company. We meaningfully differentiate pay for our top performers, who have the opportunity to earn above-market compensation that can outpace the broader market over time.

Beyond compensation, we provide the foundations that enable long-term success:

  • a flexible, employee-led remote model
  • a professional development stipend
  • comprehensive health and parental leave plans
  • and more

About the company

the company (NYSE: IOT) is the pioneer of the Connected Operations Cloud, which is a platform that enables organizations that depend on physical operations to harness Internet of Things (IoT) data to develop actionable insights and improve their operations. At the company, we are helping improve the safety, efficiency and sustainability of the physical operations that power our global economy. Representing more than 40% of global GDP, these industries are the infrastructure of our planet, including agriculture, construction, field services, transportation, and manufacturing - and we are excited to help digitally transform their operations at scale.

Working at the company means you’ll help define the future of physical operations and be on a team that’s shaping an exciting array of product solutions, including Video-Based Safety, Vehicle Telematics, Apps and Driver Workflows, and Equipment Monitoring. As part of a recently public company, you’ll have the autonomy and support to make an impact as we build for the long term., DevEx is responsible for the engineering environment that a globally distributed engineering org relies on every day, from build and deploy systems to development tooling and AI-assisted workflows that help teams move quickly and confidently.

Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate and mitigate incidents faster.

OPX is focused on raising the bar for system stability, resilience, and reliability across the company – building automated safeguards that protect production, ensuring engineers get the right signals at the right time, and giving teams the visibility they need to stay in control of their services. We’re also investing in AI-driven operational tooling and partnering directly with product engineering teams to strengthen their operational posture.

You should apply if:

  • You want to impact the industries that run our world: The software, firmware, and hardware you build will result in real-world impact - helping to keep the lights on, get food into grocery stores, and most importantly, ensure workers return home safely.
  • You want to build for scale: With over 2.3 million IoT devices deployed to our global customers, you will work on a range of new and mature technologies driving scalable innovation for customers across industries driving the world’s physical operations.
  • You are a life-long learner: We have ambitious goals. Every Samsarian has a growth mindset as we work with a wide range of technologies, challenges, and customers that push us to learn on the go.
  • You believe customers are more than a number: the company engineers enjoy a rare closeness to the end user and you will have the opportunity to participate in customer interviews, collaborate with customer success and product managers, and use metrics to ensure our work is translating into better customer outcomes.
  • You are a team player: Working on our the company Engineering teams requires a mix of independent effort and collaboration. Motivated by our mission, we’re all racing toward our connected operations vision, and we intend to win - together.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

Videos

See all

Related articles

See all