Staff Site Reliability Engineer

Loera, R Company
United States
21 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$180,000.0 - $240,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence DevOps Python (Programming Language) Reliability Engineering TypeScript Golang

Job description

  • Design and Deliver High-Impact Solutions: Design and implement systems that enhance reliability, observability, traceability, and incident management, ensuring the platform scales effectively
  • Lead Strategic Initiatives: Take ownership of cross-team collaborations and drive impactful projects by providing technical leadership and guidance
  • Partner Across Teams: Collaborate with engineers from AI/ML, Data, Platform, and Product teams to develop best-in-class services
  • Partner with engineers from AI/ML, Data, Platform, Product, and other groups to deliver best-in-class services
  • Establish Standards and Best Practices: Define and enforce production standards, processes, and tools to ensure operational excellence
  • Champion Reliability Goals: Advocate for and implement SLIs, SLOs, and other reliability-focused metrics across the engineering organization
  • Mentorship and Knowledge Sharing: Guide and mentor team members, fostering technical growth and helping to develop the next generation of engineering leaders
  • Innovate and Inspire: Drive continuous improvement by bringing creative ideas and challenging the status quo

Requirements

  • 7+ years of experience in Production Engineering, Backend Engineering, SRE, DevOps or similar role
  • Strategic visionary: Your strong technical background enables you to look beyond solving the immediate problem, planning for the future.
  • Proficient Problem-Solver: Strong coding ability in at least one language (e.g., Golang, Python, Java, Typescript) with the capability to solve complex issues through code
  • Track Record of Success: Demonstrated experience delivering medium to large-scale projects that drive meaningful improvements in platform reliability and scalability
  • Reliability Expertise: Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management
  • Strong Communicator: Excellent verbal and written communication skills with the ability to influence and collaborate across technical and non-technical teams
  • Fast-Paced Experience: Familiarity with working in dynamic, reliability-focused production environments (preferred)

Benefits & conditions

You’ll get competitive perks and benefits, from health & wellness to equity, to help you bring your best self to work.

For US based applicants:

  • The US base salary range for this full-time position is $180,000 - $240,000 annually+ equity + benefits
  • Our salary ranges are determined by role, level and location

About the company

Attentive is the AI marketing platform for 1:1 personalization redefining the way brands and people connect. We’re the only marketing platform that combines powerful technology with human expertise to build authentic customer relationships. By unifying SMS, RCS, email, and push notifications, our AI-powered personalization engine delivers bespoke experiences that drive performance, revenue, and loyalty through real-time behavioral insights.

Recognized as the #1 provider in SMS Marketing by G2, Attentive partners with more than 8,000 customers across 70+ industries. Leading global brands like Crate and Barrel, Urban Outfitters, and Carter’s work with us to enable billions of interactions that power tens of billions in revenue for our customers.

With a distributed global workforce and employee hubs in New York City, San Francisco, London, and Sydney, Attentive’s team has been consistently recognized for its performance and culture. We’re proud to be included in Deloitte’s Fast 500 (four years running!), LinkedIn’s Top Startups, Forbes’ Cloud 100 (five years running!), Inc.’s Best Workplaces, and the Human Rights Campaign Foundation’s Corporate Equality Index!

About the RoleOur Platform Infrastructure team is the backbone of everything we do at Attentive, providing a resilient and cost-effective platform that seamlessly handles billions of events from over 100 million customers daily. We own everything from compute, persistence, and networking to observability and deployments. Joining our team offers a high-growth career opportunity to collaborate with some of the world’s most talented engineers in a high-performance, high-impact culture.

As part of the Infrastructure and Platform organization, the Production Engineering Team is focused on delivering a fast and reliable platform that empowers Attentive engineers to deliver solutions quickly and safely. We build scalable systems that automate routine tasks so we can focus on other impactful efforts. Reliability, scalability, and security are our areas of expertise. We focus on release, observability, and cost optimization. Our mission is to create robust platforms and tools that allow stakeholders to concentrate on delivering exceptional products.

As a Staff Engineer, you will take a strategic role in designing and implementing solutions that enhance the reliability and scalability of our systems, while mentoring others and influencing technical roadmaps across the organization.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:00 min

Misconceptions about TypeScript safety capabilities

Simone Sanfratello · JS Congress

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all