Lead Site Reliability Engineer

ALLEN FAMILY VENTURES, LLC
New York, United States
2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$179,000.0 - $226,000.0
Working hours
Regular working hours
Job source

Tech stack

JavaScript (Programming Language) Amazon Web Services Big Data Databases Software Debugging Programming Tools Distributed Systems Python (Programming Language) Software Engineering Datadog Build Management Kubernetes
+6 more
Build Tools Cloudwatch Terraform Docker Golang Programming Languages

Job description

We’re looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees. What you’ll be doing

Reporting to the Engineering Manager of Infrastructure, you’ll:

  • Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
  • Reduce operational toil by turning manual processes into reliable, repeatable workflows
  • Build internal tooling and platforms that enable safe self-service changes for other engineers
  • Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
  • Implement and evolve systems for deploying and running applications in Kubernetes
  • Contribute to architecture decisions across infrastructure, reliability, and security
  • Write and review production-quality code
  • Participate in on-call rotations-but focus on building systems that prevent incidents, not just respond to them

Requirements

  • 10+ years of experience in infrastructure, SRE, or software engineering roles
  • Strong software engineering skills-you build systems, not just scripts
  • Experience managing production infrastructure at scale (cloud + containerized systems)
  • Experience with Infrastructure as Code (e.g., Terraform)
  • Experience running and troubleshooting distributed systems (Docker/Kubernetes)
  • Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
  • Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
  • Experience participating in on-call rotations and improving systems based on incidents
  • Strong communication and collaboration skills

You might be a great fit if you

  • Default to automation over manual processes
  • See repetitive work and immediately want to eliminate it
  • Think in terms of systems, failure modes, and long-term scalability
  • Care about building infrastructure that other engineers can use safely and confidently
  • Enjoy working in a small team with high ownership and impact

Nice to have

  • Experience running Kubernetes in production at scale
  • Deep familiarity with AWS
  • Experience building internal platforms or developer tooling
  • Background in distributed systems or large-scale data systems

We’re a lean team, so your impact will be felt immediately, and opportunities will grow as the company scales up. If this all sounds like a good fit for you, why not join us?

Benefits & conditions

Alloy is committed to fair and equitable compensation practices. Below is the anticipated starting base compensation range for this role; however, pay may vary depending on job-related knowledge, in-demand skills, relevant experience, and/or geography. In addition to a competitive base salary, this position is also eligible for equity awards in the form of stock options (ISOs) as well as a competitive total benefits package. Your recruiter will be happy to walk you through the details and what compensation could look like for you specifically!

This position has a salary range of $179,000 to $226,000. Benefits and Perks

  • Unlimited PTO and flexible work policy
  • Employee stock options
  • Medical, dental, vision plans with HSA (monthly employer contribution) and FSA options
  • 401k with 100% match up to 4% of annual employee compensation
  • Eligible new parents receive 16 weeks of paid parental leave
  • Home office stipend for new employees
  • Annual Learning & Development annual stipend
  • Well-being benefits include access to ClassPass, OneMedical, UrbanSitter, and Spring Health
  • Hybrid work environment: employees are expected to work Tuesdays through Thursdays from our HQ in Union Square, Manhattan. Tasty lunches catered from a variety of local restaurants and frequent employee-organized cultural events contribute to our positive office energy. On Monday/Friday most employees Zoom into work from home while some take advantage of the quieter office.

About the company

Alloy helps solve the identity risk problem for companies that offer financial products by enabling them to outpace fraud and confidently serve more people around the world. Over 800 of the world’s largest financial institutions and fintechs turn to Alloy to take control of fraud, credit, and compliance risk, and grow with the clearest picture of their customers.

Through our values: Be Bold, Get Scrappy, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we’ve been continuously recognized and named one of Inc. Magazine’s Best Workplaces, Forbes America’s Best Startup Employers, Best Fintech to Work for by American Banker, year after year.

Check out our investors and read more about us here. About the team

Alloy’s Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.

Our challenge isn’t just scale-it’s making that scale reliable, secure, and operable with less manual work.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

6:58 min

Building engineering communities and finding technical inspiration

Videos

See all

Related articles

See all