SITE RELIABILITY ENGINEER

Stellent IT LLC
New York, NY, United States
6 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$200,000.0 - $300,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Distributed Systems Python (Programming Language) PostgreSQL Machine Learning Redis TypeScript Workflow Management Systems Delivery Pipeline Kubernetes Apache Kafka Free and Open-Source Software
+1 more
Terraform

Requirements

Looking for 7-10 years of experience, ideally with a strong technical focus.

Must be proficient in Python and Typescript, with familiarity in Kubernetes, Helm, Terraform, and Terragrunt., Experience driving systematic application improvements across complex, distributed systems with high- availability requirements for a high- growth start- up

Extended period working at high- growth start- ups (no recent Big Tech)

Prior SWE experience and strong coding abilities (Python and TypeScript)

Hard skills

Hands- on experience automating single- tenant infrastructure provisioning and optimising deployment pipelines for application code or machine learning models

IaC and Orchestration tools including Kubernetes, Helm, Terraform, and Terragrunt

Experience with the rest of our tech stack (e. g. , AWS, PostgreSQL, Redis, and Kafka)

Soft skills

Ability to own complex systems and solve complex problems independently

Defines standards for operational excellence, e. g. , automating manual workflows

Miscellaneous

Experience in regulated industries (health tech given HIPAA complications)

Open- source contributions and LinkedIn recommendations for others

Benefits & conditions

Traits to avoid

Bootcamps and no university degrees, recent consulting, or traditional finance

Frequent career moves with multiple short, 1- year stints

Compensation and Logistics

Open to candidates in San Francisco or New York, with occasional travel required.

Timeline and Urgency

Looking to hire one person in Q1 and potentially another in Q2.

Three-stage interview process: team chat, technical interview, and on-site assessment.

Pain Points

Need for automation to reduce manual workload associated with onboarding new customers.

The current team is small, underscoring the need for individuals who can handle diverse tasks.

Ideal Candidate Profile

Prefer candidates with startup experience and a generalist, SRE background.

Background in backend development transitioning to infrastructure is highly valued. Infrastructure Responsibilities

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.

About the company

Client is solving complex challenges in healthcare through robust and reliable Medical Intelligence. Our AI-driven solutions significantly enhance workflows for hospitals and clinics, enabling patients to access medication faster while substantially boosting healthcare providers’ revenues. We are developing a suite of software solutions in conjunction with major healthcare systems to create fast, effective, and affordable access to a provider for every patient in the United States.

“Based in New York or San Francisco, offering a strong compensation range of $200-300k. The hiring team has a strong track record of successful placements (c.5 in recent months).

Position requires a balance of operational and software engineering skills.

Role involves optimizing database accesses, handling production issues, and improving application performance.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

5:37 min

Extensibility and programmability features of the PostgreSQL database

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

5:00 min

Exploring the specific workplace responsibilities of staff software engineers

Jan Giacomelli · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all