> Markdown version of [/jobs/ext/2719145-operational-focused-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2719145-operational-focused-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # operational-focused Site Reliability Engineer - **Company:** Cohere Health - **Location:** Boston, MA, United States (Remote available) - **Experience:** Starter - **Salary:** $100,000.0 - $110,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Application Programming Interfaces (APIs), Amazon Web Services, Amazon Elastic Compute Cloud, Automation of Tests, Big Data, Software as a Service, Cloud Computing, Data Architecture, Extract Transform Load (ETL), Shard (Database Architecture), Software Debugging, Distributed Data Store, Memory Management, Identity and Access Management, Python (Programming Language), Microsoft Message Queuing, MySQL, Node.Js, Queueing Systems, RabbitMQ, Reliability Engineering, DataOps, Memory Leaks, Data Streaming, TypeScript, Data Processing, Mern, Apache Spark, AWS Lambda, Indexer, Amazon Virtual Private Cloud (VPC), Backend, Event Driven Architecture, Pyspark, AWS Glue, Cloudwatch, Terraform, Stream Processing, Data Pipelines, Serverless Computing, Amazon Elastic Mapreduce (EMR) - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-ll-cohere-health-8771986 ## About the Role * SaaS Platform Experience: Minimum of 3+ years of hands-on experience operating multi-tenant, cloud-hosted, or cloud-native SaaS platforms at scale. * AWS Cloud Engineering: Deep expertise operating AWS core services, specifically AWS Lambda, Amazon ECS/EKS, Amazon EMR or AWS Glue (for Spark), EC2, VPC networking, IAM permissions, and CloudWatch. * Automation & Data Languages: Professional competency in writing, debugging, and maintaining automation scripts and data tools using Python (including PySpark APIs) and Node.js. * Data Operations: Experience managing and troubleshooting distributed data orchestration pipelines, ETL tools, message queues (e.g., AWS SQS/SNS, RabbitMQ), or stream processing frameworks. * MERN Stack Operations: Deep understanding of the operational lifecycle of JavaScript/TypeScript applications, including memory management, asynchronous runtimes, and Node.js clustering. * Database Administration: Practical experience managing, sharding, indexing, and optimizing production-grade MySQL DB & Athena (RDS or self-hosted). * Infrastructure as Code: Proven ability to deploy and maintain immutable infrastructure utilizing Terraform or OpenTofu. * Healthcare Experience: Minimum 1 year working within HIPAA-regulated environments. Direct experience securing data-at-rest and data-in-transit containing sensitive patient records is preferred. * Education & Experience: Minimum of 4 years of software/systems experience, with at least 1-2 years focused on live cloud operations and distributed data workflow management is preferred. * Crisis Management: Calm under pressure with a methodical approach to identifying and isolating PySpark driver OOM (Out of Memory) errors or data corruption during high-stress outages. Attention to detail and effective communications skills will be critical in working with clients and internal stakeholders is preferred. ## Description This is a remote-first role that may require travel to Boston, MA for new hire onboarding and occasional in-person team meetings and company events., We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will bridge the gap between AWS cloud infrastructure, MERN stack applications, and large-scale data workflows. You will spend roughly 60% of your time on live incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning, and 40% on engineering automated solutions to eliminate operational toil., * Production Operations: Maintain the continuous uptime, scalability, and security of our AWS-hosted MERN applications and backend data architectures. * Serverless Execution: Manage, optimize, and troubleshoot event-driven architectures running on AWS Lambda, focusing on cold-start mitigation, memory allocation, and execution timeouts. * Data Pipeline Execution: Monitor scheduled PySpark data workflows, execute standard operating procedures (SOPs) for large-scale data ingestion, and rapidly triage, rerun, or patch failed data processing jobs. * Incident Management: Participate in a collaborative on-call rotation to rapidly triage, debug, and mitigate live application outages and data flow bottlenecks. * Healthcare Compliance: Maintain strict HIPAA, SOC2, and HITRUST compliance profiles across all runtime environments, storage systems, and data pipelines handling Protected Health Information (PHI). * Toil Elimination: Engineer automated workflows to eliminate repetitive tasks like manual data seeding, infrastructure provisioning, and routine PySpark pipeline recovery steps. * Observability Engineering: Build specialized dashboards and alerts to monitor Node.js event loops, PySpark job execution stages, driver/worker memory leaks, and data pipeline throughput anomalies. * Post-Mortem Culture: Lead blameless post-mortems for operational and data processing failures, translating system crashes into permanent structural fixes. ## Related Videos - [Flex your Energy: Building a Cloud-Native Platform for Renewable Energy Communities](https://www.wearedevelopers.com/videos/1990-flex-your-energy-building-a-cloud-native-platform-for-renewable-energy-communities) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) - [Coding for Good: Achieving social change with an app](https://www.wearedevelopers.com/videos/1645-coding-for-good-achieving-social-change-with-an-app) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)