> Markdown version of [/jobs/ext/525694-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/525694-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** BOSTON PUBLIC HEALTH COMMISSION - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $145,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Audit Trail, Microsoft Azure, Cloud Computing, Computer Programming, Continuous Integration, DevOps, Distributed Systems, Fault Tolerance, Github, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, Network Segmentation, Octopus Deploy, Open Source Technology, Reliability Engineering, Prometheus, Runbook, Software Engineering, Datadog, Pulumi, Cloud Platform System, Real Time Systems, Grafana, Mttr, Multi-Cloud, Reliability of Systems, HybridCloud, Cloudformation, Kubernetes, Low Latency, Deployment Automation, Machine Learning Operations, Terraform, Dynatrace, Devsecops, Docker, Jenkins, Vulnerability Analysis, Microservices - **Published:** June 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5c83684cbe530e29 ## About the Role Do you have experience in Tooling?, 3+ years in Site Reliability Engineering, DevOps, or infrastructure roles Experience owning production systems at scale Strong background in incident response and postmortem processes Experience in fast-paced startup or high-growth environments Technical Skills Strong programming ability in Python, Go, or similar Deep experience with AWS, GCP, or Azure infrastructure Strong Kubernetes and Docker expertise in production environments Experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation) Hands-on experience with CI/CD systems (GitHub Actions, ArgoCD, Jenkins, etc.) Strong knowledge of observability tools (Prometheus, Grafana, Datadog, ELK) Understanding of distributed systems and system design principles Soft Skills Strong reliability-first mindset with focus on failure scenarios Calm and structured thinking during production incidents Clear communication across technical and non-technical teams Strong documentation and runbook writing ability High ownership and autonomy in system management Preferred Qualifications Experience with AI/ML infrastructure or GPU-based workloads Background in real-time systems or low-latency architectures Experience with chaos engineering practices Knowledge of AI observability (model latency, drift, throughput metrics) Multi-cloud or hybrid cloud experience Contributions to open-source infrastructure or SRE tooling Early-stage startup experience (Seed to Series A), Suggested improvements for reliability or performance Optional: GitHub, SRE projects, or technical writing samples ## Description We are hiring a proactive, systems-minded Site Reliability Engineer to ensure that LockedIn AI's production systems remain highly reliable, scalable, and performant. This role sits at the intersection of infrastructure, software engineering, and operations. You will be responsible for maintaining the reliability of real-time AI systems that serve over 1 million users globally. You will own production stability across cloud infrastructure, Kubernetes environments, AI inference systems, APIs, and real-time services. Key Responsibilities System Reliability & Performance Own reliability, availability, and performance of production systems Define and manage SLIs, SLOs, and error budgets aligned with product goals Design fault-tolerant and self-healing system architectures Continuously optimize system latency, throughput, and resource usage Identify bottlenecks and improve system performance at scale Infrastructure as Code & Cloud Systems Design and manage cloud infrastructure across AWS, GCP, or Azure Implement Infrastructure as Code using Terraform, Pulumi, or CloudFormation Manage Kubernetes clusters supporting microservices and AI workloads Ensure infrastructure is versioned, reproducible, and scalable Optimize cloud costs while maintaining performance and reliability Observability & Monitoring Build observability systems using metrics, logs, and traces (Prometheus, Grafana, Datadog, ELK, etc.) Design alerting systems that reduce noise and improve incident response Implement distributed tracing for microservices and AI pipelines Monitor AI-specific metrics including latency, GPU usage, and token throughput Provide real-time visibility into system health and performance Incident Response & Reliability Engineering Lead incident response for production outages and system failures Participate in on-call rotations and coordinate cross-team resolution efforts Conduct blameless postmortems with actionable improvements Build runbooks, playbooks, and escalation procedures Track MTTR and reliability trends to drive continuous improvement CI/CD & Release Engineering Design and maintain CI/CD pipelines for safe and automated deployments Implement canary releases, blue-green deployments, and rollback systems Ensure AI model deployments follow the same reliability standards as application code Add testing and validation gates to prevent production issues Improve deployment velocity while maintaining system stability Security & Compliance Implement secure infrastructure practices including IAM, encryption, and secrets management Maintain network segmentation and audit logging Ensure privacy-first infrastructure design across all systems Manage vulnerability scanning, patching, and system hardening Collaborate with security teams to enforce DevSecOps practices ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)