> Markdown version of [/jobs/ext/2383680-senior-site-reliability-and-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2383680-senior-site-reliability-and-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability and Infrastructure Engineer - **Company:** Treeswift Inc - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $160,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon S3, Continuous Integration, Customer Data Management, Directed Acyclic Graph (Directed Graphs), Data Infrastructure, Extract Transform Load (ETL), Software Debugging, Linux, DevOps, Machine Learning, Reliability Engineering, Web Applications, Real Time Systems, State Machines, Kubernetes, Machine Learning Operations, Functional Programming, Amazon Simple Queue Service (SQS), Terraform, Data Pipelines - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=0e863866c06ccd59 ## About the Role * You are an experienced software engineer where the last 7-10 years required significant time on observability, systems/infrastructure engineering, SRE, or DevOps (ideally in a cloud environment). * Ability to reason about architecture end-to-end and articulate your thoughts with product impact in mind (data movement, execution, failure handling, and operational visibility). * Hands-on experience with infrastructure-as-code (Terraform and similar) and using it to deliver reliable environments. * Experience with container orchestration and debugging in practice (Kubernetes and/or ECS/container-based deployments). * Strong Linux debugging skills and demonstrated ability to investigate production issues with logs/metrics and clear hypotheses. * Empathy and communication: you can collaborate effectively with engineers across teams (especially the data platform team) and explain tradeoffs clearly., * Experience working in early-stage or fast-moving environments where ownership and processes evolve quickly. * Experience with Apache Airflow and/or Astronomer. * Experience with AWS, although other cloud providers are fine. (DuploCloud experience is also helpful.) * Experience with geospatial/imagery/lidar/point-cloud style domains. * ML Ops skills (model deployment/inference reliability, packaging, CI/CD for model artifacts, and operational observability for inference pipelines). Work location This is a full-time, hybrid role based out of our Lower Manhattan, NYC office (2 days per week in person, currently pinned to Tuesdays and Wednesdays). ## Description * You'll be our first full-time SRE/infrastructure engineer, so we'll look to you for leadership on how to improve and scale our infrastructure to support each part of the platform. Our data pipeline, machine learning training platform, and web app could all benefit from further productionization. * Help us scale and harden the platform that schedules our pipelines, runs machine learning training, and hosts our web app. We run Apache Airflow on Astronomer with DAGs that orchestrate high-volume processing across AWS and Kubernetes, including machine learning inference inside pipeline tasks. You will build the observability and reliability foundations that let us run this system confidently as customer data volume grows: monitoring, alerting, performance/cost visibility, and clear operational practices. * Stay curious, collaborative, and cross-functional while also taking ownership of problems. We translate complex, real-world requirements from a critical industry into high-quality data products, so understanding the business holistically is key. We take pride in managing complexity and providing high-fidelity data that our customers can use to make better-informed decisions., * Partner with the data platform and engineering teams to understand how changes propagate across pipeline execution (Astronomer-hosted Airflow DAGs), containerized workers (Kubernetes), and AWS services (S3, SQS, Lambda, Step Functions, ECS). * Design and implement reliability and observability for high-volume pipeline operations, including: + actionable monitoring/alerting for DAG/task failures and reruns + visibility into operational workflows like flight orchestration (including DLQ/failed-message alerting and notification pathways) + dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health * Own CI/CD guardrails for production changes: build/deploy validation and safe rollout mechanics for Astronomer deployments (image builds pushed to ECR, and Airflow configuration updates via Astronomer CLI variable updates) * Make machine learning inference operations more reliable and observable: + instrument inference runs executed inside pipeline runners (model checkpoint resolution, S3 sync behavior, thresholds and fallback behavior, and output correctness) + add operational visibility for inference outcomes (e.g., unknown classification rates, fallback usage, and failure modes) * Create operational tooling and continuously improve systems ('leave it better than you found it'), including: + runbooks, incident learnings, and engineering standards for debugging at scale + automate away toil in deployment and operations workflows as we learn what hurts most On-call / incident response There is not currently an established on-call rotation for this platform, and the pipelines do not require real-time processing. That said, you'll still help lead reliability improvements and operational readiness-so the team has faster diagnosis, better alerts, and safer releases when issues do occur. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)