> Markdown version of [/jobs/ext/1264057-devops-sre](https://www.wearedevelopers.com/jobs/ext/1264057-devops-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps/SRE - **Company:** Solidus Labs - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Databases, Continuous Integration, DevOps, Amazon DynamoDB, Virtual Private Networks (VPN), Network Troubleshooting, PostgreSQL, Redis, Prometheus, Data Logging, Scripting, Application Enhancement Tool, Load Balancing, Amazon ElastiCache, Snowflake, Grafana, Apache Spark, Git, Amazon Relational Database Service, Gitlab-ci, Kubernetes, Vertica, Functional Programming, Cloudwatch, Terraform, Docker, Databricks - **Published:** July 14, 2026 - **Apply:** https://www.comeet.com/jobs/soliduslabs/E6.007/devopssre/7E.C66 ## About the Role * Proficiency with Terraform, Helm, and GitLab CI (or similar) * Scripting experience with Bash and Python * Familiarity with pub/sub systems (SQS, Kafka, or similar) * Solid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB) * 3+ years of hands-on DevOps / SRE experience * Strong troubleshooting skills across infrastructure, CI/CD, and networking * Strong production experience with Docker and Kubernetes * Experience with monitoring, logging, and alerting systems * Willingness to participate in on-call rotations * GitOps workflows and advanced Git usage * Experience supporting databases such as Postgres, Snowflake, or ClickHouse * Experience with Redis, Airflow, Databricks, Spark/EMR ## Description * We are seeking an experienced New York-based DevOps / Site Reliability Engineer to join our DevOps team and own the reliability, stability, and operational support of our production systems * This role focuses on production ownership, monitoring, incident response, and on-call support, providing critical coverage. You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant * Own the reliability, availability, and performance of our production environments * Operate production Kubernetes (EKS), including cluster upgrades and Helm deployments * Manage scaling and capacity using KEDA, Karpenter, and HPA for resource optimization * Manage AWS Cloud environments including EC2, Lambda, AWS Batch, Elasticache, RDS, and more * Evolve infrastructure as code using Terraform and Helm with security best practices * Support GitLab CI/CD pipelines, resolving deployment issues and improving stability * Design observability systems using Prometheus, Grafana, and EFK to reduce alert fatigue * Solve networking issues involving TLS, Load Balancing, VPCs, NAT, and VPN * Support compliance initiatives and respond to security-related incidents * Leverage AI-powered tools as a standard part of your workflow for automation and productivity * Lead incident response end-to-end, including troubleshooting, mitigation, and resolution * Perform deep-dive RCA to drive long-term corrective and preventive actions * Participate in on-call rotations to provide consistent operational coverage ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)