DevOps/SRE

Solidus Labs
New York, NY, United States
about 1 month ago
Apply on www.comeet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Databases Continuous Integration DevOps Amazon DynamoDB Virtual Private Networks (VPN) Network Troubleshooting PostgreSQL Redis
+19 more
Prometheus Data Logging Scripting Application Enhancement Tool Load Balancing Amazon ElastiCache Snowflake Grafana Apache Spark Git Amazon Relational Database Service Gitlab-ci Kubernetes Vertica Functional Programming Cloudwatch Terraform Docker Databricks

Job description

  • We are seeking an experienced New York-based DevOps / Site Reliability Engineer to join our DevOps team and own the reliability, stability, and operational support of our production systems
  • This role focuses on production ownership, monitoring, incident response, and on-call support, providing critical coverage. You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant
  • Own the reliability, availability, and performance of our production environments
  • Operate production Kubernetes (EKS), including cluster upgrades and Helm deployments
  • Manage scaling and capacity using KEDA, Karpenter, and HPA for resource optimization
  • Manage AWS Cloud environments including EC2, Lambda, AWS Batch, Elasticache, RDS, and more
  • Evolve infrastructure as code using Terraform and Helm with security best practices
  • Support GitLab CI/CD pipelines, resolving deployment issues and improving stability
  • Design observability systems using Prometheus, Grafana, and EFK to reduce alert fatigue
  • Solve networking issues involving TLS, Load Balancing, VPCs, NAT, and VPN
  • Support compliance initiatives and respond to security-related incidents
  • Leverage AI-powered tools as a standard part of your workflow for automation and productivity
  • Lead incident response end-to-end, including troubleshooting, mitigation, and resolution
  • Perform deep-dive RCA to drive long-term corrective and preventive actions
  • Participate in on-call rotations to provide consistent operational coverage

Requirements

  • Proficiency with Terraform, Helm, and GitLab CI (or similar)
  • Scripting experience with Bash and Python
  • Familiarity with pub/sub systems (SQS, Kafka, or similar)
  • Solid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB)
  • 3+ years of hands-on DevOps / SRE experience
  • Strong troubleshooting skills across infrastructure, CI/CD, and networking
  • Strong production experience with Docker and Kubernetes
  • Experience with monitoring, logging, and alerting systems
  • Willingness to participate in on-call rotations
  • GitOps workflows and advanced Git usage
  • Experience supporting databases such as Postgres, Snowflake, or ClickHouse
  • Experience with Redis, Airflow, Databricks, Spark/EMR

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.comeet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all