DevOps/SRE
Solidus Labs
New York, NY, United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.comeet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source
Tech stack
Airflow
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Databases
Continuous Integration
DevOps
Amazon DynamoDB
Virtual Private Networks (VPN)
Network Troubleshooting
PostgreSQL
Redis
+19 more
Prometheus
Data Logging
Scripting
Application Enhancement Tool
Load Balancing
Amazon ElastiCache
Snowflake
Grafana
Apache Spark
Git
Amazon Relational Database Service
Gitlab-ci
Kubernetes
Vertica
Functional Programming
Cloudwatch
Terraform
Docker
Databricks
Job description
- We are seeking an experienced New York-based DevOps / Site Reliability Engineer to join our DevOps team and own the reliability, stability, and operational support of our production systems
- This role focuses on production ownership, monitoring, incident response, and on-call support, providing critical coverage. You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant
- Own the reliability, availability, and performance of our production environments
- Operate production Kubernetes (EKS), including cluster upgrades and Helm deployments
- Manage scaling and capacity using KEDA, Karpenter, and HPA for resource optimization
- Manage AWS Cloud environments including EC2, Lambda, AWS Batch, Elasticache, RDS, and more
- Evolve infrastructure as code using Terraform and Helm with security best practices
- Support GitLab CI/CD pipelines, resolving deployment issues and improving stability
- Design observability systems using Prometheus, Grafana, and EFK to reduce alert fatigue
- Solve networking issues involving TLS, Load Balancing, VPCs, NAT, and VPN
- Support compliance initiatives and respond to security-related incidents
- Leverage AI-powered tools as a standard part of your workflow for automation and productivity
- Lead incident response end-to-end, including troubleshooting, mitigation, and resolution
- Perform deep-dive RCA to drive long-term corrective and preventive actions
- Participate in on-call rotations to provide consistent operational coverage
Requirements
- Proficiency with Terraform, Helm, and GitLab CI (or similar)
- Scripting experience with Bash and Python
- Familiarity with pub/sub systems (SQS, Kafka, or similar)
- Solid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB)
- 3+ years of hands-on DevOps / SRE experience
- Strong troubleshooting skills across infrastructure, CI/CD, and networking
- Strong production experience with Docker and Kubernetes
- Experience with monitoring, logging, and alerting systems
- Willingness to participate in on-call rotations
- GitOps workflows and advanced Git usage
- Experience supporting databases such as Postgres, Snowflake, or ClickHouse
- Experience with Redis, Airflow, Databricks, Spark/EMR
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.comeet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
4 days ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
about 2 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago