SRE/DevOps Engineer

Versana LLC
New York, United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Shift work
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) .NET Framework Amazon Web Services Microsoft Azure Software as a Service Continuous Integration Linux DevOps Distributed Systems Elasticsearch Github
+19 more
Monitoring of Systems Python (Programming Language) Reliability Engineering Data Streaming Datadog Google Cloud Spring Cloud Grafana Reliability of Systems Cloudformation Gitlab-ci Infrastructure Automation Frameworks Bicep Apache Kafka Terraform Docker Jenkins Golang Programming Languages

Job description

Design, implement and enhance system observability and monitoring tools Monitor system performance, create incident response plans, and implement observability practices to gain insights into system behavior. Implement and monitor service-level objectives (SLOs) and indicators. Improve system reliability and resiliency. Conduct post-incident reviews and implement necessary changes to prevent system failures. Assist teams in implementing observability tools and leveraging available telemetry data to troubleshoot and resolve incidents and problems. Leverage observability and event management to improve key incident management metrics, such as mean time to detect and mean time to restore services. Continually optimize systems and workflows by improving architecture, infrastructure, automation, CI/CD, and observability. Collaborate with developers to ensure applications are designed with DevOps best practices in mind. Participate in a rotating on-call schedule for weekend releases and being available to respond to production issues outside of regular working hours, including weekends and holidays.

Requirements

Versana is seeking a motivated SRE/DevOps Engineer with strong observability experience to join our growing Platform Engineering squad. The squad’s goal is to manage public cloud, improve DevOps practices, and monitor Versana’s real-time syndicated loan data platform. The ideal candidate will have a deep understanding of cloud-native applications, distributed computing, CI/CD implementation, observability tools and practices., 5+ years of experience as a Site Reliability Engineer or similar role. 3+ years of work experience with public cloud (Azure, AWS or Google Cloud Platform). 3+ years of direct experience with observability tools like Datadog, Elasticsearch, and Grafana Labs, etc. 3+ years of experience with containerization and orchestration technologies like Docker and Kubernetes. 2+ years of experience in development and management of CI/CD pipelines (e.g., Azure DevOps, Gitlab CI/CD, Github Actions, Jenkins, etc). 2+ years of experience with Infrastructure-as-code tools like Terraform, Azure Bicep, Cloud Formation, etc. 1+ years of experience with site reliability tools like Gremlin, Chaos Mesh, or similar. Proven track record leveraging core observability concepts, end-user monitoring, and infrastructure monitoring with SaaS solutions. Experience with messaging services like Kafka or Azure Event Hubs. Good understanding of the Linux operating system.

Nice to Have: Experience in at least one coding language such as Java, JavaScript, Python, GoLang, or .NET. Certifications in cloud technologies. Experience with Azure cloud or Azure DevOps. Experience with Datadog or similar modern observability tools.

About the company

Versana is an industry-backed data and technology company on a mission to make the syndicated loan market better. By digitally capturing agent banks’ data on a real-time basis and centralizing it onto a single platform, Versana provides unprecedented transparency into loan level details and portfolio positions, bringing efficiency and velocity to the entire market. Through our platform, participants can rest assured they are accessing the loan market’s most credible source of deal

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · World Congress 2024

Videos

See all

Related articles

See all