> Markdown version of [/jobs/ext/2718389-sre-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2718389-sre-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE/DevOps Engineer - **Company:** Versana LLC - **Location:** New York, United States - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), .NET Framework, Amazon Web Services, Microsoft Azure, Software as a Service, Continuous Integration, Linux, DevOps, Distributed Systems, Elasticsearch, Github, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Data Streaming, Datadog, Spring Cloud, Grafana, Reliability of Systems, Cloudformation, Containerization, Gitlab-ci, Infrastructure Automation Frameworks, Bicep, Apache Kafka, Terraform, Docker, Jenkins, Golang, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/sre-devops-engineer-versana-7994248 ## About the Role Versana is seeking a motivated SRE/DevOps Engineer with strong observability experience to join our growing Platform Engineering squad. The squad's goal is to manage public cloud, improve DevOps practices, and monitor Versana's real-time syndicated loan data platform. The ideal candidate will have a deep understanding of cloud-native applications, distributed computing, CI/CD implementation, observability tools and practices., * 5+ years of experience as a Site Reliability Engineer or similar role. * 3+ years of work experience with public cloud (Azure, AWS or GCP). * 3+ years of direct experience with observability tools like Datadog, Elasticsearch, and Grafana Labs, etc. * 3+ years of experience with containerization and orchestration technologies like Docker and Kubernetes. * 2+ years of experience in development and management of CI/CD pipelines (e.g., Azure DevOps, Gitlab CI/CD, Github Actions, Jenkins, etc). * 2+ years of experience with Infrastructure-as-code tools like Terraform, Azure Bicep, Cloud Formation, etc. * 1+ years of experience with site reliability tools like Gremlin, Chaos Mesh, or similar. * Proven track record leveraging core observability concepts, end-user monitoring, and infrastructure monitoring with SaaS solutions. * Experience with messaging services like Kafka or Azure Event Hubs. * Good understanding of the Linux operating system. Nice to Have: * Experience in at least one coding language such as Java, JavaScript, Python, GoLang, or .NET. * Certifications in cloud technologies. * Experience with Azure cloud or Azure DevOps. * Experience with Datadog or similar modern observability tools. ## Description * Design, implement and enhance system observability and monitoring tools * Monitor system performance, create incident response plans, and implement observability practices to gain insights into system behavior. * Implement and monitor service-level objectives (SLOs) and indicators. * Improve system reliability and resiliency. * Conduct post-incident reviews and implement necessary changes to prevent system failures. * Assist teams in implementing observability tools and leveraging available telemetry data to troubleshoot and resolve incidents and problems. * Leverage observability and event management to improve key incident management metrics, such as mean time to detect and mean time to restore services. * Continually optimize systems and workflows by improving architecture, infrastructure, automation, CI/CD, and observability. * Collaborate with developers to ensure applications are designed with DevOps best practices in mind. * Participate in a rotating on-call schedule for weekend releases and being available to respond to production issues outside of regular working hours, including weekends and holidays. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Valven Atlas: Engineering Intelligence That Delivers](https://www.wearedevelopers.com/videos/1660-valven-atlas-engineering-intelligence-that-delivers) - [Monoskope: Developer Self-Service Across Clusters](https://www.wearedevelopers.com/videos/329-monoskope-developer-self-service-across-clusters) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)