> Markdown version of [/jobs/ext/912849-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/912849-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Veloc Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Layers, Microsoft Azure, Backup Devices, Bash Shell, Software as a Service, Cloud Computing, Cloud Engineering, Continuous Integration, Linux, DevOps, Disaster Recovery, Distributed Systems, Failover, Python (Programming Language), Performance Tuning, Windows PowerShell, Role-Based Access Control, Reliability Engineering, Ansible, Prometheus, Datadog, Data Logging, Google Cloud, Cloud Platform System, Delivery Pipeline, Grafana, Infrastructure as Code (IaC), Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Bicep, Terraform, Splunk, Azure Resource Manager, Docker, Golang - **Published:** June 30, 2026 - **Apply:** https://www.dice.com/job-detail/afec95df-a0b8-49e7-8e83-267ab8d87f9a ## About the Role 7+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Production Operations roles. Strong experience operating workloads in cloud environments such as Microsoft Azure, AWS, or Google Cloud. Hands-on experience with Kubernetes, Docker, CI/CD pipelines, and Infrastructure as Code tools. Strong scripting and automation skills using Python, Bash, PowerShell, Go, or similar languages. Experience with observability and monitoring platforms such as Datadog, Grafana, Prometheus, or Splunk. Strong understanding of networking, Linux/Windows administration, distributed systems, and cloud-native architectures. Experience with incident response, production troubleshooting, and operational governance. Strong communication skills and ability to collaborate across engineering and business teams. Preferred Qualifications Experience supporting multi-tenant SaaS environments. Experience with Terraform, Bicep, ARM templates, or Ansible. Familiarity with GitOps and modern deployment strategies such as canary or blue/green deployments. Experience working within regulated or compliance-driven environments. ## Description Senior Site Reliability Engineer - combination of deep operational expertise and hands-on engineering ability. The majority of your time (~70%) will be focused on owning the reliability, availability, scalability, and operational excellence of the cloud infrastructure and SaaS platforms powering our business. The remaining ~30% puts you directly in the platform engineering flow: building automation, improving deployment pipelines, and driving reliability initiatives from conception through production. You will write and review automation code, contribute to architecture and deployment discussions, and collaborate closely with product engineering teams to ensure operational and reliability decisions are made correctly the first time. Key Responsibilities Reliability Engineering & Operations (~40% of role) Own day-to-day monitoring, alerting, operational health, and on-call support for mission-critical SaaS platforms and cloud infrastructure. Lead major incident response activities including escalation coordination, root cause analysis, and postmortem reviews. Design and maintain high-availability, failover, backup, and disaster recovery procedures; validate RTO/RPO targets regularly. Investigate and resolve production incidents end-to-end across infrastructure, platform, and application layers. Automation & Platform Engineering (~30% of role) Design, implement, and maintain Infrastructure as Code (IaC), deployment automation, and CI/CD pipeline improvements. Develop tooling and automation to reduce operational toil and improve engineering productivity. Partner with development teams to improve deployment safety, release reliability, and operational scalability. Drive standardization of cloud infrastructure, operational engineering practices, and deployment governance. Observability & Performance Optimization (~15% of role) Build and maintain monitoring, logging, tracing, and alerting capabilities across distributed systems. Establish service-level objectives (SLOs), SLIs, and error budget policies. Identify and remediate performance bottlenecks, scaling issues, and infrastructure inefficiencies. Analyze operational telemetry and trends to improve reliability and capacity planning. Security, Compliance & Architecture (~15% of role) Implement operational security best practices including RBAC, least privilege access, and infrastructure hardening. Ensure compliance with SOC 2, HIPAA, GDPR, and organizational security standards. Participate in architecture reviews and operational readiness assessments for new services and platforms. Mentor junior engineers on reliability engineering, cloud operations, automation, and incident management best practices. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)