> Markdown version of [/jobs/ext/173788-platform-reliability-engineer-azure-job-in-irving](https://www.wearedevelopers.com/jobs/ext/173788-platform-reliability-engineer-azure-job-in-irving). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Reliability Engineer, Azure job in Irving - **Company:** Wellfit Technologies Inc. - **Location:** Irving, TX, United States - **Salary:** $130,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** .NET Framework, Application Programming Interfaces (APIs), Application Layers, Application Performance Management, Application Services, Microsoft Azure, Command-Line Interface, Cloud Computing, DevOps, Windows PowerShell, Reliability Engineering, Cloud Services, Prometheus, Runbook, Software Deployment, Software Engineering, SQL Databases, Systems Integration, Datadog, Scripting, Microsoft Power Automate, Cloud Monitoring, System Availability, Grafana, Software Troubleshooting, AngularJS, Google Cloud Functions, Webhooks, Dynatrace, Azure Resource Manager - **Published:** May 25, 2026 - **Apply:** https://jobs.diversity.com/career/2354826/platform-reliability-engineer-azure-texas-tx-irving ## About the Role &bull Hands-on experience supporting production systems in Azure. &bull Strong working knowledge of Azure App Services, Azure Monitor, Application Insights, and Azure production troubleshooting. &bull Experience with DevOps, cloud operations, site reliability, platform engineering, or production support in a hands-on environment. &bull Strong troubleshooting instincts and the ability to work through ambiguous production issues. &bull Comfort working across logs, metrics, traces, alerts, configurations, deployments, and service dependencies. &bull Ability to collaborate with software engineering teams across the stack, including .NET, Angular, SQL, APIs, and cloud services. &bull Experience building or improving dashboards, alerts, runbooks, incident workflows, or operational playbooks. &bull Working knowledge of scripting or automation, preferably with PowerShell, Logic Apps, CLI tooling, or similar technologies. &bull Clear communication skills with the ability to document findings, explain issues, and drive follow-through after incidents. &bull A high-ownership mindset with the ability to create structure, improve processes, and operate effectively in a fast-moving environment. Preferred Experience &bull Azure certifications. &bull Grafana, Prometheus, DataDog, Dynatrace, or similar observability/APM tools. &bull Slack integrations, webhooks, Logic Apps, or incident routing workflows. &bull Azure Front Door, CDN, Function Apps, WebJobs, Service Bus, Event Hub, Event Grid, SQL Pools, App Service Plans, or related Azure services. &bull Experience in healthcare, fintech, payments, or other high-availability environments. &bull Experience in startup, SMB, or scale-up environments where ownership is broad and hands-on. ## Description We are seeking a hands-on Platform Reliability Engineer, Azure to help strengthen the reliability, visibility, and operational maturity of our Azure-based platforms. This role is ideal for someone who enjoys working directly in Azure, improving production systems, troubleshooting issues across infrastructure and application layers, and building practical monitoring and alerting solutions that help teams respond faster and operate more confidently. You do not need to be an expert in every part of the stack on day one. We are looking for someone with strong Azure experience, solid troubleshooting instincts, a DevOps/reliability mindset, and the ability to collaborate closely with engineering teams across systems, services, and applications. What You'll Do &bull Own and improve monitoring, alerting, and observability across Azure-based production systems. &bull Work directly in Azure Monitor, Application Insights, App Services, logs, metrics, traces, and related Azure tooling to troubleshoot reliability and performance issues. &bull Build and refine practical alerting workflows, including Sev0/Sev1 alert routing, escalation paths, and runbook integration. &bull Create and maintain clear, actionable runbooks that help on-call engineers respond confidently to production incidents. &bull Partner with engineering teams to investigate issues across infrastructure, configuration, deployments, services, and application behavior. &bull Support release readiness by improving visibility into critical Azure resources before, during, and after production deployments. &bull Build and maintain dashboards in tools such as Grafana, Azure Monitor, Application Insights, or similar observability platforms. &bull Help configure incident routing integrations, including Slack/webhook-based alert delivery to the appropriate team channels. &bull Automate repeatable operational tasks using PowerShell, Logic Apps, Azure tooling, or similar workflow automation methods. &bull Contribute to RCA documentation, incident follow-up, reliability improvements, and operational playbook development., &bull You understand how production systems are monitored, where alerting gaps exist, and how to improve them. &bull You can work directly in Azure to investigate issues, improve visibility, and support reliable operations. &bull You build practical runbooks, dashboards, and alerting workflows that teams actually use. &bull You collaborate well with engineers, ask strong troubleshooting questions, and help drive issues to resolution. &bull You bring ownership, curiosity, and a builder mindset to a growing platform environment. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)