> Markdown version of [/jobs/ext/577443-sre-lead-monitoring-consultant](https://www.wearedevelopers.com/jobs/ext/577443-sre-lead-monitoring-consultant). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Lead & Monitoring Consultant - **Company:** VALUE SPECTRUM TECHNOLOGIES LLC - **Location:** Phoenix, AZ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Test Suite, Amazon Web Services, Audit Trail, Microsoft Azure, Bash Shell, Cloud Computing, DevOps, Disaster Recovery, Fraud Prevention and Detection, Monitoring of Systems, Python (Programming Language), Machine Learning, Node.Js, PCI Data Security Standards, Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Runbook, Datadog, Scripting, Google Cloud, Grafana, Multi-Cloud, HybridCloud, Kubernetes, Infrastructure Automation Frameworks, Performance Monitor, Apache Kafka, Cloud Optimization, Terraform, Splunk, Docker, Elk Stack, Golang - **Published:** June 12, 2026 - **Apply:** https://www.dice.com/job-detail/2b7dbe03-67a6-4f6a-a8cb-6fbdd23d2da6 ## About the Role Experience: * * 7+ years in Site Reliability Engineering, DevOps, or infrastructure engineering * 3+ years in SRE leadership roles. * The ideal candidate will possess strong expertise in Java, Node.js, Kafka, AWS Cloud, and modern AIOps/Observability practices. * Implement proactive monitoring and predictive alerting using AIOps platforms and machine learning-driven insights. * 3+ years hands-on experience with Datadog, Splunk, Grafana, and Prometheus. * Strong hands-on experience with Java and Node.js application architectures. * Previous experience in fintech or regulated industries. * Proven track record building SRE practices from scratch. Technical Skills * Deep understanding of SRE principles, error budgets, and SLO/SLI frameworks. * Expertise with cloud platforms (AWS, Azure, or Google Cloud Platform). * Proficiency with Kubernetes, Docker, and infrastructure as code (Terraform, Ansible). * Strong programming/scripting skills (Python, Go, Bash). * Experience with incident management and post-mortem culture. * Knowledge of compliance requirements (SOC 2, PCI-DSS, ISO 27001). Soft Skills * Exceptional leadership and mentoring abilities. * Strong communication and stakeholder management. * Data-driven decision-making approach. * Collaborative mindset with ability to drive cultural change. Preferred Qualifications * Cloud certifications (AWS, Google Cloud Platform, Azure) or Kubernetes certifications (CKA/CKAD). * Experience with ELK stack. * Background in cloud cost optimization. * Multi-cloud or hybrid cloud experience. ## Description SRE Practice Development * Assess operational maturity and build SRE transformation roadmap * Establish SLOs, SLIs, and error budgets for critical services * Design incident management processes and on-call strategies * Implement chaos engineering and resilience testing * Mentor teams on SRE principles and best practices Monitoring & Observability * Deploy and configure Datadog, Splunk, Grafana, and Prometheus * Implement metrics collection, log aggregation, and APM * Build custom dashboards and alerting configurations * Set up anomaly detection and intelligent alerting * Configure automated health checks and remediation * Establish golden signals monitoring (latency, traffic, errors, saturation) Reliability & Compliance * Conduct reliability reviews and performance optimization * Design disaster recovery and failover procedures * Implement security monitoring and audit logging * Configure fraud detection and transaction monitoring * Create runbooks and operational documentation, * Fully configured monitoring stack with Datadog, Splunk, Grafana, and Prometheus * SLO/SLI definitions and error budgets * Custom dashboards, alerting, and automated remediation * Incident management framework and runbooks * Chaos engineering test suite ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)