SRE Lead & Monitoring Consultant
VALUE SPECTRUM TECHNOLOGIES LLC
Phoenix, AZ, United States
2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
Test Suite
Amazon Web Services
Audit Trail
Microsoft Azure
Bash Shell
Cloud Computing
DevOps
Disaster Recovery
Fraud Prevention and Detection
Monitoring of Systems
Python (Programming Language)
+25 more
Machine Learning
Node.Js
PCI Data Security Standards
Performance Tuning
Reliability Engineering
Site Reliability Engineering Practices
Ansible
Prometheus
Runbook
Datadog
Scripting
Google Cloud
Grafana
Multi-Cloud
HybridCloud
Kubernetes
Infrastructure Automation Frameworks
Performance Monitor
Apache Kafka
Cloud Optimization
Terraform
Splunk
Docker
Elk Stack
Golang
Job description
SRE Practice Development
- Assess operational maturity and build SRE transformation roadmap
- Establish SLOs, SLIs, and error budgets for critical services
- Design incident management processes and on-call strategies
- Implement chaos engineering and resilience testing
- Mentor teams on SRE principles and best practices
Monitoring & Observability
- Deploy and configure Datadog, Splunk, Grafana, and Prometheus
- Implement metrics collection, log aggregation, and APM
- Build custom dashboards and alerting configurations
- Set up anomaly detection and intelligent alerting
- Configure automated health checks and remediation
- Establish golden signals monitoring (latency, traffic, errors, saturation)
Reliability & Compliance
- Conduct reliability reviews and performance optimization
- Design disaster recovery and failover procedures
- Implement security monitoring and audit logging
- Configure fraud detection and transaction monitoring
- Create runbooks and operational documentation, * Fully configured monitoring stack with Datadog, Splunk, Grafana, and Prometheus
- SLO/SLI definitions and error budgets
- Custom dashboards, alerting, and automated remediation
- Incident management framework and runbooks
- Chaos engineering test suite
Requirements
Experience:
-
- 7+ years in Site Reliability Engineering, DevOps, or infrastructure engineering
- 3+ years in SRE leadership roles.
- The ideal candidate will possess strong expertise in Java, Node.js, Kafka, AWS Cloud, and modern AIOps/Observability practices.
- Implement proactive monitoring and predictive alerting using AIOps platforms and machine learning-driven insights.
- 3+ years hands-on experience with Datadog, Splunk, Grafana, and Prometheus.
- Strong hands-on experience with Java and Node.js application architectures.
- Previous experience in fintech or regulated industries.
- Proven track record building SRE practices from scratch.
Technical Skills
- Deep understanding of SRE principles, error budgets, and SLO/SLI frameworks.
- Expertise with cloud platforms (AWS, Azure, or Google Cloud Platform).
- Proficiency with Kubernetes, Docker, and infrastructure as code (Terraform, Ansible).
- Strong programming/scripting skills (Python, Go, Bash).
- Experience with incident management and post-mortem culture.
- Knowledge of compliance requirements (SOC 2, PCI-DSS, ISO 27001).
Soft Skills
- Exceptional leadership and mentoring abilities.
- Strong communication and stakeholder management.
- Data-driven decision-making approach.
- Collaborative mindset with ability to drive cultural change.
Preferred Qualifications
- Cloud certifications (AWS, Google Cloud Platform, Azure) or Kubernetes certifications (CKA/CKAD).
- Experience with ELK stack.
- Background in cloud cost optimization.
- Multi-cloud or hybrid cloud experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago
LM
Luis Minvielle
What’s the Difference between a Junior, Mid, and Senior Developer?
about 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago