software engineer and SRE.

VALUE SPECTRUM TECHNOLOGIES LLC
Phoenix, AZ, United States
3 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Border Gateway Protocol Cloud Computing Continuous Integration DDoS Mitigation Domain Name System (DNS) Fault Tolerance Monitoring of Systems IP Routing Subnetting Virtual Private Networks (VPN)
+20 more
Routing Network Segmentation Network Service Prometheus Zero Trust Network Access TCP/IP Datadog Load Balancing Computer Network Operations Cloud Platform System System Availability Grafana Reliability of Systems Amazon Virtual Private Cloud (VPC) Cloudformation Low Latency Cloudwatch Firewall Services Module Terraform Software Version Control

Job description

We are seeking a Network SRE to ensure the reliability, scalability, and performance of cloud and hybrid network platforms.

This role applies SRE principles to networking by shifting from manual network operations to automated, observable, and resilient network services.

The ideal candidate is a network engineer who thinks like a software engineer and SRE., Network Reliability Engineering

  • Define SLIs, SLOs, and Error Budgets for network services.

Design networks for:

  • High availability
  • Fault tolerance
  • Low latency
  • Predictable performance

Improve network reliability while reducing operational toil.

Cloud & Hybrid Networking

Architect and operate AWS networking:

  • VPCs, Subnets, Route Tables
  • Transit Gateway
  • NAT, IGW
  • PrivateLink, VPC Endpoints

Design hybrid connectivity:

  • VPN
  • Direct Connect

Support multi-account and multi-region architectures.

Network Observability & Monitoring

Build deep network observability using:

  • VPC Flow Logs
  • CloudWatch
  • Datadog
  • Prometheus / Grafana

Analyze packet loss, latency, and throughput.

Implement proactive alerting based on SLOs.

Correlate network signals with application performance.

Automation & Infrastructure as Code

Automate network provisioning and changes using:

  • Terraform / CloudFormation

Implement CI/CD for network changes.

Reduce manual configuration and human error.

Version-control network definitions.

Incident Response & Troubleshooting

Lead network-related incident response.

Perform deep root-cause analysis for:

  • Packet drops
  • Routing issues
  • DNS failures
  • Load balancer degradation

Participate in on-call rotation and post-incident reviews.

Drive permanent fixes rather than workarounds.

Security & Traffic Management

Design and enforce:

  • Network segmentation
  • Zero-Trust principles
  • Firewall rules (Security Groups, NACLs)

Implement secure ingress/egress patterns.

Support DDoS protection (AWS Shield, WAF).

Work with Security teams on audits and remediation.

Performance & Capacity Planning

Conduct traffic modeling and capacity forecasting.

Tune load balancers (ALB, NLB).

Optimize routing and failover strategies.

Validate resilience through failure testing.

Collaboration & Enablement

Partner with:

  • Cloud Platform teams
  • Application SREs
  • Security & Infra teams

Enable application teams with network best practices.

Requirements

Must-Have:

  • Strong networking fundamentals (TCP/IP, DNS, BGP, routing)
  • AWS networking expertise
  • SRE concepts & practices
  • Network observability & monitoring
  • Infrastructure as Code
  • Production incident handling experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all