Senior Site Reliability Engineer

JUUL Labs, Inc.
Durham, NC, United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$185,000.0 - $227,000.0
Working hours
Regular working hours
Job source

Tech stack

Link Aggregation (Ethernet) Amazon Web Services Amazon S3 Bash Shell Cloud Computing Cloud Storage Configuration Management Disaster Recovery Data Flow Control Github Identity and Access Management Virtual Private Networks (VPN)
+37 more
Python (Programming Language) Log Analysis Windows Servers Routing Network Segmentation Windows PowerShell Prism (Software) Role-Based Access Control Reliability Engineering Aws Command Line Interface (CLI) Session Manager SubSystems Security Information and Event Management TCP/IP Virtual Local Area Networks Data Logging Scripting Transport Layer Security Load Balancing Cloud Platform System Data Ingestion Autoscaling Istio System Availability Boto3 Multi-Cloud HybridCloud Amazon Virtual Private Cloud (VPC) Cloudformation Containerization Kubernetes Information Technology Patch Management Nutanix Cloudwatch Restful APIs Terraform Splunk

Job description

A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance of Juul’s hybrid cloud infrastructure (Nutanix, AWS/GCP). This involves leading automation efforts, architecting for reliability, and acting as the final escalation point for critical incidents to ensure the platform is scalable and efficient., * Design, deploy, and maintain enterprise-scale Nutanix AHV clusters and Prism Central for multi-cluster management

  • Expert-level proficiency with Nutanix CLI (nCLI and acli) for advanced operations, troubleshooting, and automation
  • Develop automation scripts using Nutanix REST APIs, Python SDK, PowerShell, and Terraform for infrastructure-as-code
  • Create and manage VM templates, golden images, and standardized deployment catalogs for consistent provisioning
  • Design disaster recovery solutions using Leap, Protection Domains, cross-cluster replication, and metro clustering
  • Implement network micro-segmentation using Nutanix Flow and configure RBAC, encryption, and security hardening
  • Lead L3 troubleshooting using advanced diagnostics, log analysis (CVM, Genesis), NCC health checks, and cluster service resolution
  • Configure high availability, VM affinity rules, QoS policies, and optimize performance for mission-critical workloads
  • Manage AHV networking with OVS bridges, VLANs, bonds, LACP and implement resource reservations and workload balance.
  • Design, deploy, and maintain hybrid cloud infrastructure across Nutanix HCI, AWS, and GCP platforms
  • Architect and implement multi-cloud solutions ensuring high availability, scalability, and disaster recovery

Cloud Platform Engineering

  • Architect and deploy enterprise-scale, highly available multi-cloud solutions across AWS and GCP with multi-region/multi-account strategies
  • Expert-level proficiency with AWS CLI, GCP CLI, SDK, boto3, and Python for advanced automation and infrastructure orchestration
  • Design AWS Organizations and GCP Organization hierarchies with consolidated billing, IAM policies, and centralized governance
  • Configure and manage AWS Systems Manager (SSM) including Session Manager, Run Command, State Manager, and Automation for centralized fleet operations
  • Implement centralized logging using CloudWatch/CloudTrail and GCP Cloud Logging with S3/Cloud Storage aggregation
  • Integrate AWS and GCP with Splunk using HEC, CloudWatch subscriptions, Pub/Sub, Dataflow, and cloud-specific add-ons for SIEM correlation
  • Design and deploy advanced load balancing solutions with AWS ALB/NLB/ELB and GCP Cloud Load Balancing including SSL termination and auto-scaling
  • Develop infrastructure-as-code using Terraform, CloudFormation, CDK for repeatable multi-cloud deployments and CI/CD pipelines
  • Configure AWS SSO, cross-account IAM roles, GCP Workload Identity, and federated access for centralized identity management
  • Design VPC architectures with AWS Transit Gateway/PrivateLink and GCP Shared VPC/VPC peering for hybrid connectivity
  • Manage containerized workloads using EKS, GKE, ECS, Cloud Run with service mesh, observability, and security best practices
  • Implement disaster recovery using AWS Backup, Cross-Region Replication, GCP snapshots, and multi-region failover strategies
  • Lead L3 troubleshooting using CloudWatch Insights, GCP Cloud Trace, VPC Flow Logs, X-Ray, and vendor support escalation
  • Perform cost optimization through Reserved Instances, Committed Use Discounts, rightsizing, and automated resource lifecycle management

System Administration

  • Administer and support Windows Server and Unix/Linux environments in production and non-production settings
  • Perform OS-level hardening, patch management, and security compliance across heterogeneous systems
  • Automate routine administrative tasks using PowerShell, Bash, Python, or similar scripting languages
  • Manage GitHub organization settings, user permissions, repository access controls, and monitor GitHub Actions workflows and repository health across multiple teams
  • Configure Splunk forwarders, heavy forwarders and other integrations for data ingestion from cloud and on-premises sources

Requirements

  • 8-12+ years infrastructure experience with 8+ years in Nutanix HCI and enterprise cloud AWS/GCP)
  • Expert-level skills in Python, PowerShell, Bash scripting, infrastructure-as-code (Terraform/CloudFormation), and container orchestration (Kubernetes, EKS/GKE)
  • Proven experience managing enterprise-scale environments, hybrid cloud migrations, disaster recovery, and L3 critical incident management
  • Strong networking knowledge (TCP/IP, VLANs, routing, VPN), security hardening, and compliance frameworks (ITIL)
  • Strategic thinker with exceptional analytical and troubleshooting abilities for complex multi-layer infrastructure issues
  • Excellent communication skills to translate technical concepts to executives and non-technical stakeholders
  • Calm under pressure during critical outages with meticulous attention to security, compliance, and configuration management
  • Self-motivated continuous learner committed to staying current with evolving cloud technologies and automation opportunities
  • Available for on-call rotations with strong documentation skills and customer service orientation
  • Certifications (plus): Nutanix NCP/NCAP, AWS Solutions Architect Professional, AWS DevOps
  • Professional, GCP Professional Cloud Architect, Terraform, * Bachelor’s or master’s degree in computer science/IT

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Dental insurance
  • Life insurance
  • Employee assistance program
  • Disability insurance, JUUL LABS PERKS & BENEFITS:
  • People. Work with talented, committed and supportive teammates
  • Equity and performance bonuses. Every employee is a stakeholder in our success
  • Cell phone subsidy, commuter benefits and discounts on JUUL products
  • Excellent medical, dental and vision, disability, and life insurance, plus family support, wellness, legal, and employee assistance program benefits
  • 401(k) plan with company matching
  • Plus biannual discretionary performance bonuses

About the company

Juul Labs’s mission is to transition the world’s billion adult smokers away from combustible cigarettes, eliminate their use, and combat underage usage of our products. We have the opportunity to address one of the world’s most intractable challenges through a commitment to exceptional quality, research, design, and innovation. Backed by leading technology investors, we are committed to the same excellence when it comes to hiring great talent.

We are a diverse team that is united by this common purpose and we are hiring the world’s best engineers, scientists, designers, product managers, operations experts, and customer service and business professionals. If the opportunity to build your career is compelling, read on for more details.

MUST LIVE in either Mountain View, CA, Washington D.C., Austin, TX, or Durham, NC.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

1:12 min

Validating risky service requests safely via dry runs

Modood Alvi · World Congress 2025

3:46 min

Navigating a career in cloud transformation consulting

Piet Van Dongen · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all