Member of Technical Staff (Software Engineer, Cloud Infrastructure)

Perplexity AI
New York, United States
3 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Disaster Recovery Distributed Systems Python (Programming Language) Routing Peering Runbook Software Engineering Load Balancing Amazon Virtual Private Cloud (VPC)
+3 more
Kubernetes Low Latency Terraform

Job description

The Cloud Infrastructure team owns the foundational cloud primitives and deployment models that power Perplexity’s products, from multi-tenant public cloud to single-tenant and on-premises solutions for enterprise customers., * Own the roadmap and technical strategy for agent-driven cloud infrastructure management.

  • Design and operate Perplexity’s cloud networking fabric, including VPC architectures, private connectivity, and peering with hyperscalers and neocloud providers to support low-latency, high-throughput AI workloads.
  • Architect and scale compute platforms (Kubernetes/EKS, autoscaling groups, and mixed CPU/GPU fleets) to efficiently serve online request traffic and background workloads across regions.
  • Build and maintain secure, isolated deployment topologies for multi-tenant, single-tenant, and customer-owned cloud (BYOC) environments, including cross-account networking, identity, and policy guardrails.
  • Implement and evolve multi-region strategies for availability, failover, and data locality, including traffic routing, regional capacity planning, and disaster recovery playbooks.
  • Partner with security to deliver enterprise controls such as BYOK/KMS integrations, network isolation, and auditability required for regulated customers.
  • Develop automation, tooling, and runbooks that make day-2 operations (provisioning, upgrades, incident response) predictable and repeatable for Perplexity products across all environments.

Requirements

  • Deep experience designing and operating cloud infrastructure on AWS (VPC design, routing, security groups, load balancing, private connectivity).
  • Strong background with Kubernetes/EKS and container orchestration, including multi-cluster, multi-region, or multi-account setups.
  • Hands-on experience with cloud networking and peering (VPC peering, Transit Gateway, private link/service endpoints, or similar constructs in neoclouds).
  • Experience building or operating secure, isolated environments for enterprise customers (single-tenant, BYOC, or on-prem), ideally including BYOK/KMS integrations and compliance constraints.
  • Proficiency with infrastructure as code (Terraform) and strong software engineering skills in at least one of Python, Go, or Rust for automation and tooling.
  • Strong debugging and incident management skills across distributed systems (networking, compute, and platform layers), with a track record of driving root-cause analysis and long-term fixes.
  • 7+ years of industry experience building and operating production cloud infrastructure, including leading the design of complex systems or migrations.

About the company

As Perplexity grows its Computer and Enterprise products, this team builds and operates the security, isolation, and compliance layers that customers depend on. We provide the deployment topologies, multi-region infrastructure, and core services that enable Perplexity to run reliably and efficiently for both consumer traffic and large enterprises.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

5:28 min

Bitcoin scaling and the transition to peer-to-peer transactions

Jad Wahab · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all