Staff Site Reliability Engineer (FedRAMP)

Okta, Inc.
San Francisco, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 267K

Job location

San Francisco, United States of America

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Apache HTTP Server
Tomcat
Bash
Cloud Computing
System Configuration
Continuous Integration
Software Debugging
Linux
DNS
Hypertext Transfer Protocols (HTTP)
Apache Hypertext Transfer Protocol Server
Python
Network Protocols
Nginx
Public Key Infrastructure
Reliability Engineering
Ansible
TCP/IP
Tcpdump
Wireshark
SSL Certificate Management
Load Balancing
Okta
System Availability
GIT
Puppet
Terraform
Docker
Go

Job description

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

The Team

The Site Reliability team is dedicated to architecting and owning the foundational infrastructure tooling and CI/CD platforms that support Okta's SRE ecosystem. In this development-focused role, you will leverage a modern tech-stack to build durable, automated systems that maximize platform reliability and engineering velocity.

The ideal candidate is someone who enjoys analyzing systems and identifying areas of opportunity to improve system performance, availability and capacity. They are part systems administrator, part network administrator, and part developer.

What you'll be doing

  • Maintain a highly available cloud infrastructure edge for the Okta identity platform
  • Automate AWS infrastructure with Terraform and/or Chef
  • Evolve the system by introducing changes to improve efficiency, scalability, and velocity

Requirements

  • 8+ years of operations experience configuring, deploying, monitoring and troubleshooting applications and infrastructure in the cloud
  • 8+ years administering or operating within a Linux environment, strong experience using Linux based tooling, and ability to debug systems level problems
  • In-depth understanding of TCP/IP, HTTP, Load Balancing, DNS and other networking protocols
  • Solid understanding and experience with Apache httpd, nginx, Apache Tomcat, or similar
  • Strong problem solving and debugging skills coupled with a desire to take on ownership and responsibility
  • Proficiency in Bash, Python, Golang, or similar. Experienced with git
  • Experience working with Terraform, Ansible, Chef, Puppet or similar automation tools
  • Excellent written and verbal communication skills
  • Willingness to work on-call

And extra credit if you have experience in any of the following!

  • Experience working in a security-oriented cloud environment
  • Experience debugging software using gdb, strace, ltrace, tcpdump, Wireshark, etc.
  • Experience working with Docker and Kubernetes
  • Experience with PKI / certificate management, * Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

Benefits & conditions

3.93.9 out of 5 stars San Francisco, CA Hybrid work $194,000 - $267,000 a year, We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

Apply for this position