Senior/Staff Infrastructure & Platform Engineer (Bay Area)

Fortanix
Santa Clara, CA, United States
1 day ago
Apply on apply.workable.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure C++ (Programming Language) Cloud Computing Configuration Management Computer Programming Continuous Integration Data Centers Linux DevOps Disaster Recovery Distributed Systems
+23 more
Elasticsearch Virtual Private Networks (VPN) Python (Programming Language) Network Architecture Reliability Engineering Ansible Single Sign-On Software Engineering Ceph (Software) Private Cloud Environment Data Logging System Availability IT Architecture Kubernetes Deployment Automation Cassandra Bare Metal Apache Kafka Build Tools Hardware Infrastructure Terraform Service Stack Jenkins

Job description

We are looking for a Senior/Staff Infrastructure & Platform Engineer to help architect, build, and evolve the infrastructure and tooling behind our products and engineering organization.

This is a highly technical, hands-on role for an engineer who can work across cloud, on-premises, data center, Kubernetes, networking, CI/CD, and software development. You will help define how our infrastructure should be designed-not simply implement predefined solutions.

You’ll work across internal infrastructure, production environments, and customer deployments, solving complex infrastructure problems and building scalable, automated solutions. The ideal candidate has a broad understanding of infrastructure architecture, deep expertise in Kubernetes and Linux, strong programming skills, and the technical leadership to take ambiguous problem statements and drive them from design through implementation.

What You’ll Do

  • Design, build, and operate scalable, reliable, and secure infrastructure across cloud, Kubernetes, on-premises, and hybrid environments.
  • Identify complex infrastructure and engineering productivity challenges and drive solutions from problem definition through architecture, implementation, and production.
  • Build and improve internal developer platforms, tools, and services that simplify development, deployment, and operational workflows.
  • Design, build, and optimize CI/CD platforms and pipelines, improving build and test performance, reliability, scalability, and developer feedback cycles.
  • Automate provisioning, configuration, deployment, testing, monitoring, and operational workflows to reduce engineering toil and improve efficiency.
  • Improve the developer experience across the software lifecycle, partnering with engineering teams to identify bottlenecks and deliver high-impact improvements.
  • Develop reusable Infrastructure as Code, automation, and platform capabilities that establish consistent and scalable engineering practices.
  • Build systems with strong reliability, observability, resilience, security, and disaster recovery capabilities.
  • Troubleshoot complex infrastructure and platform issues, participate in incident response and root-cause analysis, and drive long-term solutions to systemic problems.
  • Evaluate technologies and architectural approaches, make pragmatic technical recommendations, and establish infrastructure and engineering standards and best practices.
  • Lead technical initiatives spanning multiple engineering teams, influence architecture and engineering practices, and contribute to the long-term infrastructure and developer productivity strategy.

Requirements

  • 6+ years of experience in infrastructure, platform engineering, software engineering, SRE, DevOps, or related fields.
  • Strong programming and software engineering fundamentals with experience developing production-quality tools and services.
  • Deep production Kubernetes experience, including troubleshooting complex environments.
  • Experience with cloud infrastructure such as AWS, Azure, or GCP and/or on-premises infrastructure; self-managed or bare-metal.
  • Experience with Infrastructure as Code such as Terraform, Ansible, or similar technologies.
  • Experience designing and improving CI/CD platforms and deployment automation.
  • Strong understanding of Linux, networking, distributed systems, and infrastructure architecture.
  • Strong experience with server provisioning, configuration management, patching, upgrades, and lifecycle management.
  • Hands-on experience with Ansible, Chef, or similar configuration management and automation technologies.
  • Strong programming skills in at least one language such as Go, Python, Rust, or C++.
  • Experience with observability, monitoring, logging, and reliability engineering practices.
  • Ability to lead complex technical initiatives, evaluate tradeoffs, and drive scalable, maintainable solutions through production.
  • Strong communication and collaboration skills across technical and cross-functional teams.

Nice to Have

  • Experience building internal developer platforms or developer productivity tooling.
  • Experience operating Kubernetes in on-premises, private cloud, or bare-metal environments.
  • Experience with Kubernetes operators, controllers, or Kubernetes-native tooling.
  • Experience defining engineering productivity or platform reliability metrics.
  • Kubernetes CNI/CSI and networking/storage expertise.
  • Experience with distributed technologies such as Cassandra, etcd, Ceph, Kafka, or Elasticsearch/OpenSearch.
  • Jenkins architecture or administration experience.
  • Experience with VPN, SSO, infrastructure security, or network infrastructure.
  • Hardware/server provisioning and data center operations experience.
  • Enterprise customer deployment or customer-facing infrastructure experience.

What Success Looks Like

  • You solve complex infrastructure problems across multiple layers of the technology stack.

About the company

In today’s world, where data spreads across various clouds and devices, traditional security measures aren’t enough. Businesses need a dynamic approach to defend against constant cyber threats and ensure agile data security. Fortanix leads the way in data-centric cybersecurity for hybrid multicloud environments, using advanced cryptography, encryption, and confidential AI solutions.

As data breaches become more frequent and traditional defenses fall short, we focus on data exposure management to keep your information safe. Our unified data security platform addresses vulnerabilities in hybrid multicloud environments, defends against threats, and makes it easier to discover, assess, and fix data exposure risks. Whether implementing a Zero Trust model or preparing for the post-quantum computing era, we help businesses worldwide protect their most sensitive data, wherever it is.

Our commitment to solving the world’s toughest data security challenges has earned Fortanix multiple Cybersecurity Excellence and Innovation Awards, as well as recognition from industry giants such as NVIDIA, Microsoft, Intel, ServiceNow, and Snowflake.

Our team includes industry leaders and cryptography experts, creating a culture of trust, innovation and collaboration where every voice is valued. Recognized as a Great Place to Work, we’re looking for passionate individuals to help us shape the future of data security and work towards a safer digital future., * You can move seamlessly between architecture, software development, automation, and hands-on troubleshooting.

  • Improve the reliability and scalability of shared infrastructure and CI/CD systems.

  • You build reusable solutions that improve reliability, scalability, and engineering productivity.

  • Proactively identify systemic technical problems before they become significant business or operational issues.

  • You understand how Kubernetes and the infrastructure beneath it work and use that knowledge to build reliable, scalable systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on apply.workable.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all