Ceph Cluster Development Engineer (C++ Focus)

Fortinet Inc.
Santa Clara, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$179,000.0 - $219,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon S3 Health Informatics Software Bug Management C++ (Programming Language) Cloud Computing Cloud Engineering Cloud Storage Software Documentation Profiling Computer Programming Continuous Delivery
+28 more
Data Centers Software Debugging Linux DevOps Disaster Recovery File Systems Distributed Data Store Distributed Systems Failover GNU Debuggers Python (Programming Language) Network Virtualization Performance Tuning Ansible Prometheus Memory Leaks Subsystems System Programming TCP/IP Management of Software Versions Ceph (Software) System Availability Grafana Perf (Linux) Containerization Kubernetes Database Replication Block Storage

Job description

We are seeking a highly skilled Ceph Cluster Development & Operations Engineer with strong expertise in C++ systems programming to design, extend, and maintain enterprise-scale Ceph distributed storage clusters. The role involves deep development in Ceph core subsystems (RADOS, OSD, RGW, MDS), performance optimization, and operational excellence across multi-site, multi-zone architectures.

You will work closely with system architects, SREs, and cloud infrastructure teams to ensure the reliability, scalability, and security of mission-critical storage systems deployed across multiple data centers and Kubernetes environments.

Key Responsibilities

  • Design, build, and operate large-scale Ceph clusters including RADOS, RGW, RBD
  • Contribute to or extend Ceph core components written in C++ (e.g., OSD, RGW, librados, BlueStore, MGR modules).
  • Profile and optimize performance across network, disk I/O, and replication layers (PG placement, CRUSH rules, BlueStore tuning).
  • Develop automation and tooling for cluster lifecycle management (deployment, upgrades, scaling, failover, and recovery).
  • Integrate Ceph with Kubernetes (via Rook-Ceph, CSI drivers) and CI/CD pipelines for continuous delivery.
  • Implement and validate multi-site replication and disaster recovery architectures for high availability.
  • Develop and maintain secure storage solutions using dm-crypt, KMS integration, and CephX authentication.
  • Build observability pipelines using Prometheus, Grafana, and custom exporters for metrics and health analytics.
  • Write and maintain SOPs, automation scripts, and system documentation to support production-grade operations.
  • Collaborate with upstream Ceph community or maintain in-house forks for feature development and bug fixes.

Requirements

  • Strong proficiency in C++ (C++11 or later), with experience in large-scale distributed systems or kernel-adjacent development.
  • Deep understanding of Ceph architecture and its core components: MON, OSD, MGR, RGW, MDS, and CRUSH maps.
  • Proficient in Linux systems programming, debugging (gdb, perf, valgrind), and performance profiling.
  • Experience with Python or Go for tooling and automation.
  • Strong foundation in data replication, erasure coding, and consistency models in distributed storage.
  • Hands-on experience with Kubernetes, Rook-Ceph, Helm, Ansible, and related DevOps tools.
  • Familiarity with TCP/IP, HTTP/S3 APIs, block storage (RBD/iSCSI), and object storage semantics.
  • Ability to conduct root-cause analysis and lead performance investigations under production environments.

Preferred Skills

  • Contributions to the Ceph open-source project or prior experience modifying Ceph source code.
  • Experience with multi-site replication, object versioning, compliance retention, or legal hold features.
  • Background in distributed storage systems, file systems, or cloud storage platforms.
  • Familiarity with containerized environments, network virtualization, and cloud-native observability stacks.
  • Excellent technical documentation and communication skills in English.

Must be authorized to work in the U.S. without sponsorship.

Benefits & conditions

The US base salary range for this full-time position is $179,000-$219,000. Fortinet offers employees a variety of benefits, including medical, dental, vision, life and disability insurance, 401(k), 11 paid holidays, vacation time, and sick time, as well as a comprehensive leave program.

Wage ranges are based on various factors, including the labour market, job type, and job level. Exact salary offers will be determined by factors such as the candidate’s subject knowledge, skill level, qualifications, experience, and geographic location.

All roles are eligible to participate in the Fortinet equity program. Bonus eligibility is reviewed at the time of hire and annually at the Company’s discretion.

Why Join Us:

We encourage candidates from all backgrounds and identities to apply. We offer a supportive work environment and a competitive Total Rewards package to support you with your overall health and financial well-being.

About the company

Embark on a challenging, enjoyable, and rewarding career journey with Fortinet. Join us in bringing solutions that make a meaningful and lasting impact to our 660,000+ customers around the globe., Fortinet (NASDAQ: FTNT) secures the largest enterprise, service provider, and government organizations around the world. Fortinet empowers its customers with intelligent, seamless protection across the expanding attack surface and the power to take on ever-increasing performance requirements of the borderless network - today and into the future. Only the Fortinet Security Fabric architecture can deliver security without compromise to address the most critical security challenges, whether in networked, application, cloud or mobile environments. Fortinet ranks number one in the most security appliances shipped worldwide and more than 500,000 customers trust Fortinet to protect their businesses.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all