> Markdown version of [/jobs/ext/1310177-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1310177-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Globe Engineers Inc - **Location:** Washington, DC, United States (Remote available) - **Experience:** Expert - **Salary:** $128,000.0 - $192,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Configuration Management, Linux, File Systems, Distributed Data Store, Distributed Systems, Python (Programming Language), Linux Kernel, Open Source Technology, OpenStack, Reliability Engineering, Ansible, Prometheus, Ceph (Software), Openstack Neutron, Saltstack, Grafana, Software Troubleshooting, Build Management, Kubernetes, Storage Technologies, Cinder (Software), Puppet - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/390427bc-c28d-4fb3-bca5-579870508c7e ## About the Role * 5+ years operating large-scale Linux infrastructure with significant experience supporting distributed storage systems in production environments. * 2+ years hands-on Ceph administration including OSD, MON, MDS, and RGW operations, CRUSH map management, pool design, placement groups, and performance troubleshooting. * Strong understanding of Linux internals, storage architecture, networking, filesystems, block devices, and performance analysis under production workloads. * Experience developing operational automation and tooling with Python, Shell, and configuration management platforms such as Ansible, SaltStack, Puppet, or Chef. * Proven incident response expertise, including production troubleshooting, root-cause analysis, postmortem creation, and implementation of durable corrective actions. You Might Also Have... * Advanced Ceph expertise across RGW Multisite, CephFS, RBD Mirroring, erasure coding, and large-scale cluster design. * Experience integrating storage platforms with OpenStack (Cinder, Nova, Neutron, Swift) or Kubernetes storage technologies (CSI, StorageClasses, Rook). * Familiarity with modern observability platforms including Prometheus, Mimir, Loki, Grafana, and enterprise monitoring architectures. * Experience planning and executing petabyte-scale storage migrations, capacity expansion programs, and storage hardware lifecycle management. * Contributions to open-source infrastructure or storage projects and a passion for advancing distributed systems technologies. We encourage you to apply even if your experience or skillset doesn't align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth. ## Description As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do... * Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage workloads (RADOS, RGW, RBD, CephFS). * Diagnose and resolve complex distributed systems issues including recovery/backfill events, OSD instability, PG imbalance, storage latency, and RGW performance bottlenecks. * Design and build automation using Python, Shell, SaltStack, and Ansible to reduce operational toil and improve platform resilience at scale. * Define and improve observability for the storage platform through SLIs, SLOs, PromQL, LogQL, Grafana dashboards, and proactive alerting strategies. * Lead storage lifecycle initiatives including cluster expansions, hardware refreshes, software upgrades, OpenStack integrations, and large-scale migration projects. ## Related Videos - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)