> Markdown version of [/jobs/ext/3029833-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3029833-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** CyberPoint International - **Location:** Columbia, MD, United States - **Salary:** $91,944.0 - $137,916.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Apache Accumulo, Amazon Web Services, JIRA, Bash Shell, Cloud Computing, Information Systems, Linux, Distributed Systems, Elasticsearch, Apache Hadoop, Hadoop Distributed File System, Monitoring of Systems, Python (Programming Language), Linux System Administration, Open Source Technology, OpenStack, Reliability Engineering, Ansible, Prometheus, Virtualization Technology, Scripting, Reliability of Systems, Containerization, Kubernetes, Information Technology, Free and Open-Source Software, 3-tier Architectures, Docker - **Published:** September 22, 2026 - **Apply:** https://www.jofdav.com/jobs/59820799-site-reliability-engineer ## About the Role What You'll Need * U.S. Citizenship. * Active TS/SCI Security Clearance with current polygraph. * 14 years of relevant professional experience. * Bachelor's degree in Computer Science or a related technical field is highly desired and may be considered equivalent to 2 years of experience. * A Master's degree in a technical field may be considered equivalent to 4 years of experience. * Degrees in Mathematics, Information Systems, Engineering, or similar technical disciplines will be considered related technical fields. * Strong experience troubleshooting and supporting Linux-based environments. * Experience supporting operational environments and troubleshooting production systems. * Experience working with cloud-based or distributed computing environments. * Ability to provide Tier 1 through Tier 3 technical support. * Ability to participate in an on-call support rotation. * Must possess one of the following certifications: + AWS Certified Developer - Associate + AWS Certified Solutions Architect - Associate + AWS Certified Solutions Architect - Professional + AWS Certified SysOps Administrator - Associate + Certified Kubernetes Administrator (CKA) + Elastic Certified Engineer + Elastic Certified Observability Engineer * DoD 8570 IAT Level I or higher certification/qualification is required. ## Description As a Site Reliability Engineer (SRE) at CyberPoint, you will support the day-to-day operations and stability of a cloud-based platform built on Java and Free and Open-Source Software technologies, including Kubernetes, Hadoop, and Accumulo. The platform enables the execution of data-intensive analytics across a managed infrastructure supporting mission-critical operations. As a member of the Operations Team, you will provide customer support, troubleshoot complex operational issues, and help maintain the availability, reliability, and performance of the platform. This role requires a strong Linux background and the ability to troubleshoot across a diverse technology environment. The ideal candidate is a self-motivated, proactive problem solver who thrives in a fast-paced team environment, pays close attention to detail, and can independently work through complex technical issues. This is an on-call position requiring the ability to provide Tier 1 through Tier 3 operational suppor When you become a part of CyberPoint, you are joining a dynamic, diverse, fast-growing company that welcomes creative thought and ambition. We're committed to creating an environment where each employee can thrive. What You'll Do * Support the day-to-day operations, stability, availability, and performance of a cloud-based platform. * Provide Tier 1 through Tier 3 operational and technical support. * Monitor production environments and respond to operational issues and customer requests. * Diagnose, troubleshoot, and resolve complex issues within Linux-based environments. * Support cloud infrastructure and applications built using Java and open-source technologies. * Troubleshoot and support Kubernetes, Hadoop, and Accumulo environments. * Support data-intensive analytics applications operating on managed cloud infrastructure. * Investigate system and application failures and perform root cause analysis. * Identify opportunities to improve system reliability, performance, and operational efficiency. * Develop and maintain scripts and automation using Python, Bash, or similar scripting languages. * Support containerized applications and infrastructure using Docker and Kubernetes. * Monitor systems using observability and monitoring technologies. * Collaborate with developers, system administrators, cloud engineers, and other technical teams to resolve operational issues. * Provide technical support and guidance to customers and other team members. * Document troubleshooting procedures, operational processes, technical issues, and resolutions. * Participate in an on-call rotation and respond to operational incidents as required. * Proactively identify potential operational issues and implement solutions before they impact customers. * Work effectively within a fast-paced, mission-focused team environment., * Experience with one or more of the following technologies is highly desired: + Kubernetes + Docker + Apache Hadoop + Hadoop Distributed File System (HDFS) + Apache Accumulo + Python + Bash + Prometheus + Grafana + Elasticsearch and Elastic observability technologies + Salt or Ansible + OpenStack + AWS + Virtualization technologies + Jira + Java + Free and Open-Source Software (FOSS) environments + Cloud infrastructure and platform operations + Data-intensive analytics platforms + Infrastructure monitoring and observability + Automated operational tooling and scripting + Root cause analysis and production incident management ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)