Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
As a Site Reliability Engineer (SRE) at CyberPoint, you will support the day-to-day operations and stability of a cloud-based platform built on Java and Free and Open-Source Software technologies, including Kubernetes, Hadoop, and Accumulo. The platform enables the execution of data-intensive analytics across a managed infrastructure supporting mission-critical operations.
As a member of the Operations Team, you will provide customer support, troubleshoot complex operational issues, and help maintain the availability, reliability, and performance of the platform. This role requires a strong Linux background and the ability to troubleshoot across a diverse technology environment.
The ideal candidate is a self-motivated, proactive problem solver who thrives in a fast-paced team environment, pays close attention to detail, and can independently work through complex technical issues. This is an on-call position requiring the ability to provide Tier 1 through Tier 3 operational suppor
When you become a part of CyberPoint, you are joining a dynamic, diverse, fast-growing company that welcomes creative thought and ambition. We’re committed to creating an environment where each employee can thrive.
What You’ll Do
- Support the day-to-day operations, stability, availability, and performance of a cloud-based platform.
- Provide Tier 1 through Tier 3 operational and technical support.
- Monitor production environments and respond to operational issues and customer requests.
- Diagnose, troubleshoot, and resolve complex issues within Linux-based environments.
- Support cloud infrastructure and applications built using Java and open-source technologies.
- Troubleshoot and support Kubernetes, Hadoop, and Accumulo environments.
- Support data-intensive analytics applications operating on managed cloud infrastructure.
- Investigate system and application failures and perform root cause analysis.
- Identify opportunities to improve system reliability, performance, and operational efficiency.
- Develop and maintain scripts and automation using Python, Bash, or similar scripting languages.
- Support containerized applications and infrastructure using Docker and Kubernetes.
- Monitor systems using observability and monitoring technologies.
- Collaborate with developers, system administrators, cloud engineers, and other technical teams to resolve operational issues.
- Provide technical support and guidance to customers and other team members.
- Document troubleshooting procedures, operational processes, technical issues, and resolutions.
- Participate in an on-call rotation and respond to operational incidents as required.
- Proactively identify potential operational issues and implement solutions before they impact customers.
- Work effectively within a fast-paced, mission-focused team environment., * Experience with one or more of the following technologies is highly desired:
- Kubernetes
- Docker
- Apache Hadoop
- Hadoop Distributed File System (HDFS)
- Apache Accumulo
- Python
- Bash
- Prometheus
- Grafana
- Elasticsearch and Elastic observability technologies
- Salt or Ansible
- OpenStack
- AWS
- Virtualization technologies
- Jira
- Java
- Free and Open-Source Software (FOSS) environments
- Cloud infrastructure and platform operations
- Data-intensive analytics platforms
- Infrastructure monitoring and observability
- Automated operational tooling and scripting
- Root cause analysis and production incident management
Requirements
What You’ll Need
- U.S. Citizenship.
- Active TS/SCI Security Clearance with current polygraph.
- 14 years of relevant professional experience.
- Bachelor’s degree in Computer Science or a related technical field is highly desired and may be considered equivalent to 2 years of experience.
- A Master’s degree in a technical field may be considered equivalent to 4 years of experience.
- Degrees in Mathematics, Information Systems, Engineering, or similar technical disciplines will be considered related technical fields.
- Strong experience troubleshooting and supporting Linux-based environments.
- Experience supporting operational environments and troubleshooting production systems.
- Experience working with cloud-based or distributed computing environments.
- Ability to provide Tier 1 through Tier 3 technical support.
- Ability to participate in an on-call support rotation.
- Must possess one of the following certifications:
- AWS Certified Developer - Associate
- AWS Certified Solutions Architect - Associate
- AWS Certified Solutions Architect - Professional
- AWS Certified SysOps Administrator - Associate
- Certified Kubernetes Administrator (CKA)
- Elastic Certified Engineer
- Elastic Certified Observability Engineer
- DoD 8570 IAT Level I or higher certification/qualification is required.
About the company
Great people are the foundation of any great company. To attract and retain great people, CyberPoint offers fulfilling work that we like to balance with the rest of life. We accomplish this by offering outstanding benefits that allow each of our employees to live and work well. To learn more, https://www.cyberpointllc.com/benefits.php.
Be a part of CyberPoint. Be valued.
We are an inclusive community in which all types of diversity are embraced, respected and seen as true value for the company. We are committed to cultivating this community, but we realize there is still more work to do. CyberPoint will continue to evolve into a more inclusive and equitable company and is fully committed to the principles of equal employment opportunity and affirmative action.
CyberPoint is an Equal Opportunity Employer and does not discriminate against any employee or applicant for employment because of gender, gender identity or expression, sexual orientation, race, age, religion, physical or mental disability, veterans’ status or other federal, state and local protected characteristics.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Best Paying Jobs in Technology
Highest Paying Tech Companies for Developers
Top-Paying Tech Jobs (with Salaries)