> Markdown version of [/jobs/ext/1241297-site-reliability-engineer-top-secret](https://www.wearedevelopers.com/jobs/ext/1241297-site-reliability-engineer-top-secret). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Top Secret - **Company:** KATHY K'S BOOKKEEPING SERVICE LLC - **Location:** Falls Church, VA, United States - **Experience:** Experienced - **Salary:** $100,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Access, Java (Programming Language), Amazon Web Services, Amazon S3, Apache HTTP Server, Computer Vision, Build Automation, Microsoft Azure, C++ (Programming Language), Cloud Computing, Linux, Distributed Data Store, Hadoop Distributed File System, Python (Programming Language), PostgreSQL, Natural Language Processing, Network File Systems, Object-Oriented Software Development, Open Source Technology, Scrum Methodology, Reliability Engineering, Cloud Services, Ansible, Ruby, Mesos, Software Deployment, Software Engineering, Ceph (Software), Google Cloud, Apache Yarn, Deep Learning, Kubernetes Helm Charts, Infrastructure as Code (IaC), Cloudformation, Kubernetes, Information Technology, Data Analytics, Terraform, Docker, Jenkins, Microservices - **Published:** July 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=df9d27f702ad8cfc ## About the Role charts, manage services, and ensure resource allocation for optimal cluster performanceContainerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environmentsCloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)Monitoring & Incident Response: Set up monitoring solutions, define alerts, and manage the incident response process for any issues related to Jenkins or Kubernetes clustersAutomate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalentCollaborate Across Teams: Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructureSecurity & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanningActive TOP SECRET clearance or higher is requiredBachelor's degree in Computer Science or related fieldA minimum of two (2) years of experience working with on-premise and off-premise cloud environmentsExperience with AWS and/or AzureHands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, K8s, Terraform, Helm, PostgreSQL, or similar technologiesAbility to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScriptExperience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn)Proactive approach to identifying problems, performance bottlenecks, and areas for improvementAbility to lead and work independently in an Agile/Scrum environmentReal passion for developing team-oriented solutions to complex engineering problemsThrive in an autonomous, empowering and exciting environmentGreat verbal and written communication skills to collaborate multi-functionally and improve scalabilityInterest in committing to a fun, friendly, expansive and intellectually stimulating environmentDesired SkillsHands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud ServicesExperience with deep learning, natural language processing, computer vision, or reinforcement learningConveys highly technical concepts and information in written form to technical and non-technical audiencesThe ability to work on multiple concurrent projects is essential. Strong self-motivation and the ability to work with minimal supervisionMust be a team-oriented, Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability. EEO IS THE LAW. If you are an individual with a disability and would like to request a reasonable accommodation as part of the employment selection process, please contact the Recruiting Department RecruitingTeam@cathexiscorp.com#J-18808-Ljbffr, Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performanceContainerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environmentsCloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)Monitoring & Incident Response: Set up monitoring solutions, define alerts, and manage the incident response process for any issues related to Jenkins or Kubernetes clustersAutomate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalentCollaborate Across Teams: Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructureSecurity & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanningActive TOP SECRET clearance or higher is requiredBachelor's degree in Computer Science or related fieldA minimum of two (2) years of experience working with on-premise and off-premise cloud environmentsExperience with AWS and/or AzureHands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, K8s, Terraform, Helm, PostgreSQL, or similar technologiesAbility to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScriptExperience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn)Proactive approach to identifying problems, performance bottlenecks, and areas for improvementAbility to lead and work independently in an Agile/Scrum environmentReal passion for developing team-oriented solutions to complex engineering problemsThrive in an autonomous, empowering and exciting environmentGreat verbal and written communication skills to collaborate multi-functionally and improve scalabilityInterest in committing to a fun, friendly, expansive and intellectually stimulating environmentDesired SkillsHands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud ServicesExperience with deep learning, natural language processing, computer vision, or reinforcement learningConveys highly technical concepts and information in written form to technical and non-technical audiencesThe ability to work on multiple concurrent projects is essential. Strong self-motivation and the, ability to work with ## Description empower our employees to create innovative and trusted results.We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.ResponsibilitiesThe responsibilities include, but are not limited to:Monitor and Manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on KubernetesKubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm, to the team; and empower our employees to create innovative and trusted results.We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.ResponsibilitiesThe responsibilities include, but are not limited to:Monitor and Manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on KubernetesKubernetes Management: Deploy, monitor, and scale applications on ## Related Videos - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Fireside Chat with Werner Vogels, VP & CTO, Amazon.com & Daniel Gebler, CTO at Picnic](https://www.wearedevelopers.com/videos/1405-fireside-chat-with-werner-vogels-vp-cto-amazon-com-daniel-gebler-cto-at-picnic) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)