Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+43 more
Job description
Test, maintain (patching, STIGing, and upgrading), troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager (MIM), Active Directory Federation Services (ADFS), DHCP, DNS, WINS, GPOs & PKI. Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform. In addition, as you discover and document system bugs, you have the motivation to go off and fix them yourself. Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines. Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools Identify performance bottlenecks and optimize the performance of cloud infrastructure Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code. Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies. Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads. Create, script, and run performance tests to measure system behavior under varying levels of load and traffic. Identify bottlenecks, performance degradation, and areas for optimization. Design, implement, and maintain automated test suites for infrastructure and application components. Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release. Build automated systems for continuous performance testing, stress testing, and load testing. Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals. Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production. Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions. Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks. Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management. Must be willing to work nights, weekends, and provide on-call support as needed.
Requirements
Bachelor’s Degree IAT Level 2 Certification Must be able to support program execution in classified environments and access SIPRNet. Familiarity designing, configuring, and managing Active Directory, Group Policy Objects, DNS, DHCP, WINS. Experience with Windows Server 2016 / 2019 / 2022 / 2025 Experience with SQL Server 2019 / 2022 / 2025 Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred Experience in CI/CD toolsets (e.g. Jenkins, GitLab, etc.) Good command of Linux/Unix and command line knowledge Experience in application administration, configuration, and integration Familiarity with agile development methodologies Skilled and disciplined to work with a distributed team Ability to work in a highly collaborative, forward thinking, and innovation-driven environment Knowledge of Agile and DevSecOps/SRE concepts and best practices, with a desire to grow that knowledge Hand-on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.). Experience creating JIRA and/or Azure DevOps workflows, projects, custom configurations Experience administrating/maintaining SRE platform via Ansible playbooks (e.g. upgrading Jenkins) Experience in automating tasks with scripting languages like PowerShell, or Python Integrating/maintaining with various 3rd party CI/CD tools like Jenkins and Gitlab. Experience with PaaS using Red Hat OpenShift/Kubernetes and Docker containers Experience with commercial cloud infrastructure deployment environments such as AWS and Azure. Experience with automated provisioning and configuration tools like Terraform, Cloud Formation, Chef, Puppet, Ansible, or similar technologies. Working knowledge of the Risk Management Framework (RMF), DISA STIGs, Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation for automating test environments. ITILv4, Scrum Master, or Agile SAFe certification(s) or applicable experience Familiarity with designing, configuring and managing FIM 2010 R2/MIM 2016 synchronization Familiarity with VB.NET Familiarity designing, configuring, and managing Delinea Familiarity with Microsoft Active Directory Federated Services structure Familiarity with Cloud Engineering Familiarity with Public Key Infrastructure (PKI) certificates
About the company
Norfolk
Virginia
28391
Prism, Inc., PRISM is devoted to modernization and innovation within the world of technology, security, and IT enterprise solutions. We are recognized for meeting performance requirements and exceeding customer expectations since 1994. Our culture is founded on relationships, opportunity, and success. Offering comprehensive benefit plans including medical, dental, vision, and 401K along with our people - first approach sustains our reputation as a premier employer.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.
Why Upskilling And Reskilling is Important For Developers
Dev Digest 134 - Where pixels sing?
Résumé-Driven Development: How IT trends affect the job market for software developers