Principal Site Reliability Engineer

U.S. Navy
Vienna, VA, United States
about 1 month ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Computer Programming Fault Tolerance Monitoring of Systems Python (Programming Language) Reliability Engineering Software Engineering Reliability of Systems Information Technology

Job description

Maintain and enhance the reliability, availability, and performance of Navy Federal’s systems and services by implementing best practices and collaborating with development teams to ensure smooth operations. Design and implement reliable, fault-tolerant systems, as well as monitor and automate tasks to ensure system stability and efficiency. Work independently to interpret and develop solutions to complex business challenges that have a significant impact on the function or branch. Specialized skill set and proficiency with procedures and techniques. Recognized as an expert in own area within the company.

Responsibilities

  • Design and implement reliable, scalable, and highly available systems.

  • Analyze and evaluate complex variable factors to develop optimal reliability solutions.

  • Automate tasks and processes to improve efficiency and reduce human error.

  • Ensure systems are operating within established performance and reliability metrics.

  • Monitor, maintain and improve system performance, availability, and reliability.

  • Troubleshoot complex problems related to system reliability and performance.

  • Assist in leading collaboration with cross-functional teams to integrate reliability best practices.

  • Assist with the development and enhancement of practices, procedures, and instructions

  • Develop and maintain strong working relationships with team members, subject matter experts, and leaders; work with senior management on complex issues

  • Lead medium to large projects and initiatives

Requirements

  • Master’s degree in computer science, engineering, or the equivalent combination of education, training or experience.

  • 7-10 years of experience in site reliability engineering

  • Subject matter expert within business area/specialization with understanding of interrelationships of different disciplines

  • Advanced knowledge of system monitoring, incident response, and automation tools.

  • Significant experience with SRE principles and practices.

  • Significant experience with software development or site reliability engineering.

  • Advanced programming skills in languages such as Python, Java, or Go.

  • Advanced knowledge of incident response and managing production issues.

  • Advanced communication and problem-solving skills.

  • Ability to work independently.

About the company

Navy Federal provides much more than a job. We provide a meaningful career experience, including a culture that is energized, engaged and committed; and fierce appreciation for our teams, who are rewarded with highly competitive pay and generous benefits and perks.

Our approach to careers is simple yet powerful: Make our mission your passion.

  • FORTUNE 100 Best Companies to Work For® 2026

  • Yello and WayUp Top 100 Internship Programs 2025

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:22 min

Understanding software engineering as more than just coding

Lilia Gargouri Lilia Gargouri · World Congress 2026 Europe

1:05 min

Practical Byzantine Fault Tolerance in distributed computing systems

Jonan Scheffler · World Congress 2022

3:31 min

Revolutionizing computer programming through natural language code generation

Demetris Cheatham Demetris Cheatham +1 · World Congress 2024

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

3:49 min

Enhancing system resilience and fault tolerance

Michael Eder +1 · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all