Lead Site Reliability Engineer

The Federal Reserve Bank
Boston, MA, United States
3 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$146,700.0 - $190,500.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Amazon S3 Cloud Computing Code Review Databases Continuous Integration Data Structures Software Design Patterns DevOps Disaster Recovery
+38 more
Distributed Systems Amazon DynamoDB Monitoring of Systems Identity and Access Management Python (Programming Language) Key Management Open Web Application Security Scrum Methodology Reliability Engineering Runbook Software Engineering Software Vulnerability Management Datadog Large Language Models Grafana Multi-Cloud Reliability of Systems HybridCloud Amazon Virtual Private Cloud (VPC) Event Driven Architecture Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation AWS Fargate Route53 Cloudwatch Api Gateway Terraform Splunk New Relic (SaaS) Dynatrace Serverless Computing Docker Static Application Security Testing Microservices Dynamic Application Security Testing

Job description

We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

  • Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

  • Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

  • Lead incident response, conduct root cause analysis, and implement preventive measures

  • Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

  • Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

  • Automate deployment pipelines, monitoring, and operational workflows

  • Optimize cloud resource utilization and cost management

Engineering & Development

  • Build and maintain internal tools and services to improve operational efficiency

  • Collaborate with development teams to implement reliability best practices

  • Conduct code reviews and provide technical guidance on system design

  • Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

  • Integrate security practices into CI/CD pipelines (SAST/DAST)

  • Implement and maintain security controls across infrastructure and applications

  • Ensure compliance with industry standards and regulatory requirements

  • Conduct security assessments and vulnerability management

Leadership & Collaboration

  • Mentor junior SRE team members and promote SRE culture across the organization

  • Partner with software engineering teams to improve system reliability

  • Drive technical initiatives and contribute to architectural decisions

  • Document processes, runbooks, and technical specifications

Requirements

  • Strong proficiency inJava,Python, andNode.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code

Cloud Infrastructure (AWS):

  • Extensive experience with AWS services including:
  • Compute:Lambda, ECS, EC2, Fargate
  • Storage:S3, EBS, EFS
  • Database:RDS, DynamoDB, Aurora
  • Networking:VPC, Route53, CloudFront, API Gateway
  • Monitoring:CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred

DevOps & CI/CD:

  • Expert-level knowledge ofGitLab(CI/CD pipelines, runners, GitOps)
  • AdvancedTerraformskills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools

Security:

  • Hands-on experience withSAST(Static Application Security Testing) tools
  • Knowledge ofDAST(Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)

Monitoring & Observability:

  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray, * Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications:Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

Benefits & conditions

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Screening:

Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.

Sponsorship: Individuals who need immigration sponsorship now or in the future are not eligible for this position. Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.

Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)

About the company

CompanyFederal Reserve Bank of San Francisco When you join the Federal Reserve-the nation’s central bank-you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.

The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation’s financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Loading talks and stories from around this role…