Site Reliability Engineer

BC Forward
Buffalo, United States of America
2 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 229K

Job location

Buffalo, United States of America

Tech stack

Application Lifecycle Management
Application Performance Management
Azure
Cloud Computing
Cloud Engineering
Code Coverage
Computer Security
Continuous Integration
DevOps
Disaster Recovery
Distributed Systems
Fault Tolerance
Log Analysis
Performance Tuning
Systems Development Life Cycle
Regression Testing
Release Management
Reliability Engineering
Site Reliability Engineering Practices
Software Deployment
Data Logging
Performance Testing
Cloud Monitoring
Infrastructure Automation Frameworks
Deployment Automation
Cloud Optimization
Terraform
Dynatrace

Job description

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure aligned to enterprise standards and SRE best practices.
  • Lead initiatives to improve reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor SLOs, SLIs, and error budgets for critical services.
  • Develop observability strategies using Dynatrace, OpenTelemetry, distributed tracing, metrics, logs, dashboards, and alerting.
  • Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to identify performance bottlenecks, reliability risks, and capacity constraints proactively.
  • Lead incident response for high-severity events and coordinate cross-functional restoration and communications.
  • Perform and facilitate RCAs with corrective and preventive actions tracked to completion.
  • Automate repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to embed reliability and observability across the SDLC.
  • Design, develop, and execute automated regression testing to validate stability and performance after changes.
  • Review test coverage and reliability validation to ensure comprehensive risk mitigation.
  • Create, maintain, and improve Terraform-based IaC for provisioning, configuration, and standardization.
  • Support and optimize Microsoft Azure environments, including App Services, resource management, scaling, and deployment automation.
  • Use Azure Monitor, Application Insights, and Log Analytics to improve visibility and reliability.
  • Drive performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness.
  • Establish operational readiness standards and enforce requirements before production deployments.
  • Review architectures and recommend improvements for resiliency, efficiency, and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization across production environments.
  • Develop and maintain runbooks, incident playbooks, knowledge articles, and SOPs.
  • Partner with engineering, infrastructure, cybersecurity, architecture, and support teams on cross-functional improvements.
  • Communicate system health, reliability trends, risks, and remediation to technical and business stakeholders.
  • Present initiatives, metrics, and recommendations at reviews, forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational practices.
  • Adhere to risk and regulatory standards and identify issues requiring escalation.
  • Promote a culture of belonging consistent with company values and maintain internal control standards.

Requirements

We are seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex reliability initiatives., * Associate's degree with 7+ years in SRE, Cloud, Systems, Infrastructure Engineering, DevOps, or Application Support. Bachelor's degree with 5+ years. Or 9+ years combined education and experience with 5+ years in a technology engineering role.

  • Hands-on observability and monitoring experience with Dynatrace, OpenTelemetry, distributed tracing, metrics, centralized logging, alerting, and dashboards.
  • Proven ability to design and execute automated regression testing frameworks and suites.
  • Strong proficiency with Terraform and Infrastructure as Code practices.
  • Experience with CI/CD, deployment automation, and operational tooling.
  • Expertise in production monitoring, incident management, and troubleshooting of distributed systems.
  • Understanding of application performance management and modern cloud-native architectures.

Preferred Skills:

  • Microsoft Azure expertise, including App Services, Resource Groups, networking, scaling and optimization, deployment and release management, and application lifecycle management.
  • Use of Azure Monitor, Application Insights, Log Analytics, dashboards, and alerting.
  • SRE practices such as SLOs, SLIs, error budgets, incident/problem management, RCA, and reliability automation.
  • Experience with performance tuning, capacity planning, proactive issue detection, and observability-driven improvements.
  • Automated recovery mechanisms and self-healing solutions, resiliency patterns, DR planning, and high-availability architectures.

Benefits & conditions

  • Competitive compensation and benefits.
  • Opportunities for growth with global clients.
  • A supportive, inclusive culture that values innovation and people.
  • Exposure to modern technologies and projects.

About the company

BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity.

Apply for this position