Site Reliability Engineer

QGENDA, LLC
Atlanta, GA, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Data Analysis Systems Engineering Configuration Management Code Review Computer Engineering Continuous Integration DevOps
+34 more
Distributed Revision Control Domain Name System (DNS) Fault Tolerance Identity and Access Management Key Management Log Analysis Network Segmentation Scrum Methodology Systems Development Life Cycle Reliability Engineering Cloud Services Amazon Simple Notification Service (SNS) Datadog Scripting Software Modules GitHub Copilot Amazon ElastiCache System Availability Reliability of Systems AWS ECS Git Amazon Relational Database Service Kubernetes Infrastructure Automation Frameworks Information Technology Teamcity Cloudwatch Amazon Simple Queue Service (SQS) Terraform Docker Pagerduty Jenkins Amazon Redshift Programming Languages

Job description

As a Senior Site Reliability Engineer, you will work with our Infrastructure and Product Development Teams to increase the scalability, reliability, and performance of our systems and services. You will build and extend existing automation for configuration and monitoring of our AWS hosted applications. You will have the opportunity to evaluate new AWS services and tools to determine if they could be utilized in our environments. You’ll bring a focus to platform health and monitoring to allow us to deliver the best possible experience for our customers. This is an excellent opportunity to have a significant impact on the stability of our systems and contribute to the evolution of our technology stack., As a Senior Site Reliability Engineer, you will work with our Infrastructure and Product Development Teams to increase the scalability, reliability, and performance of our systems and services. You will build and extend existing automation for configuration and monitoring of our AWS hosted applications. You will have the opportunity to evaluate new AWS services and tools to determine if they could be utilized in our environments. You’ll bring a focus to platform health and monitoring to allow us to deliver the best possible experience for our customers. This is an excellent opportunity to have a significant impact on the stability of our systems and contribute to the evolution of our technology stack. Responsibilities:

System Reliability & Performance:

  • Design, implement, and manage scalable systems that ensure high availability, fault tolerance, and optimal performance.
  • Continuously monitor and enhance system health and performance through data analysis and metrics.

Automation & Tooling:

  • Develop and advocate for automation tools to eliminate repetitive manual processes and improve efficiency.
  • Build and enhance CI/CD pipelines to streamline software delivery and deployments.

Incident Management & Troubleshooting:

  • Participate in on-call rotation to respond to incidents, troubleshoot problems, and minimize downtime.
  • Conduct root cause analyses and implement permanent solutions to recurring issues.

Infrastructure Management:

  • Manage our cloud-based infrastructure environment in AWS.
  • Optimize costs and resources while maintaining robust and scalable systems.

Collaboration & Culture:

  • Serve as a technical advisor to engineering teams on infrastructure and operations best practices.
  • Actively contribute to fostering an SRE culture within the organization by promoting observability, retrospectives, and continuous improvement.

Requirements

  • Curiosity-driven mindset with a desire to continuously learn and improve systems
  • Strong sense of ownership - you see problems through to resolution, not just escalation
  • Comfortable navigating ambiguity and making pragmatic tradeoffs under pressure
  • Availability for off-hours deployment and upgrades of production systems during release and maintenance windows
  • Strong problem-solving skills and ability to work effectively under pressure.
  • Excellent communication skills for cross-functional collaboration as well as documentation creation.

Experience You Bring

  • B.S. in Computer Science, Computer Information Systems, or Computer Engineering from a major U.S. university or equivalent industry experience
  • 7+ years of experience as a DevOps, SRE or Systems Engineer
  • Advanced proficiency with at least one scripting or programming language
  • Experience with Docker and container orchestration tools such as AWS ECS and EKS/Kubernetes
  • Hands-on experience building infrastructure and supporting applications in AWS using services such as Lambda, EC2, ECS, S3, SNS, SQS, RDS, Redshift, and Elasticache
  • Strong understanding of networking and DNS
  • Strong experience with Terraform for infrastructure provisioning and module development, along with configuration management and infrastructure as code (IaC) practices
  • Firm understanding and experience with Agile and Scrum SDLC processes
  • Using distributed version control system experience (Git preferred) to check-in code, branching, merging, pull request, code review, etc
  • Knowledge of CI/CD best practices and tools such as AWS CodeBuild, Jenkins and/or TeamCity
  • Experience using AI-assisted coding tools (e.g., Claude, GitHub Copilot) to accelerate IaC development, scripting, and operational workflows
  • Familiarity with AI/ML-driven approaches to observability, anomaly detection, log analysis, or incident triage
  • Experience designing and delivering secure, high performance and highly available cloud services
  • Experience with observability platforms (e.g., Datadog, CloudWatch, PagerDuty) for monitoring, alerting, and incident response
  • Awareness of cloud security best practices including IAM policies, network segmentation, and secrets management

Benefits & conditions

We offer a comprehensive total rewards package to support our full-time employees and their family’s day-to-day needs, well-being and major life events, which includes:

  • Fully company-paid options for medical (both in-person and virtual), dental and vision insurance
  • Generous paid time off (PTO) policy to enjoy periods of uninterrupted rest and relaxation for a healthy work/life balance
  • Paid parental leave for birth, adoption or permanent placement
  • 401(k) with company match
  • Options to work in a hybrid-working model or remotely from home, depending on the position
  • Annual Costco membership, cell phone stipend, commuter benefits, in-office perks and more

QGenda delivers technology solutions to improve how healthcare is delivered and increase access - for everyone. We can only succeed by bringing together diverse minds, thoughts, ideas and team members to create better solutions for our customers and make us a better company as a whole. We are committed to creating a culture of embracing diversity, inclusion and equity for all.

About the company

QGenda is redefining healthcare workforce management everywhere care is delivered. We’re on a mission to empower the healthcare industry to better onboarding, deploy, and manage their workforce. Over 4,500 healthcare organizations have trusted us to help them make strategic workforce decisions through our unified software platform. With more than 800 employees across the US, we are united in our vision and culture to make a difference for our customers, while enjoying the day-to-day.

At QGenda, we value our employees and their contributions toward the success of the business. We strive to create a dynamic work environment that fosters growth, innovation, and collaboration, where employees can be proud of the work they do and the impact it has on the healthcare industry.

QGenda is headquartered in Atlanta.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all