Site Reliability Engineer

NCR VOYIX CORPORATION
Atlanta, GA, United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Applicant Tracking Systems Applications Architecture Computing Platforms Systems Engineering Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Continuous Integration Linux
+28 more
DevOps Github Monitoring of Systems Python (Programming Language) Windows PowerShell Reliability Engineering Prometheus Software Engineering Datadog Pulumi Scripting Google Cloud Grafana Software Troubleshooting Reliability of Systems Infrastructure as Code (IaC) Cloudformation Containerization Gitlab-ci Infrastructure Automation Frameworks Information Technology Deployment Automation Terraform Splunk Docker Jenkins Golang Programming Languages

Job description

At NCR Voyix, we’re looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our products and services. This role is ideal for an engineer who enjoys solving technical challenges, automating manual processes, and improving the reliability and performance of applications in a cloud-native environment.

As part of our Engineering organization, you’ll work closely with software engineers, cloud architects, and operations teams to ensure our platforms remain secure, available, and scalable. You’ll have the opportunity to gain hands-on experience with cloud technologies, automation, observability tools, and modern infrastructure practices while contributing to mission-critical systems used by customers around the world.

What You’ll Do

  • Support and enhance cloud-based infrastructure and applications across modern cloud platforms.
  • Implement automation solutions that improve operational efficiency, reduce manual effort, and increase system reliability.
  • Monitor application and infrastructure health using observability and monitoring tools to identify and resolve performance issues.
  • Participate in incident response and on-call support rotations for production systems.
  • Troubleshoot infrastructure, application, and deployment issues in partnership with engineering teams.
  • Assist with root cause analysis (RCA) activities and support implementation of preventive solutions.
  • Contribute to Infrastructure as Code (IaC) initiatives using modern automation and cloud provisioning tools.
  • Partner with software development teams to improve application reliability, scalability, and operational readiness.
  • Support CI/CD pipelines and deployment automation processes.
  • Create and maintain technical documentation, operational runbooks, and troubleshooting guides.
  • Stay current with emerging cloud and platform technologies and recommend improvements where applicable., * Helps reduce operational toil through automation and process improvements.
  • Responds effectively to incidents and contributes to long-term reliability improvements.
  • Builds strong partnerships with development and platform engineering teams.
  • Continuously expands technical expertise while supporting business-critical systems.

This job description outlines the primary responsibilities of the position and is not intended to be an exhaustive list of duties. Additional responsibilities may be assigned based on business needs and individual skills and experience.

Offers of employment are conditional upon passage of screening criteria applicable to the job

EEO Statement

Integrated into our shared values is NCR Voyix’s commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law. NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party Agencies To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 1-3 years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, Systems Engineering, Infrastructure Engineering, or a related technical role.
  • Experience with at least one public cloud platform such as AWS, Azure, or Google Cloud Platform (GCP).
  • Experience with scripting or programming languages such as Python, Bash, Go, or PowerShell.
  • Familiarity with Linux operating systems and basic system administration concepts.
  • Understanding of cloud infrastructure, networking, and application architecture fundamentals.
  • Familiarity with container technologies such as Docker and Kubernetes.
  • Knowledge of CI/CD concepts and experience with tools such as GitHub Actions, GitLab CI, Jenkins, or similar platforms.
  • Exposure to monitoring and observability tools such as Datadog, Grafana, Prometheus, Splunk, or ELK.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Excellent communication and collaboration skills with the ability to work effectively across technical teams.

Preferred Qualifications

  • Exposure to Infrastructure as Code tools such as Terraform, CloudFormation, or Pulumi.
  • Understanding of SRE concepts including service availability, reliability, incident management, and operational excellence.
  • Experience supporting production applications in a cloud-native environment.
  • Familiarity with Agile software development methodologies.
  • Relevant cloud certifications (AWS, Azure, or GCP) are a plus.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all