Site Reliability Engineer

TechSpace Solutions Inc.
Cincinnati, OH, United States
about 2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Amazon Web Services Cloud Computing Databases Data Systems Software Debugging Distributed Systems Payment Systems Performance Tuning Ruby on Rails Reliability Engineering
+10 more
Runbook Software Engineering SQL Databases Datadog Grafana Information Technology Splunk New Relic (SaaS) Service Stack Microservices

Job description

  • Client is looking for an enterprise-grade embedded finance platform enabling organizations to build, launch, and scale compliant banking, payments, and lending solutions.
  • We are seeking a Principal Software Engineer to join our Production Engineering team. This is a hands-on technical leadership role focused on operating, debugging, and improving highly distributed, mission-critical payment systems. The ideal candidate thrives in complex production environments and enjoys solving deep technical challenges across applications, infrastructure, and data systems., * Lead production triage and incident response across APIs, payment systems, distributed services, infrastructure, and databases.
  • Diagnose and resolve complex production issues spanning code, infrastructure, data, and third-party dependencies.
  • Partner with engineering teams to implement permanent fixes and improve platform reliability.
  • Design and implement monitoring, alerting, automation, and operational tooling.
  • Improve system observability, resiliency, and debuggability.
  • Work across a mixed technology stack including Ruby on Rails, Java, AWS, APIs, and SQL databases.
  • Develop runbooks and diagnostic workflows for operational excellence.
  • Mentor engineers and influence best practices across engineering and SRE teams.
  • Participate in architectural discussions to build highly reliable and scalable systems.

Requirements

  • 10+ years of experience in Software Engineering, Production Engineering, SRE, or Distributed Systems.
  • Strong experience debugging production issues end-to-end (application, infrastructure, data, and dependencies).

Hands-on experience with:

  • AWS and cloud-native environments
  • Ruby on Rails and/or Java
  • APIs, Microservices, and Distributed Systems
  • SQL and database troubleshooting
  • Observability tools such as Splunk, Datadog, New Relic, etc.

Deep understanding of:

  • System behavior in production
  • Fault isolation and troubleshooting
  • Performance optimization and resiliency patterns
  • Excellent communication and stakeholder management skills.
  • Ability to work effectively during incidents and high-pressure situations.

Preferred Qualifications:

  • Experience in Payments, FinTech, Banking, or other regulated environments.
  • Experience building and operating large-scale, high-availability platforms.
  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all