Site Reliability Engineer (Hybrid)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
A full-time, hybrid role reporting to our Technical Support Manager, our Site Reliability Engineer will work frequently from our headquarters in Cincinnati, Ohio. Residency in the state of Ohio is preferred.
As Site Reliability Engineer, you will own the reliability, scalability, and performance of our production systems hosted on Azure Kubernetes Service (AKS), working at the intersection of software engineering and operations to build the tooling, processes, and culture that keep our services running at scale. You will be a key contributor to our observability practice - using Splunk for log analytics and alerting, Instana for APM and distributed tracing, and native Azure tools including Azure Monitor, Log Analytics, and Application Insights to provide a comprehensive, real-time view of system health.
This position will be accountable for:
Reliability & Incident Management
- Defining, tracking, and reporting on SLIs, SLOs, and error budgets for all critical services.
- Designing and maintaining runbooks, escalation paths, and on-call rotation schedules.
- Designing chaos engineering practices to proactively surface reliability weaknesses before they impact customers.
Observability - Splunk, Instana & Azure
- Building and maintaining Splunk searches, dashboards, and alert policies covering application and infrastructure logs.
- Developing KPIs and unified service-health views for engineering and leadership.
- Configuring and extending instrumentation across microservices for distributed tracing and real-time baselining.
- Creating smart alerts integrated with on call and ticketing systems for automated incident routing.
- Maintaining Azure Monitor alert rules, action groups, and workbooks across Azure subscriptions.
- Utilizing Log Analytics workspaces including data-retention policies and ingestion cost governance using KQL.
- Leveraging Application Insights for APM, availability testing.
- Driving convergence of Splunk, Instana, and Azure signals into a unified observability strategy.
- Building internal tooling in Python, Go, or Bash to eliminate toil and accelerate incident response.
- Participating in security reviews and threat-modeling sessions for new platform capabilities.
The Person
The opportunity to cast a vision for success, to blend art and science with proactive strategy and tactical execution, will be a key to success in this role.
- Youâre a self-starter. In many ways, youâll help define success metrics for this position. You intuitively understand what will be required to excel at this work and wonât expect others to produce the blueprint for you. Youâll hit the ground running as the architect for this function, helping to ideate and inform where, and how, with what, and by whom site reliability will be achieved.
- You enjoy wearing multiple hats. At this stage in Unlimitedâs growth, our thought leaders have an opportunity to speak into and influence our technology, operations, and processes. Having a green field to work in will mean as you build your role, youâll touch on othersâ work and play in othersâ courts. If we agree on successful outcomes, how we get there will be a collective strategy and effort.
- Youâre comfortable with shifting priorities, and some might say youâre an expert at context shifting. Not everyone can sustain the exercise of planning strategically but building reactively. Not every day will present the requirement to shift gears, but some days will present the need to be tactically agile. When that moment comes, youâll not only be prepared, but youâll be also ready., Applicants may be subject to a background check. Employees in this position must be able to satisfactorily perform the essential functions of the position. If requested, Unlimited Systems will make every effort to provide reasonable accommodation to enable employees with disabilities to perform the positionâs essential job duties. As markets change and the Organization grows, job descriptions may change over time as requirements and employee skill levels evolve. With this understanding, Unlimited Systems retains the right to change or assign other duties to this Senior Infrastructure Administrator position.
Requirements
Required
- 5+ years in an SRE, DevOps, or Platform Engineering role in a production cloud environment.
- Hands-on AKS experience: cluster provisioning, upgrades, CNI networking, and workload lifecycle management.
- Proficiency with Splunk: SPL authoring, dashboard creation, alert configuration, and data onboarding.
- Working knowledge of Instana APM: agent deployment, custom tracing, alerting, and performance analysis.
- Solid command of Azure Monitor, Log Analytics (KQL), and Application Insights.
- Infrastructure-as-Code experience with Terraform; familiarity with Helm and Kustomize.
- Scripting or development proficiency in at least one of: Python, Go, or Bash.
- Demonstrated ability to drive incident resolution and lead blameless post-mortems.
- Clear communicator - able to translate complex technical topics for non-technical stakeholders.
Preferred
- Microsoft Certified: Azure Administrator Associate (AZ-104) or Azure DevOps Engineer Expert (AZ-400).
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
- Splunk Certified Power User or Splunk Enterprise Certified Architect.
- Experience with service mesh technologies (Istio, Linkerd) on AKS.
- Familiarity with FinOps practices and Azure cost-management tooling.
- Experience in a regulated environment (SOC 2, ISO 27001, PCI-DSS, or HIPAA).
About the company
Unlimited Systems authors the category-leading Unlimited Financials practice management system focused on the unique revenue cycle requirements of specialty healthcare providers. Unlimited Systems customers enjoy streamlined business office workflows, reduced claim denial rates, and accelerated and amplified revenue streams. Unlimited Systems is committed to ensuring that specialty healthcare providers thrive in a dynamic reimbursement environment.
Unlimited Systems is a portfolio company of Francisco Partners, a leading technology investment firm with deep sector focus and a track record of delivering outstanding returns. Through private equity and credit funds, they provide flexible capital and partnership to growth-aspiring technology companies.
The Challenge
Unlimited Financials operates across a sophisticated cloud-native environment - Azure Kubernetes Service, event-sourced microservices, healthcare integrations, and a growing network of downstream data pipelines - all of which requires dedicated technical support to remain stable, observable, and performant for our customers. We are building the reliability practice this platform deserves. As we scale into new specialty healthcare markets and take on greater complexity, we need someone who thinks proactively - not just responding to incidents but engineering the systems and signals that prevent them. This is not a run-and-maintain role. It is a âbuild-and-shape-the-futureâ role., If youâve made it this far in the job description, weâd love to hear from you. We will try to ensure our selection process is efficient, transparent, and friendly. (If itâs not, please let us know.)
What you can expect from here:
- If we feel your experience aligns with our search, weâll reach out for a discovery call. Weâll talk about our company culture, Unlimited Systemâs Core Values, this position, and weâll learn more about your background.
- From there, weâll ask to view some of your previous work as well as request that you complete an assessment. This lets us know whether our expectations for the role actuallyalign with your unique strengths.
- If weâre still feeling good about each other, weâll pass you on to our Hiring Manager and then to a panel interview so you can meet more of the faces of Unlimited Systems before we make an offer decision.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs