DevOps / Platform Infrastructure Engineer

COPE Health Solutions
United States
about 1 month ago
Apply on jobs.jobvite.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$128,600.0 - $181,600.0
Working hours
Shift work

Tech stack

Data Analysis Microsoft Azure Backup Devices Code Review Databases Continuous Integration Software Debugging Linux DevOps Disaster Recovery Github Identity and Access Management
+18 more
Subnetting Python (Programming Language) Key Management Network Troubleshooting PostgreSQL Microsoft SQL Server Windows Servers MySQL Network Architecture Peering Performance Tuning Windows PowerShell Role-Based Access Control Runbook Data Logging Scripting Kubernetes Bicep

Job description

COPE Health Solutions is hiring a Senior DevOps / Platform Infrastructure Engineer to help run and improve the platform behind our applications. Our infrastructure lives primarily in Microsoft Azure. We deploy through GitHub Actions, run workloads on Kubernetes, manage multiple databases, and support systems that handle healthcare data. This is a hands-on role. You’ll spend more time in Bicep templates, AKS clusters, GitHub Actions workflows, PowerShell scripts, and architecture diagrams than you will in meetings about them. You’ll also be one of the people developers call when a deployment breaks at 4:30 on a Friday afternoon. Patience, curiosity, and a healthy sense of humor go a long way. Because we work with healthcare data, security, compliance, and operational discipline aren’t side projects. They’re part of the job., * Build and maintain Azure infrastructure using Bicep

  • Operate and improve our AKS environments, including upgrades, scaling, reliability, and cost optimization
  • Own CI/CD pipelines built with GitHub Actions
  • Automate repetitive operational work using PowerShell and Python
  • Implement monitoring, logging, and alerting that helps us discover problems before users do
  • Reduce cloud spend without creating new operational risks
  • Support PostgreSQL, SQL Server, and MySQL environments, including backups, maintenance, performance tuning, and disaster recovery planning
  • Design and document network architecture, including hub-and-spoke VNets, private endpoints, subnet segmentation, peering, routing, and hybrid connectivity
  • Support identity, access management, secrets management, and compliance initiatives
  • Improve developer experience through better tooling, documentation, automation, and sensible platform defaults
  • Participate in incident response, root-cause analysis, and operational reviews *

Documentation is part of the job

We expect documentation to be treated like code.

You’ll be responsible for creating and maintaining:

  • Runbooks
  • Architecture diagrams
  • Onboarding guides
  • Incident postmortems
  • Operational procedures
  • Platform documentation

If someone asks how traffic gets from Point A to Point B, there should be a diagram that answers the question., Platform engineers share a primary/secondary on-call rotation. Daytime issues get picked up by whoever’s around; the rotation exists for evenings, weekends, and holidays. It covers genuine production incidents: a service down, a data-path failure, a security event, not routine tickets or “can you look at this sometime” requests. When you’re paged, you’re paged for something that actually matters.

We expect platform engineers to participate in incident response, troubleshooting, and root-cause analysis. Just as importantly, we expect recurring operational pain to be addressed through automation, monitoring, documentation, or engineering improvements. The goal isn’t to become better at responding to the same alert every week; it’s to make sure that alert stops happening.

We’ve all been on the receiving end of bad on-call rotations. We’d rather invest in reliability than heroics.

What success looks like

Month 1

  • Get access
  • Learn the environment
  • Ship something small
  • Join incident reviews
  • Figure out where the sharp edges are

Month 2

  • Take ownership of one major platform area
  • Contribute meaningful automation or operational improvements
  • Start participating in production support activities

Month 3

  • Identify risks we haven’t seen yet
  • Propose improvements we haven’t thought of
  • Help raise the engineering standard of the platform

Interview process

We’re less interested in trivia than in how you think.

Candidates should expect practical technical discussions, including:

  • Reviewing a Bicep template
  • Debugging a GitHub Actions pipeline
  • Diagnosing an AKS issue
  • Explaining a hub-and-spoke Azure network design
  • Discussing a production incident they’ve personally handled
  • Walking through a security or HIPAA-related scenario

A strong answer doesn’t require perfection. We’re looking for engineers who can reason through problems, communicate clearly, and learn from operational experience.

How we work

You’ll join a team of 13 and work closely with engineering and security teams. Code review is required. Ego is optional. Documentation gets reviewed like code. When incidents happen, we focus on fixing root causes rather than assigning blame.

Requirements

  • 5+ years operating production infrastructure
  • Deep hands-on experience with Microsoft Azure
  • Strong Infrastructure-as-Code experience with Bicep
  • Extensive Networking experience
  • Production Kubernetes experience, preferably AKS
  • Strong GitHub Actions and CI/CD experience
  • Strong PowerShell and Python scripting skills
  • Experience supporting Linux and Windows Server environments
  • Experience supporting PostgreSQL, and SQL Server
  • Experience designing and troubleshooting network architectures
  • Experience with observability platforms such as Azure Monitor, or similar tools
  • Experience implementing RBAC, IAM, secrets management, and security controls
  • Experience working in regulated environments such as HIPAA,HiTrust, and SOC 2
  • Strong written communication and documentation skills

On-call and production support We run systems clients and their members depend on, so someone needs to be reachable when something breaks outside business hours. We try to be humane about how we do that.

Benefits & conditions

As a firm passionate about health care, we’re deeply committed to the health and wellness of our own team members. We offer comprehensive, affordable insurance plans for our team and their families, and a host of other unique benefits, such as a yearly stipend for wellness-related activities, and a paid parental leave program. You can learn more about our benefits offerings here: https://copehealthsolutions.com/careers

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.jobvite.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · World Congress 2026 Europe

Videos

See all

Related articles

See all