DevOps Engineer

Careflow LLC
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

Application Performance Management Computing Platforms Software Bug Management Cloud Computing Cloud Computing Security Cloud Engineering Databases Continuous Integration Data Infrastructure Software Debugging DevOps Disaster Recovery
+20 more
Monitoring of Systems Python (Programming Language) Key Management Linux System Administration Node.Js Performance Tuning Reliability Engineering Cloud Services TypeScript Software Vulnerability Management Data Logging Google Cloud Delivery Pipeline Usage Tracking Containerization Infrastructure Automation Frameworks Performance Monitor Terraform Data Pipelines Docker

Job description

We are looking for an experienced DevOps Engineer to own and improve our cloud infrastructure, security, observability, and operational reliability. This role is responsible for ensuring our platform remains secure, scalable, performant, and highly available as we continue to grow.

The ideal candidate is someone who enjoys wearing multiple hats-building infrastructure, improving deployment processes, monitoring production systems, and troubleshooting issues across the stack. As a bonus, we would love someone who is comfortable diving into the application codebase to diagnose and resolve bugs when needed.

This is a fully remote position. We are particularly interested in candidates who can provide weekend coverage on Saturdays and take another day off during the week in exchange.

What You’ll Do

Cloud Infrastructure & Operations

  • Manage and maintain our Google Cloud Platform (GCP) environment.
  • Design, implement, and improve infrastructure for scalability, reliability, and cost efficiency.
  • Manage networking, compute resources, databases, storage, and cloud services.
  • Monitor system health and proactively address performance bottlenecks.

Monitoring, Logging & Observability

  • Build and maintain centralized logging and monitoring solutions.
  • Create dashboards and alerts for system health, application performance, and business-critical workflows.
  • Establish operational metrics and usage tracking across the platform.
  • Lead incident response and root cause analysis efforts.
  • Monitor and manage spend

Security & Compliance

  • Implement and maintain security best practices across infrastructure and applications.
  • Manage identity and access controls, secrets management, and environment security.
  • Conduct security reviews and vulnerability remediation.
  • Assist with compliance initiatives and audit readiness.

CI/CD & Automation

  • Improve deployment pipelines and release processes.
  • Automate infrastructure provisioning and operational workflows.
  • Enhance development environments and deployment reliability.
  • Reduce manual operational tasks through automation.

Reliability Engineering

  • Improve uptime, resiliency, backup strategies, and disaster recovery processes.
  • Establish service-level objectives and operational standards.
  • Drive improvements in platform stability and performance.

Cross-Functional Support

  • Partner with engineering, product, and leadership teams to support company initiatives.
  • Provide technical guidance on infrastructure and operational considerations.
  • Participate in an on-call and operational support rotation.

Bonus Responsibilities

  • Troubleshoot and fix application-level issues when needed.
  • Contribute code improvements and bug fixes across the platform.
  • Assist with performance optimization and debugging efforts.

What Success Looks Like

Within your first 90 days, you will:

  • Gain ownership of our GCP infrastructure and environments.
  • Establish visibility into system performance, reliability, and usage metrics.
  • Improve monitoring, alerting, and incident response processes.
  • Identify and address security and operational risks.
  • Reduce infrastructure-related issues and deployment friction.
  • Become a trusted technical resource for platform reliability and operational excellence.

Requirements

Do you have experience in System performance monitoring?, This role is ideal for someone who enjoys both infrastructure ownership and hands-on problem solving, and wants to have a significant impact on the reliability, security, and scalability of a growing software platform., * 5+ years of DevOps, Site Reliability Engineering, Cloud Engineering, or related experience.

  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Experience building and maintaining CI/CD pipelines.
  • Strong understanding of infrastructure monitoring, logging, and alerting systems.
  • Experience with cloud security best practices.
  • Experience managing production environments and incident response.
  • Strong Linux administration skills.
  • Experience with Infrastructure as Code tools (Terraform preferred).
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Strong troubleshooting and problem-solving abilities.
  • Excellent written and verbal communication skills.
  • Ability to work independently in a fully remote environment.

Nice-to-Have Qualifications

  • Experience working in startup or high-growth environments.
  • Experience with healthcare technology or regulated environments.
  • Ability to read and contribute to application code.
  • Experience with Python, TypeScript, Node.js, or similar technologies.
  • Experience building internal tooling and automation.
  • Experience with data pipelines and analytics infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:30 min

Transitioning from software engineering into a DevOps trajectory

Davide Imola Davide Imola · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all