Site Reliability Engineers

Socure Inc.
Carson City, NV, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Application Performance Management Cloud Computing Continuous Integration Github Identity and Access Management Intrusion Detection and Prevention Python (Programming Language) Datadog Kubernetes Production Code Terraform

Job description

You will work at the intersection of cloud infrastructure, Kubernetes, automation, and observability, with a strong focus on preventing incidents rather than reacting to them. What You’ll Own

  • End-to-end ownership of highly available, scalable AWS infrastructure
  • Design, operation, and continuous improvement of Kubernetes (EKS) platforms
  • Reliability of production systems through strong observability, automation, and SLOs
  • CI/CD systems that enable safe, fast, and repeatable deployments
  • Infrastructure defined and enforced through Terraform and GitOps
  • Incident response, root cause analysis, and long-term remediation
  • Raising operational standards through automation, documentation, and best practices

Requirements

We’re looking for engineers who have actually built, run, and scaled real production systems in the following areas: Cloud & Infrastructure

  • Deep AWS expertise - networking, compute, IAM, scaling, security
  • Strong experience managing infrastructure using Terraform at scale

Kubernetes & Platform Engineering

  • Very strong Kubernetes fundamentals (internals, scheduling, networking, storage)
  • Hands-on experience operating Amazon EKS in production environments
  • Experience troubleshooting complex, multi-layer Kubernetes issues

Coding & Automation

  • Ability to write clean, maintainable, production-quality code in: Go/ Python
  • Strong automation mindset - eliminating toil through code

CI/CD & GitOps

  • Proven experience building and operating CI/CD pipelines
  • Hands-on experience with:
  • GitHub (Actions or integrations)
  • ArgoCD and GitOps-based deployment workflows

Observability & Reliability

  • Strong understanding of observability principles: metrics, logs, traces, and alerting
  • Hands-on experience with Datadog or similar tool for:
  • Infrastructure and Kubernetes monitoring
  • Application performance monitoring (APM)
  • Alerting, dashboards, and incident detection
  • Experience defining and using SLIs/SLOs to drive reliability decisions

Ability to turn observability data into actionable operational improvements

About the company

Socure is building the identity trust infrastructure for the digital economy - verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.

We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself - keep reading.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:36 min

Critical infrastructure and performance skills for modern developers

Andrew Holway · LIVE

Videos

See all

Related articles

See all