Site Reliability Engineer (Deployment)

Clientsolv, Inc
Austin, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Databases Data as a Services Software Debugging Disaster Recovery Domain Name System (DNS) Identity and Access Management Virtual Private Networks (VPN) Python (Programming Language)
+22 more
Key Management PostgreSQL Networking Basics OAuth OpenID Role-Based Access Control Reliability Engineering Runbook Single Sign-On Software Deployment Data Logging Scripting Transport Layer Security Google Cloud Load Balancing Performance Testing Istio Large Language Models Kubernetes Infrastructure Automation Frameworks Terraform Dynatrace

Job description

We are hiring Senior Site Reliability Engineers to support enterprise platform deployments for a fast-growing AI-focused technology company. This is a hands-on role working with customer infrastructure teams to deploy, secure, automate, and optimize Kubernetes-based platforms., Deploy and manage applications on Kubernetes environments (AWS, Azure, Google Cloud Platform, or on-prem) Automate infrastructure using Terraform and GitOps practices Integrate identity management, networking, security, and data services Implement observability solutions including monitoring, logging, and alerting Support production deployments, performance validation, and operational readiness Create deployment documentation, runbooks, and operational procedures

Requirements

Strong Kubernetes administration and troubleshooting experience Terraform, Helm, Kustomize, and GitOps tools Cloud platforms (AWS, Azure, or Google Cloud Platform) Networking fundamentals including DNS, load balancing, TLS, and VPN/private connectivity Identity and security integration (OIDC, OAuth, SSO, RBAC, secrets management) Monitoring, logging, distributed tracing, and incident response Experience with PostgreSQL or similar databases Scripting/automation skills and ability to read/debug Go or Python code Strong customer-facing communication skills

Preferred: Experience in regulated environments (Healthcare, Financial Services, etc.) Service Mesh, Disaster Recovery, or Chaos Engineering experience Exposure to AI/LLM-based platforms

Candidates should be able to work onsite in Austin at least 3 days per week. Local candidates are preferred, but candidates willing to relocate or commute from nearby cities will be considered.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:49 min

Adopting OAuth best practices and removing outdated grants

Alexander Schwartz Alexander Schwartz · WWC Europe 2026

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

4:35 min

Setting up passwordless federated identity configuring OpenID Connect patterns

Marcel Lupo · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:22 min

Adapting OpenID Connect for decentralized data sharing

Adam Larter Adam Larter · WWC 2024

1:34 min

Analyzing vulnerabilities in standard OAuth 2.0 authorization flows

Alexander Schwartz Alexander Schwartz · WWC Europe 2026

Videos

See all

Related articles

See all