Devops - SRE Architect

Purple Drive Technologies LLC
Austin, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$94,000.0 - $142,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Amazon Web Services Application Performance Management Microsoft Azure C Sharp (Programming Language) Unix Cloud Computing Cloud Foundry Databases Continuous Integration Couchbase Servers
+46 more
Data Infrastructure Database Schema DevOps Disaster Recovery Distributed Systems Elasticsearch Failover Identity and Access Management JSON Python (Programming Language) PostgreSQL Linux System Administration Enterprise Messaging Systems Oracle (Applications) Performance Tuning Redis Reliability Engineering Prometheus Distributed Caching Standard Sql Shell Script Apache Solr PL-SQL SQL Databases Data Logging Scripting Transport Layer Security Java Application Server Load Balancing Cloud Platform System System Availability Grafana Multi-Cloud Git Flow Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Cassandra Performance Monitor Azure AKS Database Replication Restful APIs Amazon Simple Queue Service (SQS) Splunk Docker

Job description

Infrastructure & Platform Engineering

  • Design, implement, and maintain enterprise Kubernetes (AKS) platforms.
  • Build and optimize containerized applications using Docker and Kubernetes.
  • Implement GitOps deployment strategies using ArgoCD and Helm.
  • Design secure cloud infrastructure following Infrastructure-as-Code (IaC) best practices.
  • Manage ingress controllers, load balancers, SSL/TLS, and mTLS certificate lifecycle.
  • Architect highly available cloud-native platforms across Azure and multi-cloud environments.

DevOps & CI/CD

  • Design and maintain CI/CD pipelines.
  • Automate deployments using GitOps methodologies.
  • Develop automation using Shell and Python scripting.
  • Manage database schema migrations using Flyway.
  • Improve deployment reliability and release automation.

Site Reliability Engineering (SRE)

  • Define and monitor SLAs, SLOs, SLIs, and error budgets.
  • Ensure production availability, scalability, and performance.
  • Lead P0/P1 incident response, troubleshooting, and Root Cause Analysis (RCA).
  • Manage on-call rotations and operational excellence.
  • Implement proactive monitoring and alerting strategies.

Observability

  • Build enterprise observability platforms using:

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Splunk
  • Design logging, tracing, metrics collection, and monitoring frameworks.

Application Infrastructure

  • Troubleshoot Java applications in distributed environments.
  • Support REST APIs, gRPC services, HTTP/JSON integrations.
  • Manage messaging systems and distributed caching.
  • Optimize cloud-native application performance.

Database & Data Platform

  • Design and maintain high-availability databases including:

  • Oracle
  • PostgreSQL
  • Cassandra
  • Couchbase
  • Redis
  • CockroachDB
  • Manage Oracle GoldenGate replication.
  • Implement disaster recovery, failover, replication, backup, and recovery strategies.
  • Perform SQL/PL-SQL development and performance tuning.

Security & Governance

  • Implement Infrastructure Security best practices.
  • Manage IAM roles and privilege management.
  • Automate password rotation.
  • Collaborate with Security, Networking, and Application teams.
  • Ensure compliance with enterprise security standards.

Required Technical Skills DevOps & Platform Engineering

  • Kubernetes (AKS)
  • Docker
  • Helm
  • ArgoCD
  • GitOps
  • CI/CD
  • Flyway

Cloud Platforms

  • Microsoft Azure
  • Azure Kubernetes Service (AKS)
  • Multi-cloud (AWS/GCP/AliCloud preferred)

Scripting

  • Shell Scripting
  • Python
  • Linux Administration
  • Unix

Observability

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Splunk

Databases

  • Oracle
  • PostgreSQL
  • Cassandra
  • Couchbase
  • Redis
  • CockroachDB

Database Technologies

  • SQL
  • PL/SQL
  • Oracle GoldenGate
  • Database Replication
  • Performance Tuning

Application Technologies

  • Java
  • REST APIs
  • gRPC
  • HTTP
  • JSON
  • Elasticsearch
  • Solr
  • SNS
  • SQS, Role Overview As a .NET Engineer, you will play a crucial role in shaping the future of AI systems by leveraging your expertise in .NET and C#. Your contributions will directly i…
  • 22 days ago

Requirements

We are seeking an experienced Principal DevOps, SRE & Application Infrastructure Architect to design, build, and manage highly available, scalable, and secure cloud-native infrastructure. The ideal candidate will have deep expertise in Kubernetes (AKS), Docker, GitOps, CI/CD, Site Reliability Engineering (SRE), Observability, Cloud Infrastructure, Database Operations, and Application Infrastructure Architecture. The candidate will lead infrastructure modernization, production reliability, DevOps automation, cloud-native platform engineering, disaster recovery, and application support across enterprise-scale environments., * Load Balancers

  • Ingress Controllers
  • SSL
  • TLS
  • mTLS
  • Disaster Recovery
  • High Availability, * Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 8-10+ years of DevOps/SRE/Infrastructure Engineering experience.
  • Experience designing enterprise cloud-native platforms.
  • Strong knowledge of Kubernetes, Docker, and GitOps.
  • Hands-on production support and incident management experience.
  • Strong communication and stakeholder management skills., * Azure Solutions Architect Certification
  • Kubernetes Certification (CKA/CKAD)
  • AWS Solutions Architect Certification
  • Azure DevOps Certification
  • Experience with AliCloud
  • Experience in Financial Services or enterprise-scale production environments

Benefits & conditions

  • $94,000-142,000 per year Why AIS? When you join AIS, you’re joining a mission-driven team that’s passionate about making a difference. You’ll work on projects that matter, alongside industry-leading expe…

  • 1 day ago +

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

Videos

See all

Related articles

See all