> Markdown version of [/jobs/ext/655585-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/655585-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Realty Professionals, LLC - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Amazon Cloudfront, Amazon Elastic Compute Cloud, Amazon S3, Bash Shell, Cloud Computing, Computer Programming, Continuous Integration, Decision Support Systems, DevOps, Disaster Recovery, Distributed Systems, Fault Tolerance, Github, Identity and Access Management, Python (Programming Language), Octopus Deploy, Reliability Engineering, Newrelic, Prometheus, Datadog, Circleci, Data Logging, Istio, System Availability, Grafana, Mttr, Reliability of Systems, Amazon Virtual Private Cloud (VPC), Cloudformation, Amazon Relational Database Service, Kubernetes, Infrastructure Automation Frameworks, AWS Fargate, Graphql, Route53, BIG-IP Access Policy Manager (APM), Functional Programming, Cloudwatch, Api Gateway, Terraform, Splunk, Dynatrace, Docker, Pagerduty, Jenkins, Servicenow - **Published:** June 26, 2026 - **Apply:** https://www.dice.com/job-detail/463b1d99-43e8-428f-9123-5619e793af41 ## About the Role * 5+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with demonstrated success improving system reliability * Bachelor's degree or equivalent experience * 3+ years hands-on experience with AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes including cluster management * Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, CloudFormation) * Production experience with observability tools (NewRelic, Datadog, Prometheus, Grafana, Splunk) and distributed systems * Experience with CI/CD platforms and GitOps workflows (CircleCI, Argo CD, Jenkins); on-call rotation and incident response * Preferred: Exposure to chaos engineering tools, API Gateway technologies (Tyk/Kong), GraphQL federation (Apollo), cost optimization initiatives, FinOps principles Technical Skills * Cloud & Infrastructure: AWS (EKS, Fargate, Lambda, VPC, Route53, CloudFront), Kubernetes, Docker, Istio Service Mesh * CI/CD & GitOps: Argo CD, CircleCI, Jenkins, GitHub Actions * Observability: NewRelic - APM, distributed tracing, metrics & logging; Splunk - logging * IaC & Automation: Terraform, CloudFormation, Helm, Kustomize, Python/Go/Bash * Platform Services: Tyk Gateway, Apollo GraphQL, AWS Secrets Manager, Vault * Incident Management: OpsGenie, PagerDuty, ServiceNow Professional Qualities * Strong communication skills with ability to explain technical concepts to diverse audiences * Collaborative approach working across engineering, product, and business teams * Self-motivated with ability to solve complex problems within established practices and policies * Data-driven decision making with customer-centric approach and empathy for developer experience ## Description We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization, reporting to the Director, Operations Excellence. This role will contribute to the reliability, observability, and operational excellence of our platform infrastructure serving millions of users. As a Senior SRE, you will be a strong technical contributor who implements best practices, solves complex problems, and enables our 600+ engineers to deliver exceptional customer experiences. You will work on critical platform systems including EKS infrastructure, Skyway (CI/CD), Frontdoor (Tyk API Gateway), Pantheon (Apollo GraphQL Federation), and our observability stack, while contributing to chaos engineering practices and cost optimization initiatives with measurable ROI. What You'll Do: Platform Reliability & Infrastructure * Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures * Support reliability of critical services: Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure * Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency * Implement reliability patterns including circuit breakers, graceful degradation, and automated failover Observability & Cost Optimization * Implement observability solutions using NewRelic for APM, distributed tracing, metrics, and logging for rapid troubleshooting * Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams * Identify infrastructure cost optimization opportunities and implement FinOps practices including rightsizing and resource lifecycle management * Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD) Chaos Engineering & Incident Response * Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create runbooks and automation for resilience * Participate in on-call rotation for critical systems; conduct post-incident reviews and implement improvements * Support incident response processes and contribute to System Health Scorecard Technical Contribution * Contribute as a strong technical individual contributor to the Operations Excellence team * Collaborate with Platform Engineering, Quality Engineering, and product teams on reliability initiatives * Support security initiatives including AWS Secrets Manager migration and compliance requirements (SOC 2, PCI, GDPR) * Contribute to Developer Experience metrics and platform adoption goals * May provide technical guidance to junior team members ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023)