> Markdown version of [/jobs/ext/3039083-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3039083-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** AUWANA HAWAII LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $84,000.0 - $90,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Software System Penetration Testing, Computing Platforms, Audit Trail, Backup Devices, Cloud Computing, Cyber Security, Databases, Data Integrity, DevOps, Disaster Recovery, Fault Tolerance, Github, Identity and Access Management, Uptime, PCI Data Security Standards, Reliability Engineering, Prometheus, Software Vulnerability Management, Datadog, Data Logging, Load Balancing, Autoscaling, Delivery Pipeline, Grafana, Cloudformation, Containerization, Kubernetes, Deployment Automation, Gsuite, Cloudwatch, Restful APIs, Terraform, Docker, Pagerduty - **Published:** September 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=0aa399fe23f2d0ae ## About the Role integrity, audit trails, and uptime carry different weight when the product touches customer money. * Compliance-aware infrastructure experience. You've worked in environments holding SOC 2 and/or PCI certification (or equivalent) and know what it actually takes to keep controls sustained day-to-day, not just pass the annual audit. * Pragmatic, not dogmatic. You can tell the difference between "this must be bulletproof" and "this can ship with known tradeoffs," and you can explain why. * Comfortable with ambiguity. You can take "we want a reliable, self-healing platform" and turn it into a concrete, sequenced plan - without waiting to be told exactly what to build. Technical background we'd expect * Deep experience with cloud infrastructure (AWS preferred, given our current stack) - EC2, networking, load balancing, managed databases, backup/restore. * Strong background in observability: monitoring, alerting, logging, tracing (e.g., Datadog, CloudWatch, Prometheus/Grafana, ELK, or similar). * Experience with Infrastructure-as-Code (Terraform, CloudFormation, or similar) and CI/CD pipelines. * On-call tooling experience (PagerDuty, Opsgenie, or similar) and incident management frameworks. * Strong experience with containerization/orchestration (Docker, Kubernetes, or ECS). * Working knowledge of SOC 2 and PCI DSS control frameworks and how they map to infrastructure (access controls, logging/monitoring, encryption, change management, vendor management, incident management). * Experience with CI/CD pipelines and GitHub Actions, including build, test, and deployment automation. * Additional experience considered a plus: * Experience with endpoint and device management (MDM), including security policies and compliance. * Experience administering Google Workspace, including identity, access, and security controls. * Experience managing penetration testing findings and vulnerability remediation across engineering teams. How we'll evaluate fit, Education: Associate's in IT ## Description our client is building core infrastructure for community banks, which means uptime, data integrity, security, compliance, and incident response aren't nice-to-haves - they're the product. Today we run on cloud infrastructure with basic monitoring and alerting, but we have not established a formal SRE practice. We're hiring a Senior Site Reliability Engineer - in practice, this is an IC-heavy SRE role, not a people-management role. You'll spend most of your time building: designing DR procedures, hardening infrastructure, and writing the automation and runbooks yourself. Part of the role is about technical leadership and multiplying your knowledge into the team, not headcount management or ceremony. You'll work alongside our two DevOps engineers, pairing with them, reviewing their work, and turning what you know into practices the whole team can run without you in the room - but you're still the one with your hands on the keyboard for most of the hard problems. If you've been the SRE who got called when a mission-critical system went down, who's designed DR plans that actually got tested, who's built on-call cultures that don't burn people out * this role gives you a real green-field problem: a company that knows it needs this and hasn't had someone to build it yet. What you'll own * Hands-on SRE work, with knowledge that spreads * Personally execute the majority (roughly 60-70%) of the platform buildout - you're the senior-most engineer on infrastructure, and it shows in the code, configs, and systems you ship, not just the docs you write. * Delegate the remaining work deliberately to our two DevOps engineers, structured to grow their skills rather than just clear your queue. * Turn what's in your head into what's in the team's hands: runbooks, architecture decision records, pairing sessions, and reviews - so practices survive without you being the single point of failure. * Reliability & disaster recovery * Design, document, and run our first real Disaster Recovery drill - then make DR testing a recurring practice, not a one-time event. * Map the gap between "what we have" and a "solid, redundant, highly available platform," and turn it into a prioritized execution plan. * Build toward self-healing infrastructure: auto-scaling, automated failover, graceful degradation, clear data backup strategy, rollbacks, and reduced dependency on manual intervention. * Incident response & on-call * Review our existing monitoring and alerting tools and procedures in order to improve and establish proper paging (PagerDuty or equivalent), escalation policies, and runbooks to support the response team on an on-call rotation schedule. * Ensure that when something breaks, the right person gets paged immediately - and has what they need (logs, dashboards, runbooks, access) to remediate fast. * Own incident command during major outages; drive blameless postmortems and follow-through on action items. * Platform architecture * Be the technical authority on infrastructure architecture: cloud topology, redundancy, networking, deployment pipelines, observability stack. * Evaluate and evolve our current cloud setup toward higher availability (multi-AZ/multi-region where it matters, proper backups, tested restore procedures). * Balance reliability work against product velocity - you know what's non-negotiable (data integrity, security, compliance posture) versus what can be deferred. * Information security & compliance posture (SOC 2, PCI) * We're already SOC 2 and PCI compliant - this role exists to make sure that stays true continuously, not just during audit season. * Guide and support the team on the infrastructure-side controls that sustain both certifications: access management, encryption in transit/at rest, audit logging, change management, vulnerability management, and evidence collection. * Work with whoever owns compliance/audit relationships (internal or external) to translate control requirements into concrete infrastructure and process changes - and make sure those changes actually stick between audits, not just before them. * Treat security posture as a continuous improvement target, not a checkbox: proactively identify where reliability work and compliance requirements overlap (e.g., audit trails, immutable logs, incident response documentation) and use one to strengthen the other. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [You can’t hack what you can’t see](https://www.wearedevelopers.com/videos/41-you-can-t-hack-what-you-can-t-see) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)