Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
our client is building core infrastructure for community banks, which means uptime, data integrity, security, compliance, and incident response aren’t nice-to-haves - they’re the product. Today we run on cloud infrastructure with basic monitoring and alerting, but we have not established a formal SRE practice. We’re hiring a Senior Site Reliability Engineer - in practice, this is an IC-heavy SRE role, not a people-management role. You’ll spend most of your time building: designing DR procedures, hardening infrastructure, and writing the automation and runbooks yourself. Part of the role is about technical leadership and multiplying your knowledge into the team, not headcount management or ceremony. You’ll work alongside our two DevOps engineers, pairing with them, reviewing their work, and turning what you know into practices the whole team can run without you in the room - but you’re still the one with your hands on the keyboard for most of the hard problems. If you’ve been the SRE who got called when a mission-critical system went down, who’s designed DR plans that actually got tested, who’s built on-call cultures that don’t burn people out
- this role gives you a real green-field problem: a company that knows it needs this and hasn’t
had someone to build it yet. What you’ll own
-
Hands-on SRE work, with knowledge that spreads
-
Personally execute the majority (roughly 60-70%) of the platform buildout - you’re the
senior-most engineer on infrastructure, and it shows in the code, configs, and systems you ship, not just the docs you write.
- Delegate the remaining work deliberately to our two DevOps engineers, structured to
grow their skills rather than just clear your queue.
- Turn what’s in your head into what’s in the team’s hands: runbooks, architecture decision
records, pairing sessions, and reviews - so practices survive without you being the single point of failure.
-
Reliability & disaster recovery
-
Design, document, and run our first real Disaster Recovery drill - then make DR testing
a recurring practice, not a one-time event.
- Map the gap between “what we have” and a “solid, redundant, highly available platform,”
and turn it into a prioritized execution plan.
- Build toward self-healing infrastructure: auto-scaling, automated failover, graceful
degradation, clear data backup strategy, rollbacks, and reduced dependency on manual intervention.
-
Incident response & on-call
-
Review our existing monitoring and alerting tools and procedures in order to improve and
establish proper paging (PagerDuty or equivalent), escalation policies, and runbooks to support the response team on an on-call rotation schedule.
- Ensure that when something breaks, the right person gets paged immediately - and
has what they need (logs, dashboards, runbooks, access) to remediate fast.
- Own incident command during major outages; drive blameless postmortems and
follow-through on action items.
-
Platform architecture
-
Be the technical authority on infrastructure architecture: cloud topology, redundancy,
networking, deployment pipelines, observability stack.
- Evaluate and evolve our current cloud setup toward higher availability
(multi-AZ/multi-region where it matters, proper backups, tested restore procedures).
- Balance reliability work against product velocity - you know what’s non-negotiable (data
integrity, security, compliance posture) versus what can be deferred.
-
Information security & compliance posture (SOC 2, PCI)
-
We’re already SOC 2 and PCI compliant - this role exists to make sure that stays true
continuously, not just during audit season.
- Guide and support the team on the infrastructure-side controls that sustain both
certifications: access management, encryption in transit/at rest, audit logging, change management, vulnerability management, and evidence collection.
- Work with whoever owns compliance/audit relationships (internal or external) to translate
control requirements into concrete infrastructure and process changes - and make sure those changes actually stick between audits, not just before them.
- Treat security posture as a continuous improvement target, not a checkbox: proactively
identify where reliability work and compliance requirements overlap (e.g., audit trails, immutable logs, incident response documentation) and use one to strengthen the other.
Requirements
integrity, audit trails, and uptime carry different weight when the product touches customer money.
- Compliance-aware infrastructure experience. You’ve worked in environments holding
SOC 2 and/or PCI certification (or equivalent) and know what it actually takes to keep controls sustained day-to-day, not just pass the annual audit.
- Pragmatic, not dogmatic. You can tell the difference between “this must be bulletproof”
and “this can ship with known tradeoffs,” and you can explain why.
- Comfortable with ambiguity. You can take “we want a reliable, self-healing platform”
and turn it into a concrete, sequenced plan - without waiting to be told exactly what to build. Technical background we’d expect
- Deep experience with cloud infrastructure (AWS preferred, given our current stack) -
EC2, networking, load balancing, managed databases, backup/restore.
- Strong background in observability: monitoring, alerting, logging, tracing (e.g., Datadog,
CloudWatch, Prometheus/Grafana, ELK, or similar).
- Experience with Infrastructure-as-Code (Terraform, CloudFormation, or similar) and
CI/CD pipelines.
- On-call tooling experience (PagerDuty, Opsgenie, or similar) and incident management
frameworks.
- Strong experience with containerization/orchestration (Docker, Kubernetes, or ECS).
- Working knowledge of SOC 2 and PCI DSS control frameworks and how they map to
infrastructure (access controls, logging/monitoring, encryption, change management, vendor management, incident management).
- Experience with CI/CD pipelines and GitHub Actions, including build, test, and
deployment automation.
- Additional experience considered a plus:
- Experience with endpoint and device management (MDM), including security
policies and compliance.
- Experience administering Google Workspace, including identity, access, and
security controls.
- Experience managing penetration testing findings and vulnerability remediation
across engineering teams. How we’ll evaluate fit, Education: Associate’s in IT
Benefits & conditions
you’ve actually built, an incident you personally led through resolution, a moment you had to decide what was non-negotiable versus deferrable under time pressure, and how you’ve split ownership with engineers more junior than you without either micromanaging or disappearing. Why this role, specifically Most SRE roles at scale-ups mean inheriting someone else’s half-finished system and untangling it for two years before you get to build anything new. This one’s different: there’s no legacy SRE debt to pay off, no prior “SRE team” whose turf you’re navigating. You’ll define what reliability means here, on infrastructure that’s still young enough to shape well - with a CTO and founding team actively looking for someone senior to take this on with full ownership. Logistics
- Location: Remote, open to candidates based in Latin America.
- Team structure: You’ll work alongside and technically lead two DevOps engineers,
reporting directly to the CTO. This is not a people-management role - no performance reviews, hiring plans, or admin overhead to pull you away from the work.
- Compensation: Competitive, commensurate with experience.
About the company
equivalent senior infra engineer) at mission-critical companies for about 5 - 10 years - fintech, payments, healthcare, infra/dev-tools, or anything where downtime meant real damage, not just annoyance.
- Hands-on is non-negotiable. You still want to write IaC, debug production, and read
logs directly - not sit one abstraction layer above the work. This role stays IC-heavy by design.
- You’ve built the practices, not just followed them. On-call rotations you designed, DR
plans you wrote and actually tested, postmortem cultures you started from nothing. Bonus points if you’ve turned chaos into a system that mostly runs itself.
- You know how to multiply yourself. You’ve raised the bar for engineers less senior
than you - through pairing, reviews, and clear documentation - without needing to manage them formally.
- Fintech or regulated-industry instinct is a strong plus. You understand why data
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Why Attend a Developer Event in 2026?
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship
What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?
Highest Paying Tech Companies for Developers