> Markdown version of [/jobs/ext/3021882-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3021882-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** STORD, Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer-Aided Design, Artificial Intelligence, Cloud Computing, Code Review, Databases, Software Debugging, DevOps, Disaster Recovery, Distributed Systems, Github, Identity and Access Management, Python (Programming Language), PostgreSQL, Machine Learning, Redis, Prometheus, TypeScript, Datadog, Grafana, Multi-Cloud, Git, Event Driven Architecture, Infrastructure Automation Frameworks, Cloudflare, Apache Kafka, Vertica, Terraform, Docker, Golang, Programming Languages - **Published:** September 21, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2963806083-2 ## About the Role * 5+ years in SRE, platform, or infrastructure engineering: you've owned complex systems and driven technical work forward with minimal supervision. * GCP depth: strong hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM). You know how these fit together in production, not just in a certification. * Containers & orchestration: you're fluent in Docker and Kubernetes and can debug, tune, and scale real workloads. * Infrastructure as Code: deep Terraform experience. You write reusable modules, reason about state, and treat infra changes with the same rigor as application code. * A programming language you're genuinely productive in: TypeScript, Python, Go, or similar, used to build tooling and automation, not just glue scripts., * Ownership & Accountability: You own features end-to-end and take pride in what you ship. You follow through from design to production and don't drop things. * Strong Communication: You can explain technical decisions and trade-offs to engineers, PMs, and stakeholders. You ask good questions and listen well. * Collaborative Approach: You work well with others, give constructive code review feedback, and actively seek input from teammates. * Production Mindset: You prioritize reliability and user impact. You think about failure modes, monitoring, and operational concerns as part of your design process. * Learning Agility: You're comfortable with rapidly evolving AI/ML technologies and tools. You stay current without chasing hype. * Directed AI-Assisted Development: You know how to use AI coding tools as a productivity multiplier while maintaining quality and your own technical judgment. Strongly Preferred * Database operations depth: PostgreSQL internals (logical replication, vacuuming, lock contention) or experience with migrations and database scaling. Familiarity with Redis, ClickHouse, or analytical stores is a plus. * Event-driven systems: Kafka/Redpanda or Pub/Sub, schema registries, and the operational realities of streaming at scale. * Cost engineering: you've meaningfully reduced cloud or observability spend without sacrificing reliability., * GCP certifications (Cloud Architect, Cloud DevOps Engineer) or demonstrably equivalent depth. * Cloudflare: experience with Workers and other Cloudflare services. * Multi-cloud or hybrid architecture exposure. What Success Looks Like * 30 days: You've ramped on our GCP environment and Terraform setup, shipped your first infrastructure or automation change to production, and are contributing in incident and design discussions. * 90 days: You independently own a meaningful slice of the infrastructure, have improved a reliability, cost, or automation pain point that was slowing the team down, and teammates lean on you in your area. * 6 months: You're a trusted owner of your part of the infrastructure, consistently delivering improvements that make the platform more reliable and efficient. ## Description * Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking. * Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on. * Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization. * Drive down cost and toil through better defaults, right-sizing, and automation rather than manual intervention. Reliability & Observability * Build monitoring, alerting, and observability in Datadog (APM, logs, RUM) that catches problems before customers do. * Define the reliability signals that matter for the services you own, and hold the line on them. * Develop and maintain disaster recovery and business-continuity strategies, and prove they work. Automation & Delivery * Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety. * Automate operational workflows and infrastructure provisioning so the platform scales smoothly as the team grows. * Build custom tooling and scripts that remove recurring operational pain. Collaboration & Incident Response * Partner with data and development teams to improve deployment practices and application reliability. * Provide escalation support for production incidents, help lead post-incident reviews, and turn findings into durable fixes. * Participate in technical design reviews and offer architectural input across teams. * Help improve SRE and infrastructure best practices across the team, and participate in on-call for critical systems., * Observability: you build monitoring and alerting that's actionable (Datadog, or equivalents like Prometheus/Grafana), and you know the difference between a noisy dashboard and a useful one. * Distributed systems fundamentals: failure modes, consistency, and how systems break at scale. * Git and collaborative development workflows: you work in shared codebases and review others' changes well. * Incident management: you've run incidents and post-mortems and can stay calm and methodical when production is on fire. ## Related Videos - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding)