> Markdown version of [/jobs/ext/1857254-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1857254-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** TENEX SECURITY, INC. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Audit Trail, Microsoft Azure, Cloud Computing, Cyber Security, Software Debugging, DevOps, Distributed Data Store, Distributed Systems, Monitoring of Systems, Reliability Engineering, Prometheus, Security Information and Event Management, Datadog, Data Logging, Pulumi, Cloud Platform System, Large Language Models, Grafana, Reliability of Systems, Infrastructure as Code (IaC), Event Driven Architecture, Kubernetes, Information Technology, Machine Learning Operations, Terraform, Microservices - **Published:** July 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c3b20247cc21c16f ## About the Role * Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale. * Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage. * Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments. * Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions. * Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems. Soft Skills * Communication: Clear, concise communication skills and a bias for collaborative problem-solving. * Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams. * Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments., * Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure. * AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization). * Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS. * Strong familiarity with Agentic Workflows such as Agno, Temporal, etc.. Education & Certifications * Bachelor's or Master's degree in Computer Science, Engineering, or a related field. * Relevant certifications (CKA/CKAD, AWS/GCP Professional Cloud Architect, etc.) are a plus. ## Description * System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform. * Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning. * Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads. * Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues. * Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP). * Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)