> Markdown version of [/jobs/ext/3572518-site-reliability-engineer-edge-security](https://www.wearedevelopers.com/jobs/ext/3572518-site-reliability-engineer-edge-security). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (Edge Security) - **Company:** Cvent - **Location:** United States - **Experience:** Expert - **Salary:** $120,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Cloudfront, Bash Shell, Cloud Computing, Configuration Management, Cyber Security, Continuous Integration, DevOps, Disaster Recovery, Domain Name System (DNS), Github, Information Retrieval, Python (Programming Language), Key Management, Log Analysis, Networking Basics, Systems Development Life Cycle, Reliability Engineering, Runbook, Server Virtualization, Software Deployment, TypeScript, Datadog, AWS Cdk, Data Logging, Scripting, Load Balancing, Retrieval-Augmented Generation, Delivery Pipeline, Large Language Models, Claude Code, Prompt Engineering, Generative AI, Firewalls (Computer Science), Cloudformation, Kubernetes, Deployment Automation, Cloudflare, Api Gateway, Splunk, Docker, Pagerduty, Jenkins - **Published:** October 3, 2026 - **Apply:** https://www.dice.com/job-detail/62aa420d-7e83-498d-b831-18f95314cbab ## About the Role * 5+ years of hands-on experience in Site Reliability Engineering, DevOps or cloud operations, with a demonstrated track record of owning reliability, security, and operational excellence in production environments. * Hands-on experience with AWS WAF, Shield Advanced, or CloudFront in production at real scale: operational ownership rather than proof-of-concept familiarity. You have authored and managed rules, resolved incidents, and understood the blast radius before deploying. * TLS and certificate lifecycle management: certificate renewals, CSR generation, domain routing, and multi-domain remediation for production systems. * Infrastructure as Code with AWS CDK (or CloudFormation). You write the infrastructure, review the diffs, and understand what a change does before it lands. * Splunk query fluency: WAF log analysis is a core on-call skill. You write the searches and build the dashboards rather than only reading them. * Incident management experience on security-adjacent events: able to act as IC, coordinate across product and security stakeholders in real time, write clear incident summaries, and drive RCA. * Change management discipline: ability to communicate changes proactively to stakeholders, document rollout strategies, and manage phased production deployments with rollback plans. * Fluent in at least one scripting language such as Python, TypeScript, or Bash, enough to automate your own triage and tooling rather than only run scripts others wrote. * On-call experience at real interrupt volume. You know what a sustainable rotation feels like and have managed a walk-up queue without losing project weeks. * SRE fundamentals: SLIs, SLOs, error budgets and toil accounting, used as working tools rather than vocabulary. * Excellent communication skills and a track record of driving alignment across multi-disciplinary teams. * Experience with SDLC methodologies and PR-based deployment workflows. AI & Automation Literacy (Must Have): Practical understanding and hands-on exposure to AI fundamentals as applied to SRE and operational workflows: * Prompt Engineering: ability to design effective prompts for LLMs to assist with incident analysis, RCA generation, runbook creation, and on-call triage. * Retrieval-Augmented Generation (RAG): basic understanding of RAG patterns; ability to leverage or contribute to RAG-based internal tools that surface relevant runbooks, past incidents, and knowledge base articles during operational events. * AI-assisted Workflow & Process Automation: experience using or building AI-powered automations in operational contexts, such as automated incident summarization, alert enrichment, WAF log analysis, change risk assessment, or post-mortem drafting using LLM integrations (for example via MCP tools, Claude Code, Slack bots, or custom pipelines). Good to Have Skills: * AWS WAF WCU modeling: understanding of WebACL cost structure, rule complexity, and the WCU thresholds that translate directly to monthly spend. * AWS Shield Advanced across multi-account environments via Firewall Manager: policy management, account onboarding, transitioning endpoints from count to block mode, and Shield event analysis. * Experience with bot mitigation strategies, including AWS Bot Control, token-based traffic classification, and evaluation of third-party vendors such as Datadome. * F5 LTM configuration and management. * Cloudflare, Fastly, or Datadome: operational experience with complementary edge security, CDN, or bot-protection platforms. * CloudFront at scale: distribution management, cache behaviors, custom origins, Origin Access Control, and dependency mapping; experience with API Gateway and ALBs as part of a layered security posture. * LLM-based agents or Claude Code as a daily engineering accelerant: built or operated agents for triage, analysis, or automation in an operational context. * Experience with APM, monitoring and logging tools (Datadog, PagerDuty, Splunk) and with CI/CD tooling such as GitHub Actions or Jenkins. * Disaster recovery planning and execution: multi-region failover, DR runbooks, and RTO/RPO management. * Good understanding of containerization concepts (Docker, ECS, EKS, Kubernetes) and of basic networking concepts (DNS, TLS, HTTP, load balancing). * Familiarity with risk assessment and management concepts and practices. ## Description As a Senior SRE on the SRE Security team, you will be a core platform owner. You will own the infrastructure, carry the on-call rotation, handle the requests that gate deploys and customer onboarding, and ship the work that keeps the production edge secure and reliable at enterprise scale. We are looking for someone with the drive, ownership and ability to take on challenging problems, both technical and process related, in a dynamic, collaborative and highly distributed, multi-disciplinary team environment. You will work closely with product development teams, Information Security, Cloud Infrastructure and other SRE teams. We use SRE principles such as blameless postmortems, error budgets and a focus on automation to ensure we're constantly improving our knowledge and maintaining a good quality of life. Overall, we're passionate about continuous improvement, learning and participating in dynamic day to day work where success is rewarded with recognition and upward mobility. In This Role, You Will: * Own the edge security and routing layer: WAF rule management, CDK deployments, Shield Advanced policy configuration via AWS Firewall Manager (FMS), and CloudFront distribution management across multiple AWS accounts. * Carry the primary on-call rotation for WAF/Shield and the production edge; triage incidents by reading WAF logs and Splunk queries, act as incident commander when needed, write the postmortem and close the loop with stakeholders. * Handle the daily walk-up queue: requests for help troubleshooting blocked requests, spikes in traffic, bot scraping, or just guidance on the best design for a new application. * Manage the full edge protection estate, including WAF rule lifecycle and WCU cost governance, TLS certificate operations, F5 LTM virtual server management, and CloudFront dependency ownership. * Ensure the scalability, performance, and resilience of edge security systems and processes; identify recurring problems and anti-patterns and turn them into guardrails and automation. * Develop build, test and deployment automation for edge infrastructure using AWS CDK and PR-based workflows; champion Cvent standards and best practices. * Apply AI tooling to accelerate incident triage, cost analysis, and knowledge retrieval, and contribute to expanding the team's AI-assisted workflows. * Build the team's knowledge infrastructure: runbooks, WAF rule catalogues, Shield posture guidelines, and platform documentation. * Work with product development teams, Information Security and other SRE teams to ensure a holistic understanding of edge security concerns and their effective and efficient resolution. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)