> Markdown version of [/videos/100237-pet-arena-season-2](https://www.wearedevelopers.com/videos/100237-pet-arena-season-2). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # PET ARENA Season 2. Static benchmarks leave privacy systems vulnerable. Now, TikTok and Oblivious have launched a live, peer-to-peer battleground where defenders and attackers co-evolve to battle-test differential privacy architectures. - **Speakers:** [Sagar Sharma .](https://www.wearedevelopers.com/@sagar-sharma) - **Event:** World Congress 2026 Europe - **Published:** July 10, 2026 - **Duration:** 20:56 - **URL:** https://www.wearedevelopers.com/videos/100237-pet-arena-season-2 ## Summary Relying on static benchmarks to test privacy systems often leads to shipping software with blind spots, as real adversaries adapt and chain queries until data leaks. To combat this false confidence, TikTok's Privacy Innovation Lab and Oblivious launched PET ARENA Season 2, a peer-to-peer capture the flag (CTF) competition. This environment allows data defenders and red-team auditors to mutually co-evolve under realistic constraints. Participants test custom privacy-preserving architectures against active adversarial probing, creating an empirical testbed for differential privacy. The competition features two distinct tracks. Defenders design custom noise mechanisms, access controls, and data suppression thresholds, which must pass hidden utility checks to ensure the data output remains usable. Attackers simultaneously leverage a browser-based Pyodide sandbox to execute membership inference, attribute inference, and record linkage exploits to bypass those defenses. A centralized challenge orchestrator manages automated zero-sum matchups, updating both factions in multi-stage cycles using a chess-style Elo rating system with dynamic K-factors. Insights gleaned from Season 1 serve as the foundational curriculum for the current arena. Competitors must systematically balance their privacy budgets and avoid classic technical traps, such as mistaking an artificial noise floor for a reliable data signal or relying on hardcoded census multipliers. Ultimately, this P2P structure forces continuous live adaptation, pushing engineers, data science enthusiasts, and researchers to design robust solutions capable of withstanding highly targeted logical differencing attacks. **Keywords:** capture the flag competition, differential privacy mechanisms, peer-to-peer adversarial testing, red teaming strategies, membership inference attacks, attribute inference attacks, record linkage exploits, elo rating algorithm, query proxy architecture, pyodide browser sandbox, logical differencing attacks, privacy budget constraints, statistical exploitation, data suppression thresholds, dynamic k-factor scoring ## Chapters 1. **Replacing static privacy benchmarks with adaptive adversarial testing** (00:26) — Peer-to-peer competitions reveal how privacy systems fail against evolving query probes. 1. **Fostering a community of privacy builders and breakers** (01:39) — Data engineers and researchers compete to identify empirical privacy leakages using basic Python and SQL. 1. **Analyzing statistical and logical differencing attack strategies** (03:04) — Prior attackers successfully exploited systems via subgroup inference and reverse-engineered noise mechanisms. 1. **Avoiding common privacy querying and budget management pitfalls** (05:04) — Failing to account for noise floors or exhausting privacy budgets leads to refused queries and lost points. 1. **System architecture orchestrating automated privacy match evaluations** (06:10) — A centralized evaluation hub manages datasets, API gateways, and query proxies to prevent overfitting. 1. **Implementing defense track noise mechanisms and query guardrails** (07:44) — Defenders must balance query suppression and custom noise mechanisms without failing hidden utility thresholds. 1. **Designing inference and record linkage attacks for red teaming** (08:27) — Attackers attempt to extract sensitive attributes and bypass dynamic defenses to reidentify database targets. 1. **Progressing through peer-to-peer cycles and final calibration rounds** (09:31) — Participants complete initial house missions before facing opponent submissions across three competitive cycles. 1. **Calculating leaderboard scores with a dynamic rating system** (10:40) — Continuous match scores scale zero-sum algorithms and adapt K factors based on track participation density. 1. **Preventing match collusion and competing for dual-track rewards** (11:47) — System guardrails involuntarily pair opponents and enforce query limits to ensure equitable grandmaster evaluations. 1. **Building and testing strategies in the browser sandbox environment** (13:24) — A user journey demonstrating how to configure LLM classifiers and execute budgeted attacker notebooks within the interface. 1. **Co-evolving submissions and managing late competition registration limits** (18:13) — Continuous strategy modifications reward participants who acclimate to evolving opponent models early in the cycle. ## Related Moments - [Exploring inference, extraction, and the adversarial threat landscape](https://www.wearedevelopers.com/videos/627-machine-learning-promising-but-perilous) (from "Machine Learning: Promising, but Perilous") - [Navigating data privacy boundaries and adversarial model reliability](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) (from "Getting Started with Machine Learning") - [Embedding data security and applied ethics into developer education](https://www.wearedevelopers.com/videos/1107-the-future-of-developer-experience-with-genai-driving-engineering-excellence) (from "The Future of Developer Experience with GenAI: Driving Engineering Excellence") - [Overcoming commercial constraints and exploring transparent multiplayer artificial intelligence](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) (from "How Data is Shaping our Games") - [Developing fail-safes and defenses through adversarial training](https://www.wearedevelopers.com/videos/627-machine-learning-promising-but-perilous) (from "Machine Learning: Promising, but Perilous") - [Iterating extensively through evaluations and red teaming](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) (from "Introduction to Responsible AI: Balancing Value and Risk") ## Related Articles - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Dev Digest 110 - XY marks the spotty security](https://www.wearedevelopers.com/magazine/411-dev-digest-110-xy-marks-the-spotty-security) ## Related Jobs - [Engineer, Offensive Security Organization](https://www.wearedevelopers.com/jobs/ext/1992296-engineer-offensive-security-organization) at **Twilio** - [Staff Engineer - Offensive Security](https://www.wearedevelopers.com/jobs/ext/1226927-staff-engineer-offensive-security) at **Twilio** - [Security Architect - AI](https://www.wearedevelopers.com/jobs/ext/1581899-security-architect-ai) at **ZEISS Group** - [Software Engineer, Platform Engineering (L2)](https://www.wearedevelopers.com/jobs/ext/1956829-software-engineer-platform-engineering-l2) at **Twilio** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/319507-principal-engineer-ai-search-vector-infrastructure) at **Redis** - [Staff Developer Advocate, GitHub Security Lab](https://www.wearedevelopers.com/jobs/ext/1921051-staff-developer-advocate-github-security-lab) at **GitHub**