What an ad resilience tabletop exercise is
An ad resilience tabletop exercise is a facilitated, discussion-based rehearsal of a written incident runbook against a realistic advertising failure scenario — run in a meeting room or call, with **no changes made to live accounts**. Participants receive scenario "injects" (a mock restriction notice, a screenshot of zeroed-out delivery, a simulated alert) and talk through exactly what they would do, in what order, using the actual documents, dashboards, and contact lists they would rely on in a real event.
It is different from a live failover test. A tabletop validates **decision-making, ownership, and evidence readiness**; a live test validates that technical failover actually works. Mature programs do both, but the tabletop comes first because it is cheaper and exposes process gaps safely.
A good tabletop is grounded in real platform workflows, not hypotheticals. For example, Google's compromised-account process expects you to gather specific evidence before reporting: timestamps from the change history showing unauthorized access, unauthorized manager account IDs linked to the hierarchy, evidence of budget increases or rules that deviate from historical management, and your current IP address. A tabletop asks a blunt question: could your team produce that evidence within 30 minutes, at 2 a.m., with your most senior buyer offline? If the answer is "probably," the exercise just paid for itself. Practitioner runbook templates bundle the same components — role sheets, tiered checklists for containment/triage/recovery, dashboard wireframes, and communications templates — which map directly onto tabletop structure.
Why quarterly rehearsals matter
Practitioner incident-response guidance is explicit: test your runbook with a quarterly tabletop exercise, time the responses, and iterate. Quarterly is the sweet spot for high-spend operations for four reasons:
- **Access and security debt accumulate between drills.** Unused accounts, stale admin grants, and over-privileged users creep in. Google's Security Agent exists precisely because this drift happens continuously — it flags dormant users, unusual behavior, and access levels that exceed what a role requires, and Google recommends reviewing its suggestions at least monthly. A quarterly exercise forces the same discipline at the team level: who has admin access to what, and should they still? - **People rotate.** New hires have never seen your runbook; departed staff may still appear in your escalation tree. A drill surfaces both. - **Platform workflows change.** Recovery flows, security features (passkeys, multi-party approval, Security Agent), and enforcement behaviors evolve. A runbook written 12 months ago may reference steps that no longer exist. - **Timing data only exists if you rehearse.** You cannot know whether your team can contain a P1 incident in under two hours — a practitioner SLA target — unless you have timed a rehearsal.
One honest caveat: rehearsal improves your internal response speed and decision quality. It does not influence platform review outcomes, and no cadence of drilling guarantees faster reinstatement by a platform.
Core scenarios to rehearse
Rotate one primary scenario per quarter so the full set is covered roughly every 18 months. The core library for high-spend operators:
**1. Compromised account (Google Ads).** Walk the full flow: report the compromise with evidence, re-authenticate admins via MFA/2SV, review the change log, choose a cleanup decision, then audit users, MCC links, payment profiles, and campaigns. Two details make this scenario valuable: only one administrator can submit the cleanup decision and it cannot be changed afterward — so designate decision authority in advance — and the timeline for lifting a suspension varies depending on the compromised activity, so rehearse stakeholder communication under uncertainty. Reimbursement for unauthorized charges requires completed recovery and 2SV, and billing investigations can take 10 to 15 working days.
**2. Restriction cascade across asset types.** Rehearse distinguishing — and separately containing — four different things platforms can restrict: an individual **person's** profile/access, an **ad account**, a **Page**, and a **Business Portfolio**. Map which assets share dependencies (payment methods, domains, pixels) so one restriction doesn't surprise you by cascading. Meta does not document review timelines or outcomes for restrictions in the cited materials; your exercise should practice communicating *without* promising restoration dates.
**3. Platform outage vs. account-specific problem.** Use the four-type taxonomy: full delivery outage, attribution outage, reporting lag, and partial signal loss. The drilled rule: confirm the type before touching anything. During an attribution outage — spend pacing normally but conversions showing zero — do not pause; delivery is working and pausing triggers an unnecessary learning-phase reset. Community signals (DownDetector, r/FacebookAds) have preceded official acknowledgment by 1–3 hours in past major outages, so rehearse your early-detection stack.
**4. Billing / payment failure.** Rehearse what happens when a payment profile is unlinked: Google warns that impacted accounts stop running ads because their billing setups are deactivated and need a new billing setup to resume. Drill your separate-payment-methods-per-account discipline.
**5. Automation misfire.** A bad automated rule or a budget keyed in wrong. Meta can spend up to 75% over a daily budget on a given day (weekly spend stays within 7x the daily budget), and Ads Manager now requires confirmation for extreme budget inputs — rehearse kill-switch and rollback paths, and the habit of previewing rules before saving.
**6. Ownership and offboarding event.** A departing admin or an agency exit. Drill access revocation, credential rotation, and ownership transfer. Two hard rules from Meta's own guidance anchor this scenario: storing customer passwords is not an approved model, and shared/fake user logins risk suspension as spam.
How to run a tabletop exercise
**Two weeks before:** Pick one scenario. Appoint a facilitator (runs injects, not a participant) and a note-taker. Freeze scope: which accounts, platforms, and runbooks are in play.
**Prepare injects:** Realistic artifacts — a mock suspension email, a doctored screenshot of a delivery collapse, a simulated Security Agent flag, a "client Slack message." The more realistic the artifact, the more honest the rehearsal.
**Run a 60–90 minute session:** Walk the timeline in order:
1. **0–30 minutes (containment):** Declare the incident, assign roles, run baseline checks, decide containment actions. Practitioner checklists treat these first-30-minute steps as non-negotiable. 2. **30–120 minutes (triage):** Classify the incident, scope the impact, open the platform ticket, draft the first stakeholder update. 3. **4–48 hours (recovery):** Progressive re-enablement, monitoring, reconciliation.
**Hard rules for the exercise itself:**
- **No live account changes during a tabletop.** Discussion only. If you want to validate automation behavior, do it separately: preview rules before saving them (preview makes no permanent changes), run new rules once before scheduling recurrence, and reserve chaos/fault-injection experiments for sandboxed environments with kill-switches — not for the meeting. - **Time every decision.** The output is not "we talked about it" — it is "detection to declared incident: 11 minutes; first stakeholder update drafted: 34 minutes." - **Debrief within 72 hours, blameless.** Assign each gap an owner and a deadline. Update the runbook before the quarter closes; re-run any failed segment in the next exercise.
Ready to upgrade your ad account infrastructure?
AdsInfra provides certified agency accounts for Meta, TikTok, and Google. Setup in 2-5 business days.
Talk to a SpecialistRoles and escalation paths
Assign named owners **and backups** for each role before the exercise; the drill's job is to find out whether the backups can actually perform.
- **Incident Commander (IC):** owns the timeline, triage decisions, and stakeholder updates; escalates to executives and legal. - **Ad Ops Lead:** executes campaign-level containment (pauses, exclusions, budget moves). - **Analytics Lead:** confirms impact and scope from governed dashboards. - **Platform Liaison:** opens support tickets, works platform reps, tracks case IDs. - **Legal/Compliance and PR/Comms:** review policy notices; pre-approve holding statements. - **Finance:** quantifies revenue at risk and manages reallocation.
**Escalation paths need three things nailed down in advance:**
1. **Thresholds.** Define what triggers a P1 (e.g., RPM drop >25% vs. rolling baseline, ad served rate collapsing, or account-level disapprovals across multiple campaigns) and who can declare it. Practitioner targets worth drilling against: first stakeholder update under 30 minutes; containment under 2 hours for P1 incidents. 2. **Authority for irreversible steps.** Google's compromise flow allows only one administrator to submit the cleanup decision, and it cannot be reverted. Decide now which admins hold that authority — including off-hours. 3. **Ownership, access, billing, and offboarding maps.** Document who legally owns each Business Portfolio, ad account, and Page; who holds admin vs. partner access; who owns each payment method; and the exact offboarding sequence for departing staff or agencies (revoke access, rotate credentials, transfer assets). Meta maintains official best-practice guidance for asset management — use it as the reference frame when you build this ownership map. Meta's developer guidance reinforces the hygiene underneath it: employees should use their own logins rather than shared or fake users (which risk suspension as spam), and storing customer passwords is not an approved model. A tabletop that reveals "only Priya knows the MCC password" has found a single point of failure before an incident did.
Measuring and improving resilience
Treat resilience as a measured capability, not a vibe. Track a small scorecard across quarters:
- **Time to detect** (inject to declared incident). - **Time to first stakeholder update** — target under 30 minutes. - **Time to containment** — practitioner target: under 2 hours for P1 incidents. - **Time to recovery plan** — practitioner guidance uses recovery to baseline or stable alternate within 72 hours as a P1 target. - **Gap closure rate:** percentage of debrief action items completed before the next quarter's drill. This is the metric that separates programs that improve from programs that perform theater. - **Hygiene trend:** stale access entries, untested backups, and outdated contacts found per drill should trend toward zero. Pair quarterly drills with monthly security hygiene reviews — Google explicitly recommends checking Security Agent suggestions at least once a month.
For the automation layer, borrow SRE practice: golden-run regression suites that replay known inputs and compare outputs, and controlled chaos experiments (propagation delays, telemetry outages) with kill-switches that automatically revert to the previous safe policy if pacing or exclusion thresholds are violated. These validate that your technical safeguards behave the way the tabletop assumes they do.
Finally, keep the claims honest: these metrics measure *your* response. They cannot measure or guarantee platform behavior — suspension-lift timelines vary by case, Meta does not publish restriction-review timelines in the cited materials, and reimbursement investigations run on the platform's schedule (Google states 10–15 working days for billing investigations). Resilience work narrows your blast radius and shortens your decision time; the platform side remains outside your control.