System Architecture

MAD Platform

The full system in one picture. Walkthrough below.
MAD Platform, process flow with distributed state. Ingestion (scale-to-zero): a public scan request web form enqueues onto a Cloud Tasks scan queue, which dispatches a worker. That worker runs inside the Scan Worker zone (Cloud Run, Google ADK + Gemini): Orchestrator selects the high-risk page subset via Playwright discovery, Analyst Agent runs three parallel checks -- rules, visual AI, semantic AI -- producing findings, Editor Agent independently verifies them, dismissing false positives and scoring confidence, a Retry Gate sends not-good-enough results back to Analyst (capped at one loop) and good-enough ones to Reporter Agent, which ranks by real-world risk and drafts the report, then an Escalation Gate asks whether any finding is low-confidence or critical: no routes autonomously to Sinks and Delivery for a CSV export and emailed report, yes escalates to an SME Review Queue in that same zone, whose confirmation also reaches the CSV export. Underneath, a Firestore State and Guardrails zone coordinates all of this: a transactional Job Lease Lock stops two workers claiming the same job, and a Stage Checkpoints DB saves state after the lock is granted, after the pipeline's findings stage, and after the final report, so a crash or redeploy resumes from the last checkpoint instead of from zero; the lock is released once the report reaches its sink. A separate Grounding and Memory Loops zone runs asynchronously, not per scan: a daily Cloud Scheduler tick drives scan-wcag-poller, which refreshes the WCAG vector knowledge base that Analyst and Editor cite; a weekly Cloud Scheduler tick drives a pattern-miner Job, a Gemini analysis of past dismissals that grounds Editor's prompts with confirmed false-positive patterns.
The full process flow, including the Firestore-backed job lease and checkpoint state that make a scan resumable, and the two async loops that keep it grounded without running on every scan.

How it works

The diagram above shows the mechanism; this section walks through why each piece exists.
Entry point
Web app: scan a site on demand
A URL submitted through the live front end enqueues a scan (Cloud Tasks) and returns immediately -- it does not run the pipeline itself.
Live
Queue
scan-onboarding (the public, always-on service above) only ever enqueues -- it never runs the pipeline in its own request handler. Cloud Tasks (scan-queue) sits between it and the pipeline, so a burst of submissions queues instead of the public service falling over, and a scan that dies mid-run gets retried without the visitor resubmitting anything. scan-worker, a separate, not-publicly-reachable Cloud Run service that only Cloud Tasks' dedicated invoker identity can call, is what actually runs the Review Cycle below.
decoupled executionbackpressure via queue, not failure
Review Cycle
Runs inside scan-worker, one scan per instance at a time. Sequential where order matters, parallel where it doesn't, dynamic delegation only at genuine judgment points.
Orchestrator: select pages agent
Fetches the entry page, reads nav links, decides which pages carry real risk (forms, checkout, account flows): a bounded, justified subset, not exhaustive crawling.
orchestration: dynamic delegation
Analyst: parallel per-page checks tool + agent
Independent checks with no ordering dependency, run concurrently. Tuned high-recall: over-flagging is acceptable, missing a real violation is not.
Rule checkscontrast, alt text, headings, labels, ARIA, tab order
Visual AIscreenshot review: focus indicators, visual contrast
Semantic AIgrounded in the WCAG knowledge base below
orchestration: parallel fan-out
Editor: independent verification agent
Re-checks every flag against the underlying evidence. Dismisses false positives with a documented reason; assigns a validated confidence score to what survives.
Orchestrator: retry gate agent
"Good enough, or one more pass?" Capped at exactly one retry: a bounded loop, not open-ended.
feedback loop: bounded
Reporter: rank and draft agent
Ranks confirmed findings by real-world risk (WCAG level, litigation pattern frequency, user impact) and drafts the styled report against one fixed template.
Orchestrator: route and act
One escalation gate: low confidence or critical severity routes to a human; everything else is fully autonomous.
Action AgentExports confirmed findings as a CSV with Jira's importer columns (Summary/Description), plus emails the full report to the address the scan was submitted with. No ticket-tracker account required. Idempotent: a retried step never double-exports.
idempotency
SME Review QueueConfirm → joins the same CSV, live-updates the report the customer already has open. Dismiss → never becomes a finding, with a documented reason. The same queue also holds WCAG version-change and pattern-approval requests from the two loops below.
human approval
Slack alerts and a real ticket-tracker sink (Jira) are supported in code as opt-in integrations for anyone self-hosting this — the public free deployment runs on the CSV + email path only, no external accounts needed.
Memory: WCAG knowledge base
Grounds Analyst and Editor's citations against the real standard instead of the model's unaided claim, and stays current on its own. Runs as its own always-off Cloud Run service (scan-wcag-poller), woken only by its daily Scheduler tick — it never runs inside a live scan.
Cloud Scheduler tick, daily
→
Fetch current WCAG version
→
Classify: minor or major?
→
Auto-refresh or escalate to SME Central Review Hub
memory / vector searchfeedback loop: self-healhuman approval on structural change
Memory: Analyst/Editor's own judgment
Distinct from the knowledge base above: that loop keeps reference material current, this one tunes the pipeline's own behavior from its real history. Calls Gemini Flash via Vertex AI, the same tier reserved for judgment calls elsewhere in this pipeline — call volume here is low (one call per qualifying cluster, weekly), so a managed API beats maintaining a separate self-hosted model. Runs as a Cloud Run Job (pattern-miner), not a Service — woken weekly by its own Scheduler tick, run-to-completion, then exits.
Editor's dismissal history
→
Weekly Scheduler tick
→
Cloud Run Job: Gemini Flash
→
Consistent pattern?
→
SME Central Review Hub
Confirmed → grounds Editor's prompt on every future scan.
persistent memoryfeedback loop: self-improvinghuman approval before judgment changes
Resumability
Every stage above checkpoints its completion to Firestore as it finishes. A restart, whether a crash or a redeploy, reads the last completed checkpoint and resumes from the next incomplete stage. It never blindly restarts from the top, and never assumes a stage completed just because the job record exists.
A related, distinct failure mode: Cloud Tasks retries a dispatch that runs past its deadline, and deadline expiry does not mean the first attempt has stopped -- so a slow scan can get a second worker while the first is still running. A per-job Firestore lease (transactional claim, re-entrant for the same owner) is what stops two live workers from double-spending every Gemini call and both emailing the same report.
resumabilitycrash-safe by designexactly-once execution
Google Cloud infrastructure
Cloud Run
3 scale-to-zero services (scan-onboarding, scan-worker, scan-wcag-poller) + 1 Job (pattern-miner), least-privilege per service
Cloud Tasks
scan-queue: decouples the public service from the pipeline, so a traffic burst queues instead of failing
Google ADK
LlmAgent + Runner wraps every real-time Gemini call
Vertex AI · Gemini
Flash-lite for volume, Flash for judgment calls
Firestore
Job checkpoints, findings, escalations, KB version, learned patterns
Cloud Storage
Generated HTML reports
Cloud Scheduler
WCAG freshness tick (daily), pattern-miner (weekly)
Resend
Emails the report to the address a scan was submitted with
CSV export
Default ticket sink — Jira's importer column format, no tracker account needed
Playwright
Headless render, screenshot, computed styles
Optional, opt-in for anyone self-hosting this, not required for the public deployment: an access-code gate on the scan form and review queue (Secret Manager), a real Jira sink instead of CSV, and Slack alerts alongside email.