How it works
The diagram above shows the mechanism; this section walks through why each piece exists.
Entry point
Web app: scan a site on demand
A URL submitted through the live front end enqueues a scan (Cloud Tasks) and returns immediately -- it does not run the pipeline itself.
Live
Queue
scan-onboarding (the public, always-on service above) only ever enqueues -- it never runs the pipeline in its own request handler. Cloud Tasks (scan-queue) sits between it and the pipeline, so a burst of submissions queues instead of the public service falling over, and a scan that dies mid-run gets retried without the visitor resubmitting anything. scan-worker, a separate, not-publicly-reachable Cloud Run service that only Cloud Tasks' dedicated invoker identity can call, is what actually runs the Review Cycle below.
decoupled executionbackpressure via queue, not failure
Review Cycle
Runs inside scan-worker, one scan per instance at a time. Sequential where order matters, parallel where it doesn't, dynamic delegation only at genuine judgment points.
Orchestrator: select pages
agent
Fetches the entry page, reads nav links, decides which pages carry real risk (forms, checkout, account flows): a bounded, justified subset, not exhaustive crawling.
orchestration: dynamic delegation
Editor: independent verification
agent
Re-checks every flag against the underlying evidence. Dismisses false positives with a documented reason; assigns a validated confidence score to what survives.
Orchestrator: retry gate
agent
"Good enough, or one more pass?" Capped at exactly one retry: a bounded loop, not open-ended.
feedback loop: bounded
Reporter: rank and draft
agent
Ranks confirmed findings by real-world risk (WCAG level, litigation pattern frequency, user impact) and drafts the styled report against one fixed template.
Orchestrator: route and act
One escalation gate: low confidence or critical severity routes to a human; everything else is fully autonomous.
Action AgentExports confirmed findings as a CSV with Jira's importer columns (Summary/Description), plus emails the full report to the address the scan was submitted with. No ticket-tracker account required. Idempotent: a retried step never double-exports.
idempotency
SME Review QueueConfirm → joins the same CSV, live-updates the report the customer already has open. Dismiss → never becomes a finding, with a documented reason. The same queue also holds WCAG version-change and pattern-approval requests from the two loops below.
human approval
Slack alerts and a real ticket-tracker sink (Jira) are supported in code as opt-in integrations for anyone self-hosting this — the public free deployment runs on the CSV + email path only, no external accounts needed.
Memory: WCAG knowledge base
Grounds Analyst and Editor's citations against the real standard instead of the model's unaided claim, and stays current on its own. Runs as its own always-off Cloud Run service (scan-wcag-poller), woken only by its daily Scheduler tick — it never runs inside a live scan.
Cloud Scheduler tick, daily
→
Fetch current WCAG version
→
Classify: minor or major?
→
Auto-refresh or escalate to SME Central Review Hub
memory / vector searchfeedback loop: self-healhuman approval on structural change
Memory: Analyst/Editor's own judgment
Distinct from the knowledge base above: that loop keeps reference material current, this one tunes the pipeline's own behavior from its real history. Calls Gemini Flash via Vertex AI, the same tier reserved for judgment calls elsewhere in this pipeline — call volume here is low (one call per qualifying cluster, weekly), so a managed API beats maintaining a separate self-hosted model. Runs as a Cloud Run Job (pattern-miner), not a Service — woken weekly by its own Scheduler tick, run-to-completion, then exits.
Editor's dismissal history
→
Weekly Scheduler tick
→
Cloud Run Job: Gemini Flash
→
Consistent pattern?
→
SME Central Review Hub
Confirmed → grounds Editor's prompt on every future scan.
persistent memoryfeedback loop: self-improvinghuman approval before judgment changes
Resumability
Every stage above checkpoints its completion to Firestore as it finishes. A restart, whether a crash or a redeploy, reads the last completed checkpoint and resumes from the next incomplete stage. It never blindly restarts from the top, and never assumes a stage completed just because the job record exists.
A related, distinct failure mode: Cloud Tasks retries a dispatch that runs past its deadline, and deadline expiry does not mean the first attempt has stopped -- so a slow scan can get a second worker while the first is still running. A per-job Firestore lease (transactional claim, re-entrant for the same owner) is what stops two live workers from double-spending every Gemini call and both emailing the same report.
resumabilitycrash-safe by designexactly-once execution
Google Cloud infrastructure
Cloud Run
3 scale-to-zero services (scan-onboarding, scan-worker, scan-wcag-poller) + 1 Job (pattern-miner), least-privilege per service
Cloud Tasks
scan-queue: decouples the public service from the pipeline, so a traffic burst queues instead of failing
Google ADK
LlmAgent + Runner wraps every real-time Gemini call
Vertex AI · Gemini
Flash-lite for volume, Flash for judgment calls
Firestore
Job checkpoints, findings, escalations, KB version, learned patterns
Cloud Storage
Generated HTML reports
Cloud Scheduler
WCAG freshness tick (daily), pattern-miner (weekly)
Resend
Emails the report to the address a scan was submitted with
CSV export
Default ticket sink — Jira's importer column format, no tracker account needed
Playwright
Headless render, screenshot, computed styles
Optional, opt-in for anyone self-hosting this, not required for the public deployment: an access-code gate on the scan form and review queue (Secret Manager), a real Jira sink instead of CSV, and Slack alerts alongside email.