The diagram above shows the mechanism; this section walks through why each piece exists.
Entry point
Web app: scan a site on demand
A URL submitted through the live front end kicks off a Review Cycle directly.
Live
Review Cycle
Sequential where order matters, parallel where it doesn't, dynamic delegation only at genuine judgment points.
Orchestrator: select pagesagent
Fetches the entry page, reads nav links, decides which pages carry real risk (forms, checkout, account flows): a bounded, justified subset, not exhaustive crawling.
orchestration: dynamic delegation
Analyst: parallel per-page checkstool + agent
Independent checks with no ordering dependency, run concurrently. Tuned high-recall: over-flagging is acceptable, missing a real violation is not.
Rule checkscontrast, alt text, headings, labels, ARIA, tab order
Semantic AIgrounded in the WCAG knowledge base below
orchestration: parallel fan-out
Editor: independent verificationagent
Re-checks every flag against the underlying evidence. Dismisses false positives with a documented reason; assigns a validated confidence score to what survives.
Orchestrator: retry gateagent
"Good enough, or one more pass?" Capped at exactly one retry: a bounded loop, not open-ended.
feedback loop: bounded
Reporter: rank and draftagent
Ranks confirmed findings by real-world risk (WCAG level, litigation pattern frequency, user impact) and drafts the styled report against one fixed template.
Orchestrator: route and act
One escalation gate: low confidence or critical severity routes to a human; everything else is fully autonomous.
Action AgentFiles a real Jira ticket automatically. Idempotent: a retried step never double-files.
idempotency
SME Central Review HubConfirm → files the ticket, live-updates the report the customer already has open. Dismiss → never becomes a ticket. The same hub also holds WCAG version-change and pattern-approval requests from the two loops below.
human approval
Both paths post to Slack: a real-time alert when something's escalated, a summary when the scan completes.
Memory: WCAG knowledge base
Grounds Analyst and Editor's citations against the real standard instead of the model's unaided claim, and stays current on its own.
Cloud Scheduler tick, daily
→
Fetch current WCAG version
→
Classify: minor or major?
→
Auto-refresh or escalate to SME Central Review Hub
memory / vector searchfeedback loop: self-healhuman approval on structural change
Memory: Analyst/Editor's own judgment
Distinct from the knowledge base above: that loop keeps reference material current, this one tunes the pipeline's own behavior from its real history. A self-hosted Gemma, not Vertex AI: a background batch job with no live-request latency, exactly the workload Gemma's on-device design targets.
Editor's dismissal history
→
Weekly Scheduler tick
→
Cloud Run Job: self-hosted Gemma
→
Consistent pattern?
→
SME Central Review Hub
Confirmed → grounds Editor's prompt on every future scan.
persistent memoryfeedback loop: self-improvinghuman approval before judgment changesGemma, self-hosted
Resumability
Every stage above checkpoints its completion to Firestore as it finishes. A restart, whether a crash or a redeploy, reads the last completed checkpoint and resumes from the next incomplete stage. It never blindly restarts from the top, and never assumes a stage completed just because the job record exists.
resumabilitycrash-safe by design
Google Cloud infrastructure
Cloud Run
2 scale-to-zero services + 1 Job, least-privilege per service
Google ADK
LlmAgent + Runner wraps every real-time Gemini call
Vertex AI · Gemini
Flash-lite for volume, Flash for judgment calls
Ollama · Gemma
Self-hosted on the pattern-miner Job, the one background call, not real-time