System Architecture

MAD Platform

The full system in one picture. Walkthrough below.
Lane 1 — Client & SME queues
Lane 2 — Core review pipeline (live)
Lane 3 — Databases & platform ecosystem
Lane 4 — Parallel async self-improvement loops
MAD Platform architecture diagram, four lanes. Lane 1: a Live Web App URL entry point, and an SME Central Review Hub with three sub-queues (Live Review Queue for low-confidence triage, Structural Review, Prompt Review Queue for pattern approvals). Lane 2, the core live pipeline: Orchestrator Agent fans out to Analyst Agent's three parallel checks (Rule Checks, Visual AI, Semantic AI), which merge into Editor Agent, then a Retry Gate capped at one pass, then Reporter Agent, then a Route and Act escalation gate that either goes to Action Agent directly (high confidence) or to the SME Live Review Queue (low confidence); Action Agent files Jira tickets and posts Slack notifications. Lane 3: Cloud Storage for HTML reports and Firestore for async status logs. Lane 4: two scheduled loops, a daily WCAG freshness check (fetch current WCAG version, classify the change, auto-refresh the vector knowledge base) and a weekly Cloud Run Job running a self-hosted Gemma model that mines dismissal patterns, checks them against existing patterns, and once an SME approves, grounds Editor Agent's prompt with the optimized rules.

How it works

The diagram above shows the mechanism; this section walks through why each piece exists.
Entry point
Web app: scan a site on demand
A URL submitted through the live front end kicks off a Review Cycle directly.
Live
Review Cycle
Sequential where order matters, parallel where it doesn't, dynamic delegation only at genuine judgment points.
Orchestrator: select pages agent
Fetches the entry page, reads nav links, decides which pages carry real risk (forms, checkout, account flows): a bounded, justified subset, not exhaustive crawling.
orchestration: dynamic delegation
Analyst: parallel per-page checks tool + agent
Independent checks with no ordering dependency, run concurrently. Tuned high-recall: over-flagging is acceptable, missing a real violation is not.
Rule checkscontrast, alt text, headings, labels, ARIA, tab order
Visual AIscreenshot review: focus indicators, visual contrast
Semantic AIgrounded in the WCAG knowledge base below
orchestration: parallel fan-out
Editor: independent verification agent
Re-checks every flag against the underlying evidence. Dismisses false positives with a documented reason; assigns a validated confidence score to what survives.
Orchestrator: retry gate agent
"Good enough, or one more pass?" Capped at exactly one retry: a bounded loop, not open-ended.
feedback loop: bounded
Reporter: rank and draft agent
Ranks confirmed findings by real-world risk (WCAG level, litigation pattern frequency, user impact) and drafts the styled report against one fixed template.
Orchestrator: route and act
One escalation gate: low confidence or critical severity routes to a human; everything else is fully autonomous.
Action AgentFiles a real Jira ticket automatically. Idempotent: a retried step never double-files.
idempotency
SME Central Review HubConfirm → files the ticket, live-updates the report the customer already has open. Dismiss → never becomes a ticket. The same hub also holds WCAG version-change and pattern-approval requests from the two loops below.
human approval
Both paths post to Slack: a real-time alert when something's escalated, a summary when the scan completes.
Memory: WCAG knowledge base
Grounds Analyst and Editor's citations against the real standard instead of the model's unaided claim, and stays current on its own.
Cloud Scheduler tick, daily
→
Fetch current WCAG version
→
Classify: minor or major?
→
Auto-refresh or escalate to SME Central Review Hub
memory / vector searchfeedback loop: self-healhuman approval on structural change
Memory: Analyst/Editor's own judgment
Distinct from the knowledge base above: that loop keeps reference material current, this one tunes the pipeline's own behavior from its real history. A self-hosted Gemma, not Vertex AI: a background batch job with no live-request latency, exactly the workload Gemma's on-device design targets.
Editor's dismissal history
→
Weekly Scheduler tick
→
Cloud Run Job: self-hosted Gemma
→
Consistent pattern?
→
SME Central Review Hub
Confirmed → grounds Editor's prompt on every future scan.
persistent memoryfeedback loop: self-improvinghuman approval before judgment changesGemma, self-hosted
Resumability
Every stage above checkpoints its completion to Firestore as it finishes. A restart, whether a crash or a redeploy, reads the last completed checkpoint and resumes from the next incomplete stage. It never blindly restarts from the top, and never assumes a stage completed just because the job record exists.
resumabilitycrash-safe by design
Google Cloud infrastructure
Cloud Run
2 scale-to-zero services + 1 Job, least-privilege per service
Google ADK
LlmAgent + Runner wraps every real-time Gemini call
Vertex AI · Gemini
Flash-lite for volume, Flash for judgment calls
Ollama · Gemma
Self-hosted on the pattern-miner Job, the one background call, not real-time
Firestore
Job checkpoints, findings, escalations, KB version, learned patterns
Cloud Storage
Generated HTML reports
Cloud Scheduler
WCAG freshness tick (daily), Gemma pattern-miner (weekly)
Secret Manager
Access codes, Jira + Slack credentials
Jira Cloud
Real ticket filing, not a mock sink in production
Slack
Real-time escalation alerts, scan-complete summaries
Playwright
Headless render, screenshot, computed styles