System Architecture

SHIP

The full system in one picture. Walkthrough below.
Agent step
Service / infra
Output / data artifact
Human-in-the-loop
Retry / failure path
SHIP-WEBHOOK · ALWAYS-ON LAMBDA, FAST-ACK ONLY SHIP-FRAGMENT-PROCESSOR · INDEPENDENTLY-SCALING LAMBDA, ONE JOB PER FRAGMENT GitHub PR opened / updated Verify + allowlist HMAC signature, repo check Screener whole-diff precheck any match? no Pass no model call spent yes Split + per-fragment Screener scan one file/function at a time Enqueue 1 SQS message per flagged fragment SQS fragment queue Detector RAG-grounded judgment Regulation corpus GDPR · EU AI Act · OWASP, sourced verbatim throttled → adaptive retry Triage per-category threshold ≥ threshold? no Log, continue no alert created yes fails repeatedly Dead-letter queue visible for manual triage Freeze: store alert DynamoDB, idempotent write Gate dashboard token-gated review console approve / reject? Resolved one-way, guarded
The two deployables and how they connect: the always-on webhook that only ever does cheap work, and the independently-scaling processor that does one flagged fragment's real review per invocation.

How it works

The diagram above shows the mechanism; this section walks through why each piece exists.
Entry point
GitHub webhook: a pull request opens or updates
The only trigger. No polling, no scheduled scan — SHIP reacts to the same event GitHub already fires for every PR.
ship-webhook · fast-ack Lambda
Everything in this Lambda is cheap, local, and bounded in milliseconds — no model call happens here, so GitHub's webhook delivery never has to wait on one.
Verify + allowlist tool
HMAC-SHA256 signature check against the configured webhook secret (fails closed if that secret is ever missing — there is no silent "allow unsigned" default), then the named repo against the connected-repositories table (one GetItem, point-lookup, fails closed if the table can't be read). A payload naming any repo not connected through Gate is rejected before it can spend this service's GitHub token or model budget.
Screener: whole-diff precheck tool
Plain regex/AST pattern matching, zero cost, zero network calls. If nothing in the whole diff matches any trigger term, the PR passes here and nothing downstream ever runs.
Split + per-fragment scan tool
The diff is split on file boundaries, then on top-level function/class boundaries within each file, so a PR touching several unrelated concerns doesn't get judged as one jumbled blob. Each fragment is screened individually — this is also where a fragment is isolated down to just the matched lines and their immediate context, keeping what gets queued small and focused.
Enqueue tool
One SQS message per flagged fragment — not one message for the whole PR. This is the design choice the rest of this document explains: making each fragment its own unit of work is what everything downstream depends on.
Why fragments are dispatched independently
The design decision this whole second Lambda exists to implement.

A simpler design — one Lambda invocation reviewing every flagged fragment in a PR, one after another — works fine for a PR with one or two issues. It stops working once a PR has enough flagged issues that reviewing all of them, one at a time, takes longer than that single Lambda invocation is allowed to run: AWS Lambda has a hard, non-negotiable ceiling on how long one invocation may run, and nothing resets that clock partway through. A live test against this exact failure mode confirmed it directly — a PR with several genuinely different flagged issues had its review silently cut short mid-way, with no error surfaced anywhere; the issues reviewed before the cutoff were recorded correctly, the rest simply never got a turn.

Making each fragment its own independent, queued job removes the shared clock entirely: every fragment gets its own full time budget, and fragments belonging to the same PR run concurrently rather than waiting in line behind each other. A PR with several flagged issues now takes about as long as its slowest single issue, not the sum of all of them.

ship-fragment-processor · one job per fragment
Triggered by the queue above, scaled independently from the webhook Lambda, with a concurrency cap so parallel reviews stay within the account's real model request-rate limit rather than overwhelming it.
Detector: RAG-grounded judgment agent
A Strands Agent that must call its retrieval tool against the sourced regulation corpus before judging anything — every citation it produces traces back to retrieved text, never an unaided claim. Returns a structured verdict: matched or not, which category, a 1–10 risk score, a plain-English explanation, the exact citation, and a draft fix.
RAG-grounded, never unaidedruns on Bedrock AgentCore Runtime
Triage: per-category threshold tool
Pure conditional routing, no judgment of its own. The freeze threshold is set per category rather than one shared number: a category where a confirmed violation is structurally irreversible or breaks a required safety guarantee (raw data reaching an external service, an automated decision with zero human checkpoint) freezes at a lower score than a category that's more a matter of degree.
Gate: the only way an alert resolves human
A token-gated dashboard (deliberately a separate credential from the webhook's own — compromising one can't silently disable the other) listing every frozen alert with its file, category, risk score, plain-English summary, citation, and suggested patch. The token is only ever pasted once: it establishes a signed-in session cookie, so every page after that is a plain link, not a secret sitting in a URL. Approve/Reject is the only way an alert leaves the "frozen" state, and it's a one-way, guarded transition: a duplicate webhook delivery or a retried job cannot silently reopen or overwrite a decision a human already made. The same decision is posted back to the PR as a comment and folded into a recomputed ship/compliance commit status — resolving the last blocking finding is what turns the check green. Gate also has a history view (every past disposition, with the reason a human gave) and a connected-repos view (connecting a repository is a form submission here, not a redeploy).
Failure handling
Two independent mechanisms, for two different failure shapes.
Slow, or rate-limitedA fragment that hits the model provider's own request-rate limit backs off and retries automatically (adaptive retry, configured once at the shared client rather than a hand-guessed delay) — it does not fail, it just takes longer, and it never blocks any other fragment's own review.
adaptive backoff
Genuinely brokenA fragment that fails outright (a malformed message, an unrecoverable error) is retried a bounded number of times by SQS itself, then moved to a dead-letter queue instead of being retried forever or silently dropped — visible for manual triage, not lost.
dead-letter queue

Every alert write is idempotent regardless of which path a fragment took to get there: the alert's own ID is derived from the specific violation (repo, PR, file, category, and the fragment text itself), not a random ID. A retried or redelivered fragment that reaches the same real violation overwrites the same row rather than creating a duplicate — and a write can never re-freeze or overwrite an alert a human has already resolved.

Detectors
Deliberately scoped to what a single PR diff can actually prove — no infrastructure state, no cross-file history, nothing that would require watching a repo over time. Each grounded in real, sourced text, not a model's general knowledge of the regulation.
PIIE-001
Raw, direct PII (SSN, account number, full profile) reaching an external sink with no masking
GDPR Article 32
PIIE-002
The same kind of raw PII, written to a log stream
GDPR Article 32
PIIE-003
Raw PII stored in a cache/session store with no encryption
GDPR Article 32
TLGP-002
An AI-produced decision applied as final with no human checkpoint anywhere in the fragment
EU AI Act Article 14
ALBP-001
A protected characteristic (or a clear proxy) directly driving a scoring calculation
EU AI Act Art. 10 + Annex III §5(b)
TLGP-001
A dangerous capability (shell exec, unscoped DB write) granted to an AI agent with no gate
OWASP LLM06:2025

Screener's job is to notice a candidate — a keyword, an import, a call shape — cheaply and without judgment. Detector's job is to decide whether it's real. The two are graded together: every detector above is verified against genuine violations and deliberate look-alikes engineered to trip a naive pattern match (a non-agent backup job that happens to call a dangerous function, a profile field that's merely displayed rather than scored, PII that's already hashed before use). A detector that flags the look-alike as a violation is not considered working, no matter how well it catches the real case.

The pipeline underneath is mechanically generic enough to flag other kinds of risk too, but staying scoped to AI-specific patterns is deliberate: it's the whole differentiation from a general-purpose static-analysis tool that has never heard of an LLM call. Widening scope to catch things like SQL injection would dilute that, not strengthen it.

AWS infrastructure
Lambda · ship-webhook
Always-on, fast-ack only, no model calls
Lambda · ship-fragment-processor
Independently-scaling, one flagged fragment per invocation
SQS
Fragment queue + dead-letter queue, concurrency-capped consumer
Bedrock AgentCore Runtime
Deployed Detector agent, callable remotely from either Lambda
Bedrock · Amazon Nova Lite
Primary reasoning model — picked over Claude Haiku 4.5 on a measured benchmark: same accuracy, 20× the request quota
Bedrock Titan Embeddings
Regulation-corpus retrieval
DynamoDB · ship-alerts
Idempotent alert storage, Gate's source of truth
DynamoDB · ship-repos
Connected repositories — the webhook's allowlist, point lookup
FastAPI + Mangum
Webhook route and the Gate dashboard, one deployable app
Two fallback model backends exist alongside Bedrock, switchable via one environment variable, both genuinely functional rather than unverified stretch goals: Google Gemini (for AWS credit exhaustion) and a local Ollama model (for a fully offline path with zero external dependency).