CargoLens is a shipping-document intelligence workspace for operations teams that live in their inbox. It classifies every inbound email — comparison request, SI request, invoice query, general, spam — scores its urgency, reads the attached Shipping Instruction and draft Bill of Lading, compares the seven fields that decide a shipment, and requests clearer evidence when bounded recovery cannot resolve a case. Typed judgments come from TypeSafe Jev; arithmetic, aliases, completion rules and every state transition stay in deterministic, unit-tested code.
Provider and mailbox credentials stay on the server: the
Chrome extension calls a local API that owns the TypeSafe key, the dashboard is
gated by a bearer token, and Gmail access flows through an encrypted OAuth
connection. Jev proposes; code verifies; only source evidence can clear a
document. A request for a future draft stays awaiting documents even when a
benchmark convention labels it OK.
flowchart LR
Inbox["Gmail, Outlook or dataset import"] --> Classify["Jev: category, urgency, document expectation"]
Classify --> Store["Durable case store and event stream"]
Store --> Readers["TXT, XLSX, DOCX and native PDF readers"]
Readers --> Recovery["OCR sidecar and recovery for unreadable pages"]
Recovery --> Compare["Deterministic seven-field comparison"]
Compare --> Decision{"Evidence sufficient?"}
Decision -->|Yes| Verified["Verified outcome and evidence-backed reply"]
Decision -->|No| Escalate["Request documents or escalate to a human"]
Verified -.-> Export["Benchmark projection, separate contract"]
Classification is the fast path: bounded packed Jev requests classify multiple emails. The dashboard receives server-sent events; extension previews use background runtime messages. Verification is the careful path: documents are read independently, candidates carry exact source locations, and a comparison claim is saved only after source-proof validation. Full cases and benchmark exports share the durable case store. Extension badges classify visible snippets only and cannot establish document verification.
Implemented today:
- Typed Jev category, urgency and document-expectation questions; batches of eight, bounded concurrency and retries, content caching, durable events.
- TXT/CSV/TSV, XLSX, DOCX and native PDF readers and label/value candidates with byte hashes and exact source locations. Candidate pairings are structural hypotheses for the field selector. Scanned pages are explicitly marked for OCR.
- An optional Python/Tesseract OCR sidecar and Node recovery reader that preserve native evidence separately, validate source hashes, and bound process concurrency, output size and timeouts. See the recovery reader contract. A dedicated verification harness recovers all 15 supplied scans under neutral and misleading filenames with hash-checked snapshots (details).
- Gmail OAuth client with PKCE, thread/reference lookup, durable outbound queue, four reply templates, stale-source checks and resume fixtures.
- Source-proof validation before saving a comparison claim or queueing a confirmation: bytes, excerpts and source pairs must agree, and the trusted role classifier and field comparator must establish their meaning.
- Version-bound operational drafts shared with the Gmail queue, and optional OpenRouter recovery of unresolved fields from selected server-read regions. Recovery candidates remain proposals until the field comparator validates them.
- Extension inbox triage: hide-spam, urgent pinning and plain-language smart filters judged by Jev, configured from the extension settings page (see Inbox actions and smart filters).
- A one-command organizer evaluation runner with frozen grouped manifests, export provenance, explicit failed rows, and source-backed dispute overlays.
Still assigned integration work:
- A complete official submission and OCR-aware field integration. The bounded source-only comparator runs during dataset import and evaluation, but its conservative layout/role/reference support does not yet export a valid 520-row submission. The OpenRouter vision-escalation provider exists (VISION.md) but is not yet wired into the API; automatic recovery orchestration, dashboard and workflow builder are also pending.
- Low-confidence or conflicting classifications emit recovery signals. Server
side text recovery is connected end to end through the authenticated
POST /cases/:id/recoverroute; automatic orchestration and semantic acceptance by the comparator remain pending.
A request for a future draft remains awaiting documents, never verified
solely because the benchmark labels it OK.
Follow the steps below for local API, OCR, dashboard and synthetic-demo setup.
- Node.js 20.19 or newer and npm.
- Python 3 with Tesseract for the optional OCR sidecar.
- A TypeSafe API key; Google Cloud OAuth credentials only for connector work.
Copy .env.example to .env. The running API requires TYPESAFE_API_KEY and
DASHBOARD_TOKEN; generate a long random token. An authorized operator enters this workspace token into the dashboard; provider and mailbox credentials remain server-side.
| Variable | Purpose |
|---|---|
TYPESAFE_API_KEY |
Jev inference. Never placed in the extension. |
DASHBOARD_TOKEN |
Workspace bearer token; public local preview and OAuth callback routes are listed below. |
HOST / PORT |
API bind address, default 127.0.0.1:3001. |
DATABASE_PATH |
SQLite location, default runtime/cargolens.sqlite. |
DATASET_ROOT |
Dataset root, default training_data/sdoc-hackathon-docker/extracted/data_v2. |
TYPESAFE_MODEL / JEV_PROMPT_VARIANT |
Pinned model and question variant; both feed the classifier revision. |
JEV_CONCURRENCY / JEV_BATCH_SIZE / JEV_REQUESTS_PER_MINUTE |
Throughput controls for classification. |
GMAIL_* |
OAuth registration, polling, sync query and automation switches; see Gmail configuration. |
OPENROUTER_API_KEY / OPENROUTER_TEXT_MODEL |
Optional server-side text recovery provider and model. |
RECOVERY_MAX_CASE_ATTEMPTS / RECOVERY_MAX_CASE_TOKENS / RECOVERY_MAX_CASE_USD |
Durable per-case recovery reservation ceilings, including failed attempts. |
AI_GATEWAY_API_KEY |
Reserved for gateway integration. |
npm ci
cp .env.example .env # then set TYPESAFE_API_KEY and DASHBOARD_TOKEN
npm run devThe API binds to http://127.0.0.1:3001. GET /health and POST /classify
are public local endpoints. Health includes a classifierRevision hash so
preview clients can invalidate cached results when the model, questions or
packing configuration changes. Operational routes require Authorization: Bearer <DASHBOARD_TOKEN>; /rules/evaluate and the OAuth callback/result routes are also public as listed below. The extension never receives API
keys or this token. Local environment files, SQLite databases and evaluation
outputs are ignored by Git.
POST /import imports the configured dataset and returns a job immediately.
Follow GET /events for results or use GET /emails; GET /cases/:id
includes the full operational state. SSE reconnects accept Last-Event-ID.
GET /usage is the authoritative request-level token total; packed
classifications have usage: null and reference their shared usageRequestId.
Response-level request usage can overlap across concurrent preview calls, so do
not sum it for billing.
POST /classify accepts up to 20 unique-ID preview rows:
{"emails":[{"id":"visible-row-1","subject":"Please check the draft BL","from":"sender@example.com","snippet":"Compare the attached draft with our SI."}]}Preview results use only the visible subject/snippet and cannot establish document verification. Public previews do not overwrite imported cases.
With the local API running:
npm run build:extensionIn Chrome's extension manager, enable Developer mode, choose Load unpacked,
and select apps/extension/dist. Open the CargoLens popup to check the local
API and enable previews, or open Inbox actions & smart filters for the full
settings page. The supported inbox hosts are Gmail, Outlook Live and
Outlook Office. Ctrl+Shift+L toggles previews. Reload existing mailbox tabs
after loading or updating the extension.
Badges classify visible subject/snippet text; the urgent tray provides shortcuts to native rows. No mailbox OAuth token is needed for these previews. Full-thread retrieval and outbound automation use the Gmail connector below. Gmail and Outlook Live were checked in signed-in Chrome on 2026-09-20.
The settings page turns Jev's judgments into triage, all preview-only and reversible — nothing is deleted, moved or marked in the mailbox:
- Hide spam rows at a confidence you choose. Hidden rows can be shown again from the CargoLens tray, and a confident spam label wins over any filter, so phishing that literally matches an "action needed" filter stays hidden.
- Pin urgent rows to the top of the visible list: Jev's "today" and "blocking" urgency levels sort above everything else, most urgent first.
- Smart filters: plain-language conditions Jev judges as a yes/no
probability for every visible row (
POST /rules/evaluate, one Noul question per rule, packed and cached per email). Shipping-focused presets — needs my reply now, cargo or release blocked, vessel/cut-off changes, charges disputes, automated notices, marketing — ship disabled; each filter chooses its own action (pin, highlight or hide) and sensitivity threshold. Only the wording of a condition costs a new judgment; changing actions or thresholds re-applies instantly from cached probabilities.
Preset thresholds were chosen from a stratified run over the benchmark dataset
plus hand-written shipping scenarios (npx tsx --env-file=.env tools/eval/src/rules.ts); raw scores are written to runtime/eval/rules/.
Successful previews are cached for five minutes in browser session storage, up to 200 entries. Cache keys include the mailbox context, local API endpoint, classifier revision and content fingerprint, and are stored as hashes. The cache stores classification metadata, not email subjects, senders or snippets. Changed content or classifier configuration triggers a new classification; Retry bypasses the cached result. API failures remain visible and are not cached as successful classifications.
Add GMAIL_CLIENT_ID, GMAIL_CLIENT_SECRET, GMAIL_MAILBOX_ADDRESS and
GMAIL_OAUTH_REDIRECT_URI to .env. Register the exact callback URI in Google
Cloud; the local default is http://127.0.0.1:3001/gmail/oauth/callback. The
connector requests only gmail.readonly and gmail.send.
With the server running, call authenticated POST /gmail/oauth/start, then
open the returned URL and consent with the configured mailbox. The server
checks one-use state, PKCE, required scopes and mailbox identity before saving
the refresh token encrypted in SQLite. Alternatively, an existing
GMAIL_REFRESH_TOKEN seeds the encrypted store once. A disconnected or revoked
connection is not silently re-enabled by that environment variable: reconnect
through OAuth. Encryption is bound to the client secret and mailbox; rotating
either requires reconnection. Keep both the environment file and runtime
database private.
GMAIL_POLL_ENABLED=true enables bounded polling independently of sending.
GMAIL_SYNC_QUERY defaults to inbox excluding spam and trash. The mailbox/query
cursor survives restarts; failed pages are retried, invalid cursors restart the
scan, and full rounds rescan for new replies. This is polling, not Gmail push
or history synchronization.
GMAIL_AUTOMATION_ENABLED=false permits ingestion and queued-reply inspection
without sending. Setting it to true enables eligible dispatch during sync and
through the dispatch route. Set GMAIL_DOCUMENTATION_CONTACT only to the
responsible documentation contact: requests for a draft from us are routed
there, or ask for clarification when responsibility is unknown. No BL is
fabricated.
Eligible real-time Gmail verification uses the same hybrid field extraction,
comparison fallback, placeholder policy and optional vision recovery as
npm run pipeline:full. The shared result is converted into source-bound
operational field evidence before any confirmation or amendment can be queued.
Live read, full-thread ingestion, attachment bytes and classification were verified on 2026-09-20 with outbound sending disabled. Delivery remains covered by fixtures; no live email was sent during this verification.
| Method | Route | Auth | Purpose |
|---|---|---|---|
GET |
/health |
public | Service status and classifierRevision. |
POST |
/classify |
public, budgeted | Up to 20 preview rows for the extension. |
POST |
/rules/evaluate |
public, budgeted | Smart-filter probabilities for up to 20 rows against up to 8 rules. |
GET |
/usage |
bearer | Authoritative request-level token totals. |
GET |
/dashboard |
bearer | Aggregated dashboard report snapshot and data mode. |
GET |
/runs/:id |
bearer | One import/run report. |
POST |
/import |
bearer | Import the configured dataset; returns a job. |
GET |
/emails |
bearer | Imported inbox with classification state. |
GET |
/cases/:id |
bearer | Full operational state for one case. |
GET |
/cases/:id/activity |
bearer | Recent durable events for one case. |
GET |
/cases/:id/sources |
bearer | Re-read attachment sources with hash checks. |
GET |
/cases/:id/comparison |
bearer | Stored document comparison evidence. |
POST |
/cases/:id/compare |
bearer | Run the bounded document comparison; requires verified comparison intent. |
POST |
/cases/:id/retry |
bearer | Re-run a failed or signalled case. |
POST |
/cases/:id/decision |
bearer | Record a human decision. |
POST |
/cases/:id/draft |
bearer | Persist an evidence-bound reply draft; requires source and decision versions. |
POST |
/cases/:id/recover |
bearer | Propose unresolved fields from selected attachment source regions. |
GET |
/events |
bearer | Server-sent events; honours Last-Event-ID. |
POST |
/gmail/oauth/start |
bearer | Begin the OAuth consent flow. |
GET |
/gmail/oauth/callback |
public | One-use OAuth callback (state + PKCE). |
GET |
/gmail/connection-result |
public | Generic connection-result page. |
GET |
/gmail/status |
bearer | Connector configuration and polling state. |
POST |
/gmail/sync |
bearer | Manual ingestion, e.g. {"maxMessages":1}. |
GET |
/gmail/outbox |
bearer | Inspect queued replies. |
POST |
/gmail/dispatch |
bearer | Dispatch eligible queued replies. |
POST |
/gmail/disconnect |
bearer | Revoke and clear the stored connection. |
flowchart TD
Ext["Chrome extension (no credentials)"] -->|POST /classify, /rules/evaluate| API["Hono API"]
Dash["Dashboard (bearer token)"] --> API
Gmail["Gmail connector (OAuth + PKCE)"] --> API
API --> Jev["TypeSafe Jev: typed judgments"]
API --> Code["Code: parsing, arithmetic, policy, state"]
API --> DB[("SQLite + durable events")]
API --> OCR["Python OCR sidecar + recovery"]
Three boundaries hold the system honest:
- Labels never enter inference. Ground-truth labels and corrected references are evaluation-only. Filenames and IDs support retrieval and provenance, never ground-truth lookup or a shortcut to a field value. Jev answers what the sender requests and what the source establishes — never which answer a scorer would reward.
- The model proposes, code decides. Jev returns typed choices and probabilities; parsing, arithmetic, alias handling, completion rules and state transitions are code-owned and unit-tested.
- Product state is separate from benchmark output. A versioned export
adapter projects operational state into the organizer schema; a benchmark
OKcannot authorize sending, clear a blocker or mark documents verified.
npm test # vitest suites across api, shared and extension
npm run typecheck # strict TypeScript across workspaces
npm run lint # eslintnpm run eval -- --prepare freezes the official dataset, scorer, grouped splits
and configuration without model calls. npm run eval runs the API classification
pipeline, exports supported decisions and invokes the unchanged organizer scorer.
npm run eval -- --offline exercises the plumbing with no provider calls and
deliberately exits nonzero. Incomplete comparison states are explicit export
failures, so a partial run is never presented as a valid headline score. Live
evaluation makes paid model calls.
See the evaluation guide for offline checks, dispute
adjudication, and final-run protection. npm run eval:final republishes the
frozen final scorecards with no provider calls, and npm run eval:triage
reproduces the offline miss triage; both refuse existing output directories.
npm run eval:classify retains the separate
category tuning harness. OCR recovery has a dedicated verification harness
(tools/eval/src/ocr-verification.ts plus a frozen label audit) proving all 15
supplied scans recover under neutral and misleading filenames; see
its README.
On 19 September 2026, the cold HTTP API import classified and saved 520 emails in 7.80 seconds using 64 Jev requests (505 distinct content states). Category agreement was 513/520 (98.65%); macro-F1 was 0.9878. All seven misses carried an uncertainty/recovery signal in this run. This does not establish perfect recovery. Estimated input cost was $0.0351 at the supplied $0.042/M rate, excluding any output charge. A persisted re-import took 0.56 seconds with no new provider calls.
These are individual measured runs, not latency guarantees, and exclude attachment extraction and comparison.
On 21 September 2026, the shared API/evaluation pipeline processed all 520 rows including conservative document comparison in 12.58 seconds. Category macro-F1 was 0.97938; 359 rows exported and 161 failed explicitly. Exact defect detection was 0/46 and review recall 0/20. These targets are not met, and the partial submission has no valid headline score. See the issue #37 report and frozen artifacts for actual results, failure accounting, environment versions and reproducibility hashes.
The frozen final run is published as the issue #60
scorecards: the unchanged
organizer scorer replayed the submission exactly, the official headline remains
unavailable (official.valid: false; the raw 0.230045 diagnostic includes
scorer defaults for absent rows and is not a headline), and the issue #39
triage accounts for all 520
rows: 284 agreements without a lossy convention, 75 future-draft label
conventions, 8 inference failures, 1 reader failure and 152 semantic/intent or
document-role blockers, with three pending disputes and zero accepted
corrections. A twelve-case independently authored document challenge measured
12/12 exact operational workflow and blocker checks, 6/12 exact exported
status, 0/12 false match confirmations and 4/12 proof-validated
confirmation/amendment decisions; the six blocked exports remain misses.
Reproduce with npm run eval:final.
Smart-filter presets were calibrated on 20 September 2026: 40 dataset rows
(stratified, 8 per category, rendered as inbox snippets) plus 12 hand-written
shipping scenarios. The tightened cargo_blocked preset no longer fires on
phishing spam (0.03 mean on SPAM) while scoring 0.92–0.98 on real customs
holds and unreleased delivery orders; every hand-written scenario fires
exactly its expected presets. Raw scores: runtime/eval/rules/presets.json.
The downloaded sdoc-hackathon-docker.zip archive is stored at
training_data/sdoc-hackathon-docker/. The extracted bundle contains:
data_v2/: 520 synthetic shipping-document inbox records, attachments, sample submissions, and ground truth.server/: FastAPI serving, loading, bundle-building, and scoring utilities.docker-compose.yml: optional local scoring server configuration.
The dataset includes data_v2/ground_truth.json, so this repository must
remain private until the answer key is intentionally released. See
the dataset README
for the schema, categories, edge cases, and regeneration instructions.
apps/api/ Hono service: pipeline, document readers, OCR recovery, Gmail connector
apps/dashboard/ React/Vite operations dashboard
apps/extraction-dashboard/ Extraction trace research workspace (Vite)
apps/extension/ MV3 Chrome extension for Gmail and Outlook web
packages/shared/ Zod schemas plus Jev question and smart-filter contracts
tools/eval/ Evaluation harness: official runner, smart filters, OCR verification
tools/ocr-sidecar/ Python/Tesseract OCR CLI
docs/ Connector, backend and acceptance verification notes
training_data/ SDOC dataset, attachments and organizer scorer
- The source-only comparator now runs during dataset import and evaluation, but its conservative layout/role/reference support does not yet produce a complete official submission. Failures remain explicit and block a valid score.
- Live Gmail read, threading and attachment ingestion were verified on 2026-09-20 with sending disabled; delivery is covered by fixtures only.
- Text recovery is available through the authenticated API; automatic recovery orchestration and semantic acceptance by the comparator remain pending.
- The final scorecards are published (issue #60); a valid 520-row official submission, the defect/review performance targets and human adjudication of the three pending disputes remain outstanding.
After rebuilding and reloading the extension, choose Open document verification in its popup. Connect to the configured local CargoLens API using the dashboard token, then select an imported message. The page displays all seven field outcomes, source excerpts and locators, known mismatches, blockers, and the current decision version. The token remains in page memory and is never saved to Chrome sync storage. Use Refresh evidence after new documents arrive. Inbox snippet labels remain intent previews, not document-clearance decisions.
AI_PROVIDER=openrouter selects the OpenRouter Jev transport for API classification; the default is typesafe. Smart filters use the direct TypeSafe key when configured. Extraction research commands retain source candidates and vision proposals separately; unsupported units and incomplete evidence cannot establish a verified match. The Fly image includes the Python/Tesseract OCR runtime and excludes the additional evaluation dataset from the production image.