Skip to content

Latest commit

 

History

131 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CargoLens

Every shipping email understood. Every document decision evidence-backed.
Issue board · Document readers · Evaluation harness · Final scorecards · SI/BL extraction architecture

CargoLens is a shipping-document intelligence workspace for operations teams that live in their inbox. It classifies every inbound email — comparison request, SI request, invoice query, general, spam — scores its urgency, reads the attached Shipping Instruction and draft Bill of Lading, compares the seven fields that decide a shipment, and requests clearer evidence when bounded recovery cannot resolve a case. Typed judgments come from TypeSafe Jev; arithmetic, aliases, completion rules and every state transition stay in deterministic, unit-tested code.

Provider and mailbox credentials stay on the server: the Chrome extension calls a local API that owns the TypeSafe key, the dashboard is gated by a bearer token, and Gmail access flows through an encrypted OAuth connection. Jev proposes; code verifies; only source evidence can clear a document. A request for a future draft stays awaiting documents even when a benchmark convention labels it OK.

The flagship journey: from inbox to verified documents

flowchart LR
  Inbox["Gmail, Outlook or dataset import"] --> Classify["Jev: category, urgency, document expectation"]
  Classify --> Store["Durable case store and event stream"]
  Store --> Readers["TXT, XLSX, DOCX and native PDF readers"]
  Readers --> Recovery["OCR sidecar and recovery for unreadable pages"]
  Recovery --> Compare["Deterministic seven-field comparison"]
  Compare --> Decision{"Evidence sufficient?"}
  Decision -->|Yes| Verified["Verified outcome and evidence-backed reply"]
  Decision -->|No| Escalate["Request documents or escalate to a human"]
  Verified -.-> Export["Benchmark projection, separate contract"]
Loading

Classification is the fast path: bounded packed Jev requests classify multiple emails. The dashboard receives server-sent events; extension previews use background runtime messages. Verification is the careful path: documents are read independently, candidates carry exact source locations, and a comparison claim is saved only after source-proof validation. Full cases and benchmark exports share the durable case store. Extension badges classify visible snippets only and cannot establish document verification.

What works now — and what still needs a partner

Implemented today:

  • Typed Jev category, urgency and document-expectation questions; batches of eight, bounded concurrency and retries, content caching, durable events.
  • TXT/CSV/TSV, XLSX, DOCX and native PDF readers and label/value candidates with byte hashes and exact source locations. Candidate pairings are structural hypotheses for the field selector. Scanned pages are explicitly marked for OCR.
  • An optional Python/Tesseract OCR sidecar and Node recovery reader that preserve native evidence separately, validate source hashes, and bound process concurrency, output size and timeouts. See the recovery reader contract. A dedicated verification harness recovers all 15 supplied scans under neutral and misleading filenames with hash-checked snapshots (details).
  • Gmail OAuth client with PKCE, thread/reference lookup, durable outbound queue, four reply templates, stale-source checks and resume fixtures.
  • Source-proof validation before saving a comparison claim or queueing a confirmation: bytes, excerpts and source pairs must agree, and the trusted role classifier and field comparator must establish their meaning.
  • Version-bound operational drafts shared with the Gmail queue, and optional OpenRouter recovery of unresolved fields from selected server-read regions. Recovery candidates remain proposals until the field comparator validates them.
  • Extension inbox triage: hide-spam, urgent pinning and plain-language smart filters judged by Jev, configured from the extension settings page (see Inbox actions and smart filters).
  • A one-command organizer evaluation runner with frozen grouped manifests, export provenance, explicit failed rows, and source-backed dispute overlays.

Still assigned integration work:

  • A complete official submission and OCR-aware field integration. The bounded source-only comparator runs during dataset import and evaluation, but its conservative layout/role/reference support does not yet export a valid 520-row submission. The OpenRouter vision-escalation provider exists (VISION.md) but is not yet wired into the API; automatic recovery orchestration, dashboard and workflow builder are also pending.
  • Low-confidence or conflicting classifications emit recovery signals. Server side text recovery is connected end to end through the authenticated POST /cases/:id/recover route; automatic orchestration and semantic acceptance by the comparator remain pending.

A request for a future draft remains awaiting documents, never verified solely because the benchmark labels it OK.

Quick start

Follow the steps below for local API, OCR, dashboard and synthetic-demo setup.

Prerequisites

  • Node.js 20.19 or newer and npm.
  • Python 3 with Tesseract for the optional OCR sidecar.
  • A TypeSafe API key; Google Cloud OAuth credentials only for connector work.

Environment variables

Copy .env.example to .env. The running API requires TYPESAFE_API_KEY and DASHBOARD_TOKEN; generate a long random token. An authorized operator enters this workspace token into the dashboard; provider and mailbox credentials remain server-side.

Variable Purpose
TYPESAFE_API_KEY Jev inference. Never placed in the extension.
DASHBOARD_TOKEN Workspace bearer token; public local preview and OAuth callback routes are listed below.
HOST / PORT API bind address, default 127.0.0.1:3001.
DATABASE_PATH SQLite location, default runtime/cargolens.sqlite.
DATASET_ROOT Dataset root, default training_data/sdoc-hackathon-docker/extracted/data_v2.
TYPESAFE_MODEL / JEV_PROMPT_VARIANT Pinned model and question variant; both feed the classifier revision.
JEV_CONCURRENCY / JEV_BATCH_SIZE / JEV_REQUESTS_PER_MINUTE Throughput controls for classification.
GMAIL_* OAuth registration, polling, sync query and automation switches; see Gmail configuration.
OPENROUTER_API_KEY / OPENROUTER_TEXT_MODEL Optional server-side text recovery provider and model.
RECOVERY_MAX_CASE_ATTEMPTS / RECOVERY_MAX_CASE_TOKENS / RECOVERY_MAX_CASE_USD Durable per-case recovery reservation ceilings, including failed attempts.
AI_GATEWAY_API_KEY Reserved for gateway integration.

Run the API

npm ci
cp .env.example .env   # then set TYPESAFE_API_KEY and DASHBOARD_TOKEN
npm run dev

The API binds to http://127.0.0.1:3001. GET /health and POST /classify are public local endpoints. Health includes a classifierRevision hash so preview clients can invalidate cached results when the model, questions or packing configuration changes. Operational routes require Authorization: Bearer <DASHBOARD_TOKEN>; /rules/evaluate and the OAuth callback/result routes are also public as listed below. The extension never receives API keys or this token. Local environment files, SQLite databases and evaluation outputs are ignored by Git.

POST /import imports the configured dataset and returns a job immediately. Follow GET /events for results or use GET /emails; GET /cases/:id includes the full operational state. SSE reconnects accept Last-Event-ID. GET /usage is the authoritative request-level token total; packed classifications have usage: null and reference their shared usageRequestId. Response-level request usage can overlap across concurrent preview calls, so do not sum it for billing.

POST /classify accepts up to 20 unique-ID preview rows:

{"emails":[{"id":"visible-row-1","subject":"Please check the draft BL","from":"sender@example.com","snippet":"Compare the attached draft with our SI."}]}

Preview results use only the visible subject/snippet and cannot establish document verification. Public previews do not overwrite imported cases.

Load the extension

With the local API running:

npm run build:extension

In Chrome's extension manager, enable Developer mode, choose Load unpacked, and select apps/extension/dist. Open the CargoLens popup to check the local API and enable previews, or open Inbox actions & smart filters for the full settings page. The supported inbox hosts are Gmail, Outlook Live and Outlook Office. Ctrl+Shift+L toggles previews. Reload existing mailbox tabs after loading or updating the extension.

Badges classify visible subject/snippet text; the urgent tray provides shortcuts to native rows. No mailbox OAuth token is needed for these previews. Full-thread retrieval and outbound automation use the Gmail connector below. Gmail and Outlook Live were checked in signed-in Chrome on 2026-09-20.

Inbox actions and smart filters

The settings page turns Jev's judgments into triage, all preview-only and reversible — nothing is deleted, moved or marked in the mailbox:

  • Hide spam rows at a confidence you choose. Hidden rows can be shown again from the CargoLens tray, and a confident spam label wins over any filter, so phishing that literally matches an "action needed" filter stays hidden.
  • Pin urgent rows to the top of the visible list: Jev's "today" and "blocking" urgency levels sort above everything else, most urgent first.
  • Smart filters: plain-language conditions Jev judges as a yes/no probability for every visible row (POST /rules/evaluate, one Noul question per rule, packed and cached per email). Shipping-focused presets — needs my reply now, cargo or release blocked, vessel/cut-off changes, charges disputes, automated notices, marketing — ship disabled; each filter chooses its own action (pin, highlight or hide) and sensitivity threshold. Only the wording of a condition costs a new judgment; changing actions or thresholds re-applies instantly from cached probabilities.

Preset thresholds were chosen from a stratified run over the benchmark dataset plus hand-written shipping scenarios (npx tsx --env-file=.env tools/eval/src/rules.ts); raw scores are written to runtime/eval/rules/.

Successful previews are cached for five minutes in browser session storage, up to 200 entries. Cache keys include the mailbox context, local API endpoint, classifier revision and content fingerprint, and are stored as hashes. The cache stores classification metadata, not email subjects, senders or snippets. Changed content or classifier configuration triggers a new classification; Retry bypasses the cached result. API failures remain visible and are not cached as successful classifications.

Gmail configuration

Add GMAIL_CLIENT_ID, GMAIL_CLIENT_SECRET, GMAIL_MAILBOX_ADDRESS and GMAIL_OAUTH_REDIRECT_URI to .env. Register the exact callback URI in Google Cloud; the local default is http://127.0.0.1:3001/gmail/oauth/callback. The connector requests only gmail.readonly and gmail.send.

With the server running, call authenticated POST /gmail/oauth/start, then open the returned URL and consent with the configured mailbox. The server checks one-use state, PKCE, required scopes and mailbox identity before saving the refresh token encrypted in SQLite. Alternatively, an existing GMAIL_REFRESH_TOKEN seeds the encrypted store once. A disconnected or revoked connection is not silently re-enabled by that environment variable: reconnect through OAuth. Encryption is bound to the client secret and mailbox; rotating either requires reconnection. Keep both the environment file and runtime database private.

GMAIL_POLL_ENABLED=true enables bounded polling independently of sending. GMAIL_SYNC_QUERY defaults to inbox excluding spam and trash. The mailbox/query cursor survives restarts; failed pages are retried, invalid cursors restart the scan, and full rounds rescan for new replies. This is polling, not Gmail push or history synchronization.

GMAIL_AUTOMATION_ENABLED=false permits ingestion and queued-reply inspection without sending. Setting it to true enables eligible dispatch during sync and through the dispatch route. Set GMAIL_DOCUMENTATION_CONTACT only to the responsible documentation contact: requests for a draft from us are routed there, or ask for clarification when responsibility is unknown. No BL is fabricated.

Eligible real-time Gmail verification uses the same hybrid field extraction, comparison fallback, placeholder policy and optional vision recovery as npm run pipeline:full. The shared result is converted into source-bound operational field evidence before any confirmation or amendment can be queued.

Live read, full-thread ingestion, attachment bytes and classification were verified on 2026-09-20 with outbound sending disabled. Delivery remains covered by fixtures; no live email was sent during this verification.

API routes

Method Route Auth Purpose
GET /health public Service status and classifierRevision.
POST /classify public, budgeted Up to 20 preview rows for the extension.
POST /rules/evaluate public, budgeted Smart-filter probabilities for up to 20 rows against up to 8 rules.
GET /usage bearer Authoritative request-level token totals.
GET /dashboard bearer Aggregated dashboard report snapshot and data mode.
GET /runs/:id bearer One import/run report.
POST /import bearer Import the configured dataset; returns a job.
GET /emails bearer Imported inbox with classification state.
GET /cases/:id bearer Full operational state for one case.
GET /cases/:id/activity bearer Recent durable events for one case.
GET /cases/:id/sources bearer Re-read attachment sources with hash checks.
GET /cases/:id/comparison bearer Stored document comparison evidence.
POST /cases/:id/compare bearer Run the bounded document comparison; requires verified comparison intent.
POST /cases/:id/retry bearer Re-run a failed or signalled case.
POST /cases/:id/decision bearer Record a human decision.
POST /cases/:id/draft bearer Persist an evidence-bound reply draft; requires source and decision versions.
POST /cases/:id/recover bearer Propose unresolved fields from selected attachment source regions.
GET /events bearer Server-sent events; honours Last-Event-ID.
POST /gmail/oauth/start bearer Begin the OAuth consent flow.
GET /gmail/oauth/callback public One-use OAuth callback (state + PKCE).
GET /gmail/connection-result public Generic connection-result page.
GET /gmail/status bearer Connector configuration and polling state.
POST /gmail/sync bearer Manual ingestion, e.g. {"maxMessages":1}.
GET /gmail/outbox bearer Inspect queued replies.
POST /gmail/dispatch bearer Dispatch eligible queued replies.
POST /gmail/disconnect bearer Revoke and clear the stored connection.

Architecture and trust boundaries

flowchart TD
  Ext["Chrome extension (no credentials)"] -->|POST /classify, /rules/evaluate| API["Hono API"]
  Dash["Dashboard (bearer token)"] --> API
  Gmail["Gmail connector (OAuth + PKCE)"] --> API
  API --> Jev["TypeSafe Jev: typed judgments"]
  API --> Code["Code: parsing, arithmetic, policy, state"]
  API --> DB[("SQLite + durable events")]
  API --> OCR["Python OCR sidecar + recovery"]
Loading

Three boundaries hold the system honest:

  1. Labels never enter inference. Ground-truth labels and corrected references are evaluation-only. Filenames and IDs support retrieval and provenance, never ground-truth lookup or a shortcut to a field value. Jev answers what the sender requests and what the source establishes — never which answer a scorer would reward.
  2. The model proposes, code decides. Jev returns typed choices and probabilities; parsing, arithmetic, alias handling, completion rules and state transitions are code-owned and unit-tested.
  3. Product state is separate from benchmark output. A versioned export adapter projects operational state into the organizer schema; a benchmark OK cannot authorize sending, clear a blocker or mark documents verified.

Commands and verification

npm test          # vitest suites across api, shared and extension
npm run typecheck # strict TypeScript across workspaces
npm run lint      # eslint

npm run eval -- --prepare freezes the official dataset, scorer, grouped splits and configuration without model calls. npm run eval runs the API classification pipeline, exports supported decisions and invokes the unchanged organizer scorer. npm run eval -- --offline exercises the plumbing with no provider calls and deliberately exits nonzero. Incomplete comparison states are explicit export failures, so a partial run is never presented as a valid headline score. Live evaluation makes paid model calls. See the evaluation guide for offline checks, dispute adjudication, and final-run protection. npm run eval:final republishes the frozen final scorecards with no provider calls, and npm run eval:triage reproduces the offline miss triage; both refuse existing output directories. npm run eval:classify retains the separate category tuning harness. OCR recovery has a dedicated verification harness (tools/eval/src/ocr-verification.ts plus a frozen label audit) proving all 15 supplied scans recover under neutral and misleading filenames; see its README.

Measured results

On 19 September 2026, the cold HTTP API import classified and saved 520 emails in 7.80 seconds using 64 Jev requests (505 distinct content states). Category agreement was 513/520 (98.65%); macro-F1 was 0.9878. All seven misses carried an uncertainty/recovery signal in this run. This does not establish perfect recovery. Estimated input cost was $0.0351 at the supplied $0.042/M rate, excluding any output charge. A persisted re-import took 0.56 seconds with no new provider calls.

These are individual measured runs, not latency guarantees, and exclude attachment extraction and comparison.

On 21 September 2026, the shared API/evaluation pipeline processed all 520 rows including conservative document comparison in 12.58 seconds. Category macro-F1 was 0.97938; 359 rows exported and 161 failed explicitly. Exact defect detection was 0/46 and review recall 0/20. These targets are not met, and the partial submission has no valid headline score. See the issue #37 report and frozen artifacts for actual results, failure accounting, environment versions and reproducibility hashes.

The frozen final run is published as the issue #60 scorecards: the unchanged organizer scorer replayed the submission exactly, the official headline remains unavailable (official.valid: false; the raw 0.230045 diagnostic includes scorer defaults for absent rows and is not a headline), and the issue #39 triage accounts for all 520 rows: 284 agreements without a lossy convention, 75 future-draft label conventions, 8 inference failures, 1 reader failure and 152 semantic/intent or document-role blockers, with three pending disputes and zero accepted corrections. A twelve-case independently authored document challenge measured 12/12 exact operational workflow and blocker checks, 6/12 exact exported status, 0/12 false match confirmations and 4/12 proof-validated confirmation/amendment decisions; the six blocked exports remain misses. Reproduce with npm run eval:final.

Smart-filter presets were calibrated on 20 September 2026: 40 dataset rows (stratified, 8 per category, rendered as inbox snippets) plus 12 hand-written shipping scenarios. The tightened cargo_blocked preset no longer fires on phishing spam (0.03 mean on SPAM) while scoring 0.92–0.98 on real customs holds and unreleased delivery orders; every hand-written scenario fires exactly its expected presets. Raw scores: runtime/eval/rules/presets.json.

Training data

The downloaded sdoc-hackathon-docker.zip archive is stored at training_data/sdoc-hackathon-docker/. The extracted bundle contains:

  • data_v2/: 520 synthetic shipping-document inbox records, attachments, sample submissions, and ground truth.
  • server/: FastAPI serving, loading, bundle-building, and scoring utilities.
  • docker-compose.yml: optional local scoring server configuration.

The dataset includes data_v2/ground_truth.json, so this repository must remain private until the answer key is intentionally released. See the dataset README for the schema, categories, edge cases, and regeneration instructions.

Project map

apps/api/           Hono service: pipeline, document readers, OCR recovery, Gmail connector
apps/dashboard/     React/Vite operations dashboard
apps/extraction-dashboard/  Extraction trace research workspace (Vite)
apps/extension/     MV3 Chrome extension for Gmail and Outlook web
packages/shared/    Zod schemas plus Jev question and smart-filter contracts
tools/eval/         Evaluation harness: official runner, smart filters, OCR verification
tools/ocr-sidecar/  Python/Tesseract OCR CLI
docs/               Connector, backend and acceptance verification notes
training_data/      SDOC dataset, attachments and organizer scorer

Known limitations and next proof points

  • The source-only comparator now runs during dataset import and evaluation, but its conservative layout/role/reference support does not yet produce a complete official submission. Failures remain explicit and block a valid score.
  • Live Gmail read, threading and attachment ingestion were verified on 2026-09-20 with sending disabled; delivery is covered by fixtures only.
  • Text recovery is available through the authenticated API; automatic recovery orchestration and semantic acceptance by the comparator remain pending.
  • The final scorecards are published (issue #60); a valid 520-row official submission, the defect/review performance targets and human adjudication of the three pending disputes remain outstanding.

Extension document verification

After rebuilding and reloading the extension, choose Open document verification in its popup. Connect to the configured local CargoLens API using the dashboard token, then select an imported message. The page displays all seven field outcomes, source excerpts and locators, known mismatches, blockers, and the current decision version. The token remains in page memory and is never saved to Chrome sync storage. Use Refresh evidence after new documents arrive. Inbox snippet labels remain intent previews, not document-clearance decisions.

AI_PROVIDER=openrouter selects the OpenRouter Jev transport for API classification; the default is typesafe. Smart filters use the direct TypeSafe key when configured. Extraction research commands retain source candidates and vision proposals separately; unsupported units and incomplete evidence cannot establish a verified match. The Fly image includes the Python/Tesseract OCR runtime and excludes the additional evaluation dataset from the production image.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages