The knowledge retiring technicians carry isn't in any manual.
A voice AI mentor that captures the tacit expertise of senior technicians and delivers it to juniors on the shop floor.
Capture, qualify and leverage all staff's unwritten knowledge. Just ask.
Demo · How it works · How answers are graded · What's real · Getting started
Built during the Activate Your Voice hackathon — Speechmatics × The AI Collective Paris
Track 1: Communication & Human Experience · 28 February – 1 March 2026
Rule #1 — Lore never contradicts a SOP. It completes it.
The FAA issued around 9,000 new mechanic certificates in 2024. More than 68,000 certificated mechanics — one in three — reach retirement age within ten years, about 6,800 a year (ATEC Pipeline Report, United States).
The headcount roughly balances. What each side carries does not: thirty years on an engine type leaves, zero arrives.
Every retiring senior takes decades of contextual knowledge that was never written down — the exceptions, the quirks of specific airframes, the patterns a manual cannot hold. Zymbly, LexX and AWS Q all do RAG over explicit documents: manuals, SOPs, service bulletins.
RAG retrieves what you put in. Lore extracts what seniors never thought to write down — through active dialogue, not passive ingestion.
flowchart LR
subgraph IN["Voice in"]
S["Senior debrief"]
J["Junior question"]
end
STT["Speechmatics<br/>real-time STT"]
subgraph API["Routes"]
C["/api/capture"]
Q["/api/query"]
L["/api/log"]
end
subgraph MEM["Backboard memory"]
TT["Thread per technician<br/>Marc · 26 yrs on CFM56"]
TA["Thread per aircraft<br/>F-GKXA"]
DOC["Auto-indexed SOPs"]
end
SYN["GPT-4o synthesis<br/>SOP first, always"]
OUT["OpenAI TTS"]
S --> STT
J --> STT
STT --> C
STT --> Q
STT --> L
C --> TT
C --> TA
L --> TA
TT --> SYN
TA --> SYN
DOC --> SYN
Q --> SYN
SYN --> OUT
OUT --> J
classDef voice fill:#1e293b,stroke:#64748b,color:#e2e8f0
classDef route fill:#312e81,stroke:#6366f1,color:#e0e7ff
classDef mem fill:#134e4a,stroke:#14b8a6,color:#ccfbf1
classDef synth fill:#7c2d12,stroke:#f97316,color:#ffedd5
class S,J,STT,OUT voice
class C,Q,L route
class TT,TA,DOC mem
class SYN synth
Three modes, one loop:
| Mode | Who | What happens |
|---|---|---|
| Capture | Senior, after an intervention | Lore interviews by voice — asks follow-ups, flags ambiguity — and stores the knowledge against airframe, component and conditions |
| Query | Junior, hands in the machine | Three layers come back, procedure first |
| Log | Junior, after the job | "Lore, log: N1 vibration 2.4 units, 6 °C, no escalation." Straight into the aircraft's memory |
What a query answer is made of:
| Layer | Source |
|---|---|
| 1 · The procedure | RAG over the official SOP / AMM documents. Always cited first. |
| 2 · Oral knowledge | What a named senior actually said — dated, with the conditions attached. |
| 3 · This airframe | The maintenance history of the tail number in front of you. |
Hangar, 23h. Thomas — junior, two years in — is alone on an Airbus A320, tail F-GKXA, CFM56-5B engine. He finds an anomaly that isn't on the job card.
Thomas: "Lore, I'm on F-GKXA, CFM56-5B, fan section. I'm getting an unusual
low-frequency vibration on the N1 shaft. It's not in the job card.
What do I know about this?"
Lore: "According to SOP 72-21-00, N1 vibration above 4 units requires
escalation. However, Marc noted in October that F-GKXA specifically
shows a harmonic resonance between 2-3 units in cold conditions —
below 8°C. It's a known characteristic of this airframe, not a defect.
He recommended logging it and monitoring across the next two cycles
before escalating."
Thomas: "Got it. Logging: vibration at 2.4 units, temperature 6°C, no
escalation per Marc's note. Lore, add this to F-GKXA's memory."
Lore: "Logged. F-GKXA memory updated."
Every maintenance answer closes the same way: «Vérifie toujours la procédure AMM avant d'intervenir.»
The failure mode here is not a crash. It is a well-formed answer carrying the wrong threshold, and test coverage cannot see that. So the answers are graded, not just the code.
64 cases · 12 categories · 10 graders, every one a pure function. A grader that needs a model to decide cannot gate a safety property, because it fails in the same way as the thing it grades.
Not every failure is equal, so they are tiered:
| Tier | Threshold | What it covers |
|---|---|---|
| safety | 100%, no budget | invented figures, a refusal followed by a guess, contradicting the procedure, manufactured consensus, misclassified bands |
| trust | 95% | named attribution, procedure cited first |
| form | 90% | wording and closing discipline |
The headline pass rate decides nothing on its own. 63 of 64 is a clean WARN when the miss sits in the form tier — and a failed run when it sits in safety. Across every canary run so far, the safety tier has held at 100%, 0.00pp.
npm run check # offline gate: unit tests, golden answers, regressions. No API keys, no cost.What the gate actually covers, and what it costs
npm run check=npm test+npm run evals+npm run evals:regression. Deterministic, offline, free.- 21 of the 64 cases have a reference answer, so that is what the offline run grades. The rest only run live.
- Offline targets measure the ruler. Only
npm run evals:synthesis(lib/llm.ts) andnpm run evals:live(/api/query→ Backboard) measure the product. - CI: the offline gate blocks every PR. The coverage gate blocks on FAIL, reports on WARN. The live gate is advisory and opt-in — a PR has to carry the
run-live-evallabel, because a full run is ~70 model calls at roughly €0.25–0.30 and CI should not be able to spend money on its own initiative. - Beyond the absolute floors, a per-tier drop of more than 3pp against the frozen baseline is a FAIL by itself.
The invariants and the case set: frontend/evals/README.md. What "green" means, tier by tier: frontend/evals/ACCEPTANCE.md. The protocol behind it, and where it came from: Harnesses, graders, closed loops.
| Layer | Technology |
|---|---|
| Voice in | Speechmatics real-time STT — browser-side WebSocket, noise-robust |
| LLM | OpenAI GPT-4o |
| Memory + RAG | Backboard — thread per aircraft, thread per technician, auto-indexed SOP documents |
| Voice out | OpenAI TTS (gpt-4o-mini-tts, fallback tts-1) |
| Frontend | Next.js 14 App Router · TypeScript · Tailwind CSS |
| Deploy | Vercel |
Eleven API routes live in frontend/app/api/; the core loop runs on three — capture, query, log.
Install, configure, seed, run
# Install dependencies
npm install
# Configure environment
cp frontend/.env.example frontend/.env.local
# Fill in API keys in frontend/.env.local
# Create/validate Backboard assistant + threads
npm run setup-backboard
# Optional: seed demo memory
npm run seed-backboard
# Run frontend dev server from root
npm run devOpen http://localhost:3000.
To get exactly what was demoed on stage:
git checkout v0.1-hackathonV1 was built in 24 hours by a team of 4 — 43 commits, from 2026-02-28 17:47 to 2026-03-01 16:37, first line of code to final demo. That build is frozen at the v0.1-hackathon tag. Development continues on main, which may diverge from the demo.
This is a prototype, not production software.
The demo-vs-production boundary, stated plainly
- Synthetic SOPs and mock aircraft data. No real EASA-regulated documents, no real tail numbers, no real technician data.
- API routes have no authentication.
- Contradiction detection — flagging it when two seniors disagree rather than averaging them — is design intent, not shipped code. The eval set now seeds a second expert written to disagree with Marc, which is what makes the claim testable at all.
- Multi-source confidence scoring — weighting an observation confirmed by four technicians over three years above an isolated note — is likewise a direction, not an implemented store. Same for knowledge-graph relationships across airframe × component × condition × expert.
- Capture and log extraction are ungraded; the harness covers query synthesis only.
Full framework: docs/trust-safety.md.
Human-centered layer — and explicitly not a decision twin
Digital twins in most industries model an asset: a building, an engine, a warehouse. The extension toward people is more recent, and the reference work on human-centered digital twins frames it as modelling a worker in order to support human judgement rather than replace it — which is exactly the constraint here, since in Part-145 maintenance a licensed technician signs the release and nothing else can.
Lore sits on that human-centered layer: it models what a technician knows, not what a machine does. It is not a decision twin. That label implies simulating a policy before applying it — "if I change this threshold, what happens" — and Lore has no such capability, nor the ground truth it would need. Nothing tells it whether a captured observation was right, and the moment a technician escalates the counterfactual disappears. Building simulation on top of what Lore captures is a plausible next system. It is not this one.
Ivan de Murard — @IvandeMurard
V1 (v0.1-hackathon) was built in 24 hours by a team of 4. Development since the hackathon is mine.
Full contributor list: AUTHORS.md · Licensed under MIT.