CourseForge turns a course idea, or the material you already have, into a tested, accessible e-learning course you can share as a single file, with every fact traced back to its source.
Building a one-hour course by hand takes weeks. Someone researches the subject, writes learning objectives, storyboards every screen, writes the assessment, builds it in an authoring tool, then tests that it works and that people using a keyboard or a screen reader can complete it.
General-purpose AI tools draft that material in minutes, and the draft cannot be trusted. Facts arrive without sources. Nothing is reviewed except by the person who asked for it. Accessibility is whatever the model happened to produce, and there is no record of what was checked, what was changed, or who approved it. That may be tolerable for onboarding trivia, but safety, compliance, clinical and financial training have to stand up to scrutiny later.
CourseForge sits between those two. The agent does the reading and writing; the system decides what must be checked, who checks it, when a person has to approve, and whether the result is allowed to ship.
Give it any of these:
- a one-line description of the course you want;
- a folder of your own material (
.md,.txt,.docx,.pdf,.html,.json); - work already in progress at any stage: a research dossier, a design blueprint, a storyboard, or a finished HTML course you want reviewed or rebuilt.
It runs an eleven-stage production line: research brief, sourced research dossier, instructional design, storyboard, editorial pass, visual direction, build, browser testing, release. What comes back is one self-contained HTML file, along with a QA report, a source report and a full record of how the course was made.
- Every claim carries its source. Research produces a dossier where each substantive statement gets a claim ID tied to a source, and those IDs follow the claim through design, storyboard and the finished screen. A course cannot be released if its citations do not hold together.
- Thirty-six specialist reviewers. They cover accuracy, objective alignment, assessment quality, reading level, accessibility, learner experience and visual design. Each one has a written rubric and a defined remit, and returns structured findings rather than a general opinion on the draft.
- Repairs are bounded and audited. Findings are deduplicated by code, adjudicated where reviewers disagree, then repaired by stable ID. A repair that touches anything outside its plan is rolled back. Repair cycles are capped, so a course cannot spin in review forever.
- You decide where a human approves. Run it with no gates, with the recommended gates, or stop at every stage. Higher-risk subjects raise the minimum automatically, and a course cannot quietly lower it.
- Accessibility decides whether the course ships. Every interaction is keyboard-operable, every figure has a text equivalent, and a real browser drives every screen with axe-core. Serious violations block release.
- Your own work is preserved. Anything you import is kept byte-for-byte and re-verified at release, so a rebuild cannot silently rewrite the source you gave it.
- No lock-in. The deliverable is one HTML file with no external requests. Email it, put it on a shared drive, or host it anywhere. A SCORM 1.2 package and simple completion tracking are available if you want them.
- No API keys. It drives the Claude Code or Codex CLI you are already signed in to.
Sixteen content block types (concepts, examples, comparisons, processes, evidence, warnings, misconceptions, scenarios and more) and six interaction types: single and multiple choice, matching, categorization, sequencing and reveal. Twenty-two diagram and chart archetypes rendered as real SVG rather than pictures of text. Six visual design families with light and dark themes, each contrast-checked. Knowledge checks throughout, a graded assessment with a pass mark, a glossary and a references list.
- Learning and development teams who need more courses than they have developer time for.
- Compliance, safety and onboarding training, where traceable sources and accessibility are requirements.
- Subject-matter experts who have the knowledge, and the documents, but no course developer.
- Consultancies and agencies producing training for clients, who need an audit trail of what was checked.
- Anyone modernizing an old course:
improvemode reviews existing HTML and rebuilds it, with a before-and-after report of what changed.
Code owns the process; the agent owns the words. Registries in config/ decide which stage runs, which reviewers
sit on it, which tools they may use, which renderer draws each figure, when a person must approve, and whether
the course may be released. Prompts are data files, not code. Every agent call runs against an execution plan
written down in advance, returns output checked against a schema, and is logged, so any run can be explained,
resumed or repeated. See the pipeline below.
This is version 0.1.0, with the gaps that implies:
- No video, no audio or narration, and no AI-generated imagery. Figures are diagrams, charts and icons.
- Navigation is linear. There are scenario screens, but no branching paths.
- The accessibility checks are automated. They are not a formal WCAG conformance audit.
- SCORM export has not yet been verified against a live learning management system.
- A full live run takes hours and costs model tokens. The tests and the demo run offline on fixtures.
- Scoring happens in the learner's browser, which suits training records rather than high-stakes exams.
The full list is under Limitations.
git clone <repo-url> CourseForge
cd CourseForge
./setup.sh # or .\setup.ps1 on WindowsThen run the offline demo, which replays committed fixtures from tests/fixtures/harness/demo and needs no
agent and no network:
export COURSEFORGE_HARNESS=fake COURSEFORGE_FIXTURES=tests/fixtures/harness/demo
./courseforge new "Spotting Phishing Emails" --to releaseMore in Quick start.
CourseForge was specified, tested and directed by its author; the implementation is AI-assisted, built with Claude Code.
The rest of this document is the technical reference.
- Architecture · Pipeline · Features · Install · Quick start · Example: importing existing work · Human review · Outputs · Troubleshooting · Limitations · Development
CourseForge is a local, repository-contained course-engineering harness. It takes a one-line concept, or an existing artifact from any production stage (research dossier, instructional design, storyboard, finished HTML course), and drives it through a deterministic eleven-stage pipeline to a tested, accessible, single-file HTML course. Claude Code or Codex does the semantic work (research, authoring, review, repair); TypeScript code decides everything else: which stage runs, which reviewers and tools are used, which renderer draws each figure, when a human must approve, and whether the course may be released.
It is not a prompt collection. Prompts are data files routed by code. Every agent call runs against a saved execution plan, returns schema-validated JSON, and is audited: writes outside the plan are rolled back, locked content cannot be changed, repair cycles are capped, and every decision is logged so a run can be explained and resumed.
flowchart TD
CLI["CLI<br/>src/cli"] --> ENV["Environment<br/>setup · doctor · smoke"]
CLI --> API["Pipeline API<br/>src/pipeline/api.ts"]
API --> REG["Course registry + ingestion<br/>course.yaml · state.json · artifacts.json<br/>src/artifacts · src/ingestion"]
API --> SM["State machine + stage loop<br/>src/pipeline"]
SM --> ROUTER["Deterministic router<br/>config/*.json → execution plan<br/>src/routing"]
ROUTER --> HAR["Harness adapters<br/>Claude Code · Codex · fake<br/>src/harness"]
HAR --> AG["Generators · reviewers ·<br/>adjudicator · repairer"]
SM --> DET["Validators · compilers ·<br/>graphics · HTML build<br/>src/renderer · src/graphics"]
SM --> QA["Browser QA<br/>Playwright + axe<br/>src/qa"]
SM --> REL["Release gate + reports<br/>src/release"]
AG --> ART[("Course folder<br/>artifacts · versions · logs · provenance")]
DET --> ART
QA --> ART
REL --> ART
Only src/harness/ knows that Claude Code or Codex exist, and only src/core/proc.ts spawns processes. Both
rules are enforced by a boundary test. See docs/architecture/overview.md.
flowchart LR
C[CONCEPT] --> RB[RESEARCH_BRIEF] --> RD[RESEARCH_DOSSIER] --> ID[INSTRUCTIONAL_DESIGN]
ID --> SB[STORYBOARD] --> ED[EDITORIAL] --> VD[VISUAL_DIRECTION]
VD --> CM[COURSE_MODEL] --> CB[COURSE_BUILD] --> CQ[COURSE_QA] --> R[RELEASE]
| Stage | Work | Review / gate |
|---|---|---|
| CONCEPT | Agent normalises the concept, classifies risk tier | 1 reviewer |
| RESEARCH_BRIEF | Agent writes the brief | 3 reviewers, section validator |
| RESEARCH_DOSSIER | Agent researches per section (web tools); code mints claim IDs | 5 reviewers, citation/ID integrity |
| INSTRUCTIONAL_DESIGN | Agent writes design JSON; code renders Markdown | 5 reviewers, alignment validators |
| STORYBOARD | Agent writes one module per task; code merges | 6 reviewers, 9 validators, hybrid gate by default |
| EDITORIAL | Agent edits prose only | 2 reviewers, structural diff + polarity guard |
| VISUAL_DIRECTION | Agent picks bounded design enums; code expands tokens, checks contrast, routes visuals | 5 reviewers |
| COURSE_MODEL | Deterministic compile of course.json + trace graph |
validators only |
| COURSE_BUILD | Deterministic graphics + single-file HTML | build checks (single file, IDs, size, text equivalents) |
| COURSE_QA | Playwright functional QA, axe, screenshots, then a 10-reviewer panel | repair loop, regression report |
| RELEASE | Pure release gate, reports, manifest, licences | gate reasons; human gate on elevated/high-stakes courses |
Every agent stage runs the same loop: plan → generate → validators + reviewer panel → deterministic pre-adjudication → optional AI adjudication → repair plan → scoped repair → targeted re-review (at most 3 cycles by default) → human gate → lock. Details: state machine, review and repair.
- Deterministic routing and saved execution plans.
config/*.jsonregistries decide stages, reviewers, skills, tools, renderers and fallbacks. The plan is written tologs/execution-plans/and every decision tologs/routing-decisions.jsonlbefore any agent runs. Identical inputs produce identical plans. - Reviewer ensembles, adjudication and bounded repair. Findings are deduplicated and merged by code; an AI adjudicator runs only for contradictions, blockers or low-confidence findings. Repairs replace whole objects by stable ID, are checked against the plan, locks and schemas, and are rolled back on violation.
- Human gates: auto, hybrid, human. Per stage, per course, or per run, with risk-tier floors that high-stakes courses cannot silently lower.
- Any-stage ingestion. Import Markdown, text, HTML, JSON, DOCX or PDF at any stage in
preserve,review-only,improveorrebuildmode. Originals are kept byte-for-byte and re-verified at release. - Traceability. source → claim → learning objective → block → item → component → HTML element, queryable
with
courseforge traceand rendered into the DOM asdata-cf-*attributes. - Single-file accessible HTML. No external requests, hashed Content-Security-Policy, keyboard-first interactions, light/dark themes, reduced motion, progress persistence.
- Structured graphics. 22 visual archetypes routed to Mermaid, native SVG builders, SVG.js, Vega-Lite or D3, with a bounded fallback chain that always ends in a text equivalent.
- Real browser QA. Playwright drives every screen and interaction; axe-core serious/critical violations block release.
- Existing-HTML improvement. Arbitrary HTML courses are crawled, reviewed and rebuilt through CourseForge components, with a before/after regression report.
- Claude Code and Codex. Same pipeline, same schemas, same fixtures; the backend is a per-run choice.
- Node.js 22.12 or newer (24 LTS recommended) with npm
- Git (to clone; optional afterwards)
- Optional, for live generation: Claude Code (
claude) or Codex CLI (codex), installed and signed in - About 1 GB of disk for dependencies plus Playwright Chromium; 2 GB of free RAM or more is recommended
No global npm packages are installed and no user-level configuration is modified.
# Windows (PowerShell 5.1 or 7)
git clone <repo-url> CourseForge
cd CourseForge
.\setup.ps1# macOS / Linux (or Git Bash on Windows)
git clone <repo-url> CourseForge
cd CourseForge
./setup.shsetup.ps1 / setup.sh check the Node.js version (printing the exact install command if it is missing or too
old), then run scripts/bootstrap.mjs, which:
- records the environment (platform, RAM, cloud-synced folder advisory);
- runs
npm ci(skipped when the lockfile is unchanged; retried on Windows file locks); - installs Playwright Chromium into the shared per-user browser cache (never inside the repository);
- compiles TypeScript to
dist/; - runs
doctor --repairfor repository-local problems; - runs the smoke fixture: build a small course, render diagrams, launch Chromium, drive an interaction, run axe, take a screenshot.
It prints READY only when doctor and the smoke fixture pass. Re-running is idempotent. Flags: --no-smoke,
--offline, --ci, --json. See docs/user-guide/installation.md.
Run the CLI without a global install: .\courseforge <command> (Windows), ./courseforge <command>,
npm run courseforge -- <command>, or node bin/courseforge.mjs <command>.
A longer walkthrough is in docs/user-guide/quick-start.md; every command and flag is listed in the CLI reference.
The fake harness replays a committed, schema-valid fixture set for a micro-course, so the whole pipeline runs without a model:
$env:COURSEFORGE_HARNESS='fake'; $env:COURSEFORGE_FIXTURES='tests/fixtures/harness/demo'
.\courseforge new "Spotting Phishing Emails" --to releaseexport COURSEFORGE_HARNESS=fake COURSEFORGE_FIXTURES=tests/fixtures/harness/demo
./courseforge new "Spotting Phishing Emails" --to releaseThe run pauses at STORYBOARD (exit code 10) because that stage has a hybrid gate by default. Review the files
under courses/spotting-phishing-emails/storyboard/, then:
./courseforge gate approve --course spotting-phishing-emails --stage storyboard
./courseforge continue --course spotting-phishing-emailsFor one command from whatever you have (a description, documents, or a half-finished course) to a finished
course, use courseforge make "<what the course is about>" [files or folders] --review one-shot. It runs
unattended and stops once, before release, for your approval (see
docs/user-guide/human-review.md). Pass --gate auto to new to run without
pausing (standard-risk courses only). If QA finds blocking issues,
the run pauses at COURSE_QA with a consolidated review in review/consolidated-review.md.
./courseforge doctor # confirm claude and/or codex are installed and signed in
./courseforge new "Safe Ladder Use" --audience "warehouse staff" --duration 30 --to storyboard
./courseforge status --course safe-ladder-use
./courseforge run --course safe-ladder-use --to release --backend codexBackend selection: --backend flag, else pipeline.agent_backend in course.yaml, else auto (first
available and signed-in backend in config/fallbacks.json order). Live runs cost model tokens and take time;
see Limitations.
Start from any stage. The example lineage in examples/chemical-risk/ is the author's own
portfolio work, used as a regression fixture:
# a storyboard: review, adjudicate and repair it, then continue to release
./courseforge ingest examples/chemical-risk/04_storyboard.md --course chem-demo --stage storyboard --mode improve
./courseforge run --course chem-demo --to release
# a finished HTML course: reconstruct the model, crawl, review, rebuild through CourseForge components
./courseforge ingest examples/chemical-risk/06_interactive_course.html --course chem-html --stage course_build --mode improve
./courseforge run --course chem-html --to release --gate autoThe stage is inferred when --stage is omitted: deterministic heuristics first, then (only if they are not
decisive) a closed-enum agent classifier; if confidence is still low the import stops with a clear message
unless --conservative (review-only) is given. Every import writes input/intake-report.json (inferred stage and
evidence, contract gaps, IDs found, warnings, next legal targets). See
docs/user-guide/ingestion.md.
| Mode | Generate | Review | Repair |
|---|---|---|---|
preserve |
no | no (validators only) | no |
review-only |
no | yes | no |
improve |
no | yes | yes |
rebuild |
yes (import is source material) | yes | yes |
./courseforge status --course <id> # stage table, gates, open findings, next action
./courseforge findings list --course <id> --stage storyboard
./courseforge findings accept --course <id> --ids SB-C0-003,SB-C0-007
./courseforge findings reject --course <id> --ids SB-C0-004
./courseforge gate lock --course <id> --stage storyboard --ids M2-B03 # protect a block from any repair
./courseforge gate reject --course <id> --stage storyboard --instructions "Shorten module 2 scenarios"
./courseforge gate approve --course <id> --stage storyboard
./courseforge continue --course <id>Gate modes: auto locks when no blocking findings remain; hybrid runs the AI repair loop, then waits for
approval; human waits for approval without AI repair. A paused run exits with code 10. See
docs/user-guide/human-review.md.
courses/<course-id>/
├─ course.yaml # course manifest (risk tier, backend, human_review gates)
├─ state.json # per-stage status, gates, cycles, canonical artifact pointers
├─ artifacts.json # registry: every artifact version with hash, producer, parents, locks
├─ input/ # concept, originals/ (read-only), intake-report.json
├─ research/ # research-brief, research-dossier (.json + .md), sources.jsonl, claims.jsonl
├─ design/ # instructional-design (.json + .md)
├─ storyboard/ # storyboard, storyboard-edited, editorial-diff.json
├─ visual/ # direction, design tokens, component plan, visual specs
├─ model/ # course.json, trace.json, build-manifest.json
├─ build/ # index.html, build-report.json
├─ review/ # functional-tests, accessibility-review, screenshots/, findings/, repair plan, regression
├─ release/ # course.html, qa-report.md, source-report.md, release-manifest.json, licenses/
├─ versions/ # immutable snapshots (<artifact>-v<N>-<event>/)
└─ logs/ # execution-plans/, routing-decisions.jsonl, run-events.jsonl, tasks/
Each stage keeps its reviewer outputs, findings and repair plans under <stage-dir>/review/<stage>/c<cycle>/.
COURSE_QA runs a contract-driven Playwright pass over CourseForge builds (every screen, every interaction with the model's answer key and a mutated wrong answer, scoring, navigation, glossary, references, progress and reset, offline behaviour, console errors, overflow, keyboard focus, axe per screen, screenshots at 1440/768/390) or a heuristic crawler for arbitrary imported HTML. The release gate is a pure function that blocks on:
OPEN_BLOCKING_FINDING · AXE_BLOCKING_VIOLATION · FUNCTIONAL_FAILURE · MISSING_ARTIFACT ·
CITATION_INTEGRITY · LOCK_CONFLICT · CYCLES_EXHAUSTED · ORIGINAL_MODIFIED · BUILD_CHECK_FAILED ·
HUMAN_APPROVAL_REQUIRED
| Claude Code | Codex CLI | |
|---|---|---|
| Invocation | claude -p --output-format stream-json, prompt on stdin |
codex exec --json … -, prompt on stdin |
| Structured output | --json-schema (inline) |
--output-schema <file> + -o <file> |
| Isolation from user config | --safe-mode, --setting-sources project, --strict-mcp-config |
--ignore-user-config, --ignore-rules, --ephemeral |
| Permissions | --permission-mode dontAsk, tool allowlist, Edit(...) allow rules for writable paths |
-s read-only or workspace-write, approvals never |
| Known gap | none known | the global ~/.codex/AGENTS.md cannot be disabled; a role preamble tells the agent to ignore it |
Flags are probed from --help and used only when present. In both cases the hard guarantee is CourseForge's own
write audit. See docs/architecture/harness.md.
| Symptom | Fix |
|---|---|
| Anything looks wrong | ./courseforge doctor (add --repair to fix repository-local issues, --json for machine output) |
EPERM / EBUSY during install or runs |
The repo is in a OneDrive/Dropbox/iCloud folder with sync active. Pause sync or clone elsewhere |
pw.chromium fails |
npx playwright install chromium (Linux: npx playwright install --with-deps chromium) |
| "running scripts is disabled" on Windows | Set-ExecutionPolicy -Scope Process Bypass, then .\setup.ps1; ZIP downloads: Get-ChildItem -Recurse *.ps1 | Unblock-File |
| Slow QA or browser crashes | Close other browsers; doctor warns below 2 GB free RAM |
No agent backend available |
Install and sign in: run claude once, or codex login; or use the fake harness |
| Codex output influenced by personal instructions | Your ~/.codex/AGENTS.md is always loaded by Codex; isolation is partial (see above) |
More in docs/user-guide/troubleshooting.md.
- Automated accessibility checks (axe, keyboard and focus checks) are not a formal WCAG conformance audit.
- Live generation of a full course takes a long time and costs model tokens; tests and the demo use fixtures.
- PPTX import is not supported (DOCX and PDF are, via text extraction).
- Result tracking is optional (docs/user-guide/tracking.md): SCORM 1.2 for a training system (LMS), a Google Sheet, or CourseForge's own results dashboard. xAPI is not supported. Scores are computed in the learner's browser, so they suit training records, not high-stakes exams.
- AI-generated imagery is not included; figures are structured diagrams, charts and icons.
- Codex isolation is partial: the user-global
AGENTS.mdcannot be switched off. - No pixel-baseline visual regression in v1; screenshots are evidence for reviewers, not golden images.
- Of the four AI classifiers (import stage, visual archetype, claim category, risk tier), only the import-stage classifier is wired in (as a fallback when heuristics are not decisive). Visual archetypes and risk tier are set by the storyboard/visual-direction and concept agents; claim categories by the research agent.
- Backend fallback happens at most once per run and only for configured failure classes (off by default).
- Crash recovery restarts an interrupted stage, reusing its generated outputs; review cycles are rerun.
| Script | Purpose |
|---|---|
npm run build |
Compile TypeScript to dist/ |
npm run typecheck |
tsc --noEmit |
npm run lint / npm run format |
Biome check / format |
npm test |
Unit + integration tests (Vitest, fake harness, no network) |
npm run test:e2e |
Playwright end-to-end tests |
npm run test:acceptance |
Acceptance matrix: spec-10 requirement IDs (A1–L7) → tests |
npm run gen:schemas |
Regenerate schemas/*.schema.json from the zod sources |
npm run doctor / npm run smoke |
Environment check / smoke fixture |
npm run smoke:clean |
Fresh-clone setup test in a temp directory |
npm run coverage |
Coverage report |
npm run ci |
typecheck + lint + test + e2e |
Read CONTRIBUTING.md and the developer guide. Design decisions are recorded as ADRs. Security: SECURITY.md.
MIT, see LICENSE. Third-party components and notices: THIRD-PARTY.md.
