Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CourseForge

CourseForge turns a course idea, or the material you already have, into a tested, accessible e-learning course you can share as a single file, with every fact traced back to its source.

A screen from a course built by CourseForge: module navigation with per-screen progress, a titled content screen, and cards listing the module's objectives and contents.

The problem

Building a one-hour course by hand takes weeks. Someone researches the subject, writes learning objectives, storyboards every screen, writes the assessment, builds it in an authoring tool, then tests that it works and that people using a keyboard or a screen reader can complete it.

General-purpose AI tools draft that material in minutes, and the draft cannot be trusted. Facts arrive without sources. Nothing is reviewed except by the person who asked for it. Accessibility is whatever the model happened to produce, and there is no record of what was checked, what was changed, or who approved it. That may be tolerable for onboarding trivia, but safety, compliance, clinical and financial training have to stand up to scrutiny later.

CourseForge sits between those two. The agent does the reading and writing; the system decides what must be checked, who checks it, when a person has to approve, and whether the result is allowed to ship.

What it does

Give it any of these:

  • a one-line description of the course you want;
  • a folder of your own material (.md, .txt, .docx, .pdf, .html, .json);
  • work already in progress at any stage: a research dossier, a design blueprint, a storyboard, or a finished HTML course you want reviewed or rebuilt.

It runs an eleven-stage production line: research brief, sourced research dossier, instructional design, storyboard, editorial pass, visual direction, build, browser testing, release. What comes back is one self-contained HTML file, along with a QA report, a source report and a full record of how the course was made.

Why it is different

  • Every claim carries its source. Research produces a dossier where each substantive statement gets a claim ID tied to a source, and those IDs follow the claim through design, storyboard and the finished screen. A course cannot be released if its citations do not hold together.
  • Thirty-six specialist reviewers. They cover accuracy, objective alignment, assessment quality, reading level, accessibility, learner experience and visual design. Each one has a written rubric and a defined remit, and returns structured findings rather than a general opinion on the draft.
  • Repairs are bounded and audited. Findings are deduplicated by code, adjudicated where reviewers disagree, then repaired by stable ID. A repair that touches anything outside its plan is rolled back. Repair cycles are capped, so a course cannot spin in review forever.
  • You decide where a human approves. Run it with no gates, with the recommended gates, or stop at every stage. Higher-risk subjects raise the minimum automatically, and a course cannot quietly lower it.
  • Accessibility decides whether the course ships. Every interaction is keyboard-operable, every figure has a text equivalent, and a real browser drives every screen with axe-core. Serious violations block release.
  • Your own work is preserved. Anything you import is kept byte-for-byte and re-verified at release, so a rebuild cannot silently rewrite the source you gave it.
  • No lock-in. The deliverable is one HTML file with no external requests. Email it, put it on a shared drive, or host it anywhere. A SCORM 1.2 package and simple completion tracking are available if you want them.
  • No API keys. It drives the Claude Code or Codex CLI you are already signed in to.

What a finished course contains

Sixteen content block types (concepts, examples, comparisons, processes, evidence, warnings, misconceptions, scenarios and more) and six interaction types: single and multiple choice, matching, categorization, sequencing and reveal. Twenty-two diagram and chart archetypes rendered as real SVG rather than pictures of text. Six visual design families with light and dark themes, each contrast-checked. Knowledge checks throughout, a graded assessment with a pass mark, a glossary and a references list.

Who it is for

  • Learning and development teams who need more courses than they have developer time for.
  • Compliance, safety and onboarding training, where traceable sources and accessibility are requirements.
  • Subject-matter experts who have the knowledge, and the documents, but no course developer.
  • Consultancies and agencies producing training for clients, who need an audit trail of what was checked.
  • Anyone modernizing an old course: improve mode reviews existing HTML and rebuilds it, with a before-and-after report of what changed.

How it works, in sixty seconds

Code owns the process; the agent owns the words. Registries in config/ decide which stage runs, which reviewers sit on it, which tools they may use, which renderer draws each figure, when a person must approve, and whether the course may be released. Prompts are data files, not code. Every agent call runs against an execution plan written down in advance, returns output checked against a schema, and is logged, so any run can be explained, resumed or repeated. See the pipeline below.

What it is not, yet

This is version 0.1.0, with the gaps that implies:

  • No video, no audio or narration, and no AI-generated imagery. Figures are diagrams, charts and icons.
  • Navigation is linear. There are scenario screens, but no branching paths.
  • The accessibility checks are automated. They are not a formal WCAG conformance audit.
  • SCORM export has not yet been verified against a live learning management system.
  • A full live run takes hours and costs model tokens. The tests and the demo run offline on fixtures.
  • Scoring happens in the learner's browser, which suits training records rather than high-stakes exams.

The full list is under Limitations.

Try it

git clone <repo-url> CourseForge
cd CourseForge
./setup.sh          # or .\setup.ps1 on Windows

Then run the offline demo, which replays committed fixtures from tests/fixtures/harness/demo and needs no agent and no network:

export COURSEFORGE_HARNESS=fake COURSEFORGE_FIXTURES=tests/fixtures/harness/demo
./courseforge new "Spotting Phishing Emails" --to release

More in Quick start.

CourseForge was specified, tested and directed by its author; the implementation is AI-assisted, built with Claude Code.


The rest of this document is the technical reference.

CourseForge is a local, repository-contained course-engineering harness. It takes a one-line concept, or an existing artifact from any production stage (research dossier, instructional design, storyboard, finished HTML course), and drives it through a deterministic eleven-stage pipeline to a tested, accessible, single-file HTML course. Claude Code or Codex does the semantic work (research, authoring, review, repair); TypeScript code decides everything else: which stage runs, which reviewers and tools are used, which renderer draws each figure, when a human must approve, and whether the course may be released.

It is not a prompt collection. Prompts are data files routed by code. Every agent call runs against a saved execution plan, returns schema-validated JSON, and is audited: writes outside the plan are rolled back, locked content cannot be changed, repair cycles are capped, and every decision is logged so a run can be explained and resumed.

Architecture

flowchart TD
  CLI["CLI<br/>src/cli"] --> ENV["Environment<br/>setup · doctor · smoke"]
  CLI --> API["Pipeline API<br/>src/pipeline/api.ts"]
  API --> REG["Course registry + ingestion<br/>course.yaml · state.json · artifacts.json<br/>src/artifacts · src/ingestion"]
  API --> SM["State machine + stage loop<br/>src/pipeline"]
  SM --> ROUTER["Deterministic router<br/>config/*.json → execution plan<br/>src/routing"]
  ROUTER --> HAR["Harness adapters<br/>Claude Code · Codex · fake<br/>src/harness"]
  HAR --> AG["Generators · reviewers ·<br/>adjudicator · repairer"]
  SM --> DET["Validators · compilers ·<br/>graphics · HTML build<br/>src/renderer · src/graphics"]
  SM --> QA["Browser QA<br/>Playwright + axe<br/>src/qa"]
  SM --> REL["Release gate + reports<br/>src/release"]
  AG --> ART[("Course folder<br/>artifacts · versions · logs · provenance")]
  DET --> ART
  QA --> ART
  REL --> ART
Loading

Only src/harness/ knows that Claude Code or Codex exist, and only src/core/proc.ts spawns processes. Both rules are enforced by a boundary test. See docs/architecture/overview.md.

Pipeline

flowchart LR
  C[CONCEPT] --> RB[RESEARCH_BRIEF] --> RD[RESEARCH_DOSSIER] --> ID[INSTRUCTIONAL_DESIGN]
  ID --> SB[STORYBOARD] --> ED[EDITORIAL] --> VD[VISUAL_DIRECTION]
  VD --> CM[COURSE_MODEL] --> CB[COURSE_BUILD] --> CQ[COURSE_QA] --> R[RELEASE]
Loading
Stage Work Review / gate
CONCEPT Agent normalises the concept, classifies risk tier 1 reviewer
RESEARCH_BRIEF Agent writes the brief 3 reviewers, section validator
RESEARCH_DOSSIER Agent researches per section (web tools); code mints claim IDs 5 reviewers, citation/ID integrity
INSTRUCTIONAL_DESIGN Agent writes design JSON; code renders Markdown 5 reviewers, alignment validators
STORYBOARD Agent writes one module per task; code merges 6 reviewers, 9 validators, hybrid gate by default
EDITORIAL Agent edits prose only 2 reviewers, structural diff + polarity guard
VISUAL_DIRECTION Agent picks bounded design enums; code expands tokens, checks contrast, routes visuals 5 reviewers
COURSE_MODEL Deterministic compile of course.json + trace graph validators only
COURSE_BUILD Deterministic graphics + single-file HTML build checks (single file, IDs, size, text equivalents)
COURSE_QA Playwright functional QA, axe, screenshots, then a 10-reviewer panel repair loop, regression report
RELEASE Pure release gate, reports, manifest, licences gate reasons; human gate on elevated/high-stakes courses

Every agent stage runs the same loop: plan → generate → validators + reviewer panel → deterministic pre-adjudication → optional AI adjudication → repair plan → scoped repair → targeted re-review (at most 3 cycles by default) → human gate → lock. Details: state machine, review and repair.

Key features

  • Deterministic routing and saved execution plans. config/*.json registries decide stages, reviewers, skills, tools, renderers and fallbacks. The plan is written to logs/execution-plans/ and every decision to logs/routing-decisions.jsonl before any agent runs. Identical inputs produce identical plans.
  • Reviewer ensembles, adjudication and bounded repair. Findings are deduplicated and merged by code; an AI adjudicator runs only for contradictions, blockers or low-confidence findings. Repairs replace whole objects by stable ID, are checked against the plan, locks and schemas, and are rolled back on violation.
  • Human gates: auto, hybrid, human. Per stage, per course, or per run, with risk-tier floors that high-stakes courses cannot silently lower.
  • Any-stage ingestion. Import Markdown, text, HTML, JSON, DOCX or PDF at any stage in preserve, review-only, improve or rebuild mode. Originals are kept byte-for-byte and re-verified at release.
  • Traceability. source → claim → learning objective → block → item → component → HTML element, queryable with courseforge trace and rendered into the DOM as data-cf-* attributes.
  • Single-file accessible HTML. No external requests, hashed Content-Security-Policy, keyboard-first interactions, light/dark themes, reduced motion, progress persistence.
  • Structured graphics. 22 visual archetypes routed to Mermaid, native SVG builders, SVG.js, Vega-Lite or D3, with a bounded fallback chain that always ends in a text equivalent.
  • Real browser QA. Playwright drives every screen and interaction; axe-core serious/critical violations block release.
  • Existing-HTML improvement. Arbitrary HTML courses are crawled, reviewed and rebuilt through CourseForge components, with a before/after regression report.
  • Claude Code and Codex. Same pipeline, same schemas, same fixtures; the backend is a per-run choice.

Requirements

  • Node.js 22.12 or newer (24 LTS recommended) with npm
  • Git (to clone; optional afterwards)
  • Optional, for live generation: Claude Code (claude) or Codex CLI (codex), installed and signed in
  • About 1 GB of disk for dependencies plus Playwright Chromium; 2 GB of free RAM or more is recommended

No global npm packages are installed and no user-level configuration is modified.

Install

# Windows (PowerShell 5.1 or 7)
git clone <repo-url> CourseForge
cd CourseForge
.\setup.ps1
# macOS / Linux (or Git Bash on Windows)
git clone <repo-url> CourseForge
cd CourseForge
./setup.sh

setup.ps1 / setup.sh check the Node.js version (printing the exact install command if it is missing or too old), then run scripts/bootstrap.mjs, which:

  1. records the environment (platform, RAM, cloud-synced folder advisory);
  2. runs npm ci (skipped when the lockfile is unchanged; retried on Windows file locks);
  3. installs Playwright Chromium into the shared per-user browser cache (never inside the repository);
  4. compiles TypeScript to dist/;
  5. runs doctor --repair for repository-local problems;
  6. runs the smoke fixture: build a small course, render diagrams, launch Chromium, drive an interaction, run axe, take a screenshot.

It prints READY only when doctor and the smoke fixture pass. Re-running is idempotent. Flags: --no-smoke, --offline, --ci, --json. See docs/user-guide/installation.md.

Run the CLI without a global install: .\courseforge <command> (Windows), ./courseforge <command>, npm run courseforge -- <command>, or node bin/courseforge.mjs <command>.

Quick start

A longer walkthrough is in docs/user-guide/quick-start.md; every command and flag is listed in the CLI reference.

Offline demo (no agent, no network)

The fake harness replays a committed, schema-valid fixture set for a micro-course, so the whole pipeline runs without a model:

$env:COURSEFORGE_HARNESS='fake'; $env:COURSEFORGE_FIXTURES='tests/fixtures/harness/demo'
.\courseforge new "Spotting Phishing Emails" --to release
export COURSEFORGE_HARNESS=fake COURSEFORGE_FIXTURES=tests/fixtures/harness/demo
./courseforge new "Spotting Phishing Emails" --to release

The run pauses at STORYBOARD (exit code 10) because that stage has a hybrid gate by default. Review the files under courses/spotting-phishing-emails/storyboard/, then:

./courseforge gate approve --course spotting-phishing-emails --stage storyboard
./courseforge continue --course spotting-phishing-emails

For one command from whatever you have (a description, documents, or a half-finished course) to a finished course, use courseforge make "<what the course is about>" [files or folders] --review one-shot. It runs unattended and stops once, before release, for your approval (see docs/user-guide/human-review.md). Pass --gate auto to new to run without pausing (standard-risk courses only). If QA finds blocking issues, the run pauses at COURSE_QA with a consolidated review in review/consolidated-review.md.

Real run

./courseforge doctor                       # confirm claude and/or codex are installed and signed in
./courseforge new "Safe Ladder Use" --audience "warehouse staff" --duration 30 --to storyboard
./courseforge status --course safe-ladder-use
./courseforge run --course safe-ladder-use --to release --backend codex

Backend selection: --backend flag, else pipeline.agent_backend in course.yaml, else auto (first available and signed-in backend in config/fallbacks.json order). Live runs cost model tokens and take time; see Limitations.

Example: importing existing work

Start from any stage. The example lineage in examples/chemical-risk/ is the author's own portfolio work, used as a regression fixture:

# a storyboard: review, adjudicate and repair it, then continue to release
./courseforge ingest examples/chemical-risk/04_storyboard.md --course chem-demo --stage storyboard --mode improve
./courseforge run --course chem-demo --to release

# a finished HTML course: reconstruct the model, crawl, review, rebuild through CourseForge components
./courseforge ingest examples/chemical-risk/06_interactive_course.html --course chem-html --stage course_build --mode improve
./courseforge run --course chem-html --to release --gate auto

The stage is inferred when --stage is omitted: deterministic heuristics first, then (only if they are not decisive) a closed-enum agent classifier; if confidence is still low the import stops with a clear message unless --conservative (review-only) is given. Every import writes input/intake-report.json (inferred stage and evidence, contract gaps, IDs found, warnings, next legal targets). See docs/user-guide/ingestion.md.

Mode Generate Review Repair
preserve no no (validators only) no
review-only no yes no
improve no yes yes
rebuild yes (import is source material) yes yes

Human review workflow

./courseforge status --course <id>                          # stage table, gates, open findings, next action
./courseforge findings list --course <id> --stage storyboard
./courseforge findings accept --course <id> --ids SB-C0-003,SB-C0-007
./courseforge findings reject --course <id> --ids SB-C0-004
./courseforge gate lock --course <id> --stage storyboard --ids M2-B03   # protect a block from any repair
./courseforge gate reject --course <id> --stage storyboard --instructions "Shorten module 2 scenarios"
./courseforge gate approve --course <id> --stage storyboard
./courseforge continue --course <id>

Gate modes: auto locks when no blocking findings remain; hybrid runs the AI repair loop, then waits for approval; human waits for approval without AI repair. A paused run exits with code 10. See docs/user-guide/human-review.md.

Outputs

courses/<course-id>/
├─ course.yaml              # course manifest (risk tier, backend, human_review gates)
├─ state.json               # per-stage status, gates, cycles, canonical artifact pointers
├─ artifacts.json           # registry: every artifact version with hash, producer, parents, locks
├─ input/                   # concept, originals/ (read-only), intake-report.json
├─ research/                # research-brief, research-dossier (.json + .md), sources.jsonl, claims.jsonl
├─ design/                  # instructional-design (.json + .md)
├─ storyboard/              # storyboard, storyboard-edited, editorial-diff.json
├─ visual/                  # direction, design tokens, component plan, visual specs
├─ model/                   # course.json, trace.json, build-manifest.json
├─ build/                   # index.html, build-report.json
├─ review/                  # functional-tests, accessibility-review, screenshots/, findings/, repair plan, regression
├─ release/                 # course.html, qa-report.md, source-report.md, release-manifest.json, licenses/
├─ versions/                # immutable snapshots (<artifact>-v<N>-<event>/)
└─ logs/                    # execution-plans/, routing-decisions.jsonl, run-events.jsonl, tasks/

Each stage keeps its reviewer outputs, findings and repair plans under <stage-dir>/review/<stage>/c<cycle>/.

QA and release gates

COURSE_QA runs a contract-driven Playwright pass over CourseForge builds (every screen, every interaction with the model's answer key and a mutated wrong answer, scoring, navigation, glossary, references, progress and reset, offline behaviour, console errors, overflow, keyboard focus, axe per screen, screenshots at 1440/768/390) or a heuristic crawler for arbitrary imported HTML. The release gate is a pure function that blocks on:

OPEN_BLOCKING_FINDING · AXE_BLOCKING_VIOLATION · FUNCTIONAL_FAILURE · MISSING_ARTIFACT · CITATION_INTEGRITY · LOCK_CONFLICT · CYCLES_EXHAUSTED · ORIGINAL_MODIFIED · BUILD_CHECK_FAILED · HUMAN_APPROVAL_REQUIRED

See docs/architecture/qa.md.

Claude Code and Codex

Claude Code Codex CLI
Invocation claude -p --output-format stream-json, prompt on stdin codex exec --json … -, prompt on stdin
Structured output --json-schema (inline) --output-schema <file> + -o <file>
Isolation from user config --safe-mode, --setting-sources project, --strict-mcp-config --ignore-user-config, --ignore-rules, --ephemeral
Permissions --permission-mode dontAsk, tool allowlist, Edit(...) allow rules for writable paths -s read-only or workspace-write, approvals never
Known gap none known the global ~/.codex/AGENTS.md cannot be disabled; a role preamble tells the agent to ignore it

Flags are probed from --help and used only when present. In both cases the hard guarantee is CourseForge's own write audit. See docs/architecture/harness.md.

Troubleshooting

Symptom Fix
Anything looks wrong ./courseforge doctor (add --repair to fix repository-local issues, --json for machine output)
EPERM / EBUSY during install or runs The repo is in a OneDrive/Dropbox/iCloud folder with sync active. Pause sync or clone elsewhere
pw.chromium fails npx playwright install chromium (Linux: npx playwright install --with-deps chromium)
"running scripts is disabled" on Windows Set-ExecutionPolicy -Scope Process Bypass, then .\setup.ps1; ZIP downloads: Get-ChildItem -Recurse *.ps1 | Unblock-File
Slow QA or browser crashes Close other browsers; doctor warns below 2 GB free RAM
No agent backend available Install and sign in: run claude once, or codex login; or use the fake harness
Codex output influenced by personal instructions Your ~/.codex/AGENTS.md is always loaded by Codex; isolation is partial (see above)

More in docs/user-guide/troubleshooting.md.

Limitations

  • Automated accessibility checks (axe, keyboard and focus checks) are not a formal WCAG conformance audit.
  • Live generation of a full course takes a long time and costs model tokens; tests and the demo use fixtures.
  • PPTX import is not supported (DOCX and PDF are, via text extraction).
  • Result tracking is optional (docs/user-guide/tracking.md): SCORM 1.2 for a training system (LMS), a Google Sheet, or CourseForge's own results dashboard. xAPI is not supported. Scores are computed in the learner's browser, so they suit training records, not high-stakes exams.
  • AI-generated imagery is not included; figures are structured diagrams, charts and icons.
  • Codex isolation is partial: the user-global AGENTS.md cannot be switched off.
  • No pixel-baseline visual regression in v1; screenshots are evidence for reviewers, not golden images.
  • Of the four AI classifiers (import stage, visual archetype, claim category, risk tier), only the import-stage classifier is wired in (as a fallback when heuristics are not decisive). Visual archetypes and risk tier are set by the storyboard/visual-direction and concept agents; claim categories by the research agent.
  • Backend fallback happens at most once per run and only for configured failure classes (off by default).
  • Crash recovery restarts an interrupted stage, reusing its generated outputs; review cycles are rerun.

Development

Script Purpose
npm run build Compile TypeScript to dist/
npm run typecheck tsc --noEmit
npm run lint / npm run format Biome check / format
npm test Unit + integration tests (Vitest, fake harness, no network)
npm run test:e2e Playwright end-to-end tests
npm run test:acceptance Acceptance matrix: spec-10 requirement IDs (A1–L7) → tests
npm run gen:schemas Regenerate schemas/*.schema.json from the zod sources
npm run doctor / npm run smoke Environment check / smoke fixture
npm run smoke:clean Fresh-clone setup test in a temp directory
npm run coverage Coverage report
npm run ci typecheck + lint + test + e2e

Read CONTRIBUTING.md and the developer guide. Design decisions are recorded as ADRs. Security: SECURITY.md.

License

MIT, see LICENSE. Third-party components and notices: THIRD-PARTY.md.

About

Agent harness that turns a course idea or existing material into a tested, accessible, single-file e-learning course, with sourced claims, specialist AI reviewers and human approval gates.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages