Skip to content

Add scorm-export: wrap a DoenetML activity as an LMS-ready SCORM 2004 package - #3019

Draft
cqnykamp wants to merge 19 commits into
Doenet:mainfrom
cqnykamp:feat-scorm-export
Draft

cqnykamp wants to merge 19 commits into
Doenet:mainfrom
cqnykamp:feat-scorm-export

Conversation

@cqnykamp

Copy link
Copy Markdown
Contributor

What

A new packages/scorm-export package that wraps a single DoenetML activity in an LMS-ready SCORM 2004 4th Edition zip — no PreTeXt toolchain, no book chrome. This is the groundwork for a "Download SCORM" button on doenet.org: an instructor exports one activity, uploads the zip to their LMS, and student scores and work persist in the LMS itself (no phone-home to doenet.org).

How it works

A package is six static files in a flat zip:

  • imsmanifest.xml — minimal SCORM 2004 4th Ed. manifest (one item, one SCO)
  • index.html — chrome-free shell hosting the activity iframe
  • activity.html — loads @doenet/standalone from CDN, embeds the DoenetML, speaks SPLICE to the parent
  • ptx_scorm_events.js — PreTeXt's SCORM bridge (vendored; LMS API discovery, scoring, state save/restore, submit)
  • lti_iframe_resizer.js — PreTeXt's SPLICE lti.frameResize handler (vendored, verbatim)
  • lz-string.min.js — pinned npm dep, copied in at build time; compresses state for suspend_data

build.mjs is template-substitution + zip (Node built-ins + the zip CLI). The design is deliberately client-side-able: the same substitute-and-zip could run in the browser behind a button.

node packages/scorm-export/build.mjs sample/sample.doenet --title "My Activity"

Reusing the PreTeXt bridge

The SCORM runtime intelligence is reused from PreTeXt (js/ptx_scorm_events.js, js/lti_iframe_resizer.js) rather than reimplemented, and kept as verbatim vendored copies (see vendor/VENDORED.md) so upstream fixes re-sync cleanly. The bridge is ~96% byte-identical to upstream; the local changes (all marked VENDOR-MOD) do one thing: persist the Doenet activity state through the SCORM data model instead of localStorage-only, so student work restores on a fresh LMS launch.

  • _doenetStates is compressed (lz-string, base64) into cmi.suspend_data; the manifest declares 4th Edition for its 64,000-char suspend_data cap. A size guard drops the blob (falling back to localStorage) rather than corrupting suspend_data if it would ever overflow.
  • Upstream only saves the Doenet state blob when Runestone is present (gated on RunestoneBase.__ptxScormHooked); in a Runestone-free standalone package the state was never captured, so the SPLICE message handler now forwards state (and the real subject) into recordInteraction.

These bridge changes are a candidate to contribute upstream to PreTeXt; if accepted there, the local mod drops and we re-copy verbatim.

Verified

End-to-end on Canvas: state save → compress → persist in suspend_data → restore on relaunch → deliver back to Doenet. cmi.entry = "resume", score restored, and the compressed state (dz) round-trips.

Debugging

debug/size-probe.html is a passive diagnostic (state-blob size, suspend_data round-trip, lti.frameResize tracking). It is not in a normal package — pass --debug to inline it.

Notes / known limitations

  • The viewer loads @doenet/standalone from jsDelivr; pin --doenet-version for reproducible packages. A fully offline package would bundle the viewer into the zip.
  • Instructor per-learner review depends on the LMS offering a review-mode launch; Canvas's native SCORM player only surfaces the grade (documented in the README).
  • In Canvas's SCORM player the activity iframe resizes correctly, but the player's own outer frame is fixed-height and not SPLICE-aware, so a scrollbar there is expected (documented).

🤖 Generated with Claude Code

cqnykamp and others added 19 commits July 23, 2026 11:45
…ckage

New workspace package `@doenet-tools/scorm-export`: the basis for a
"Download SCORM" button on doenet.org. Given one DoenetML activity, it
produces an LMS-ready SCORM 2004 (4th Edition) zip with no PreTeXt
toolchain and no book chrome.

Contents:
- build.mjs: template substitution + zip packaging (Node built-ins only).
- templates/: chrome-free index.html shell, activity.html embedding the
  activity and loading @doenet/standalone from CDN, and a minimal SCORM
  2004 4th Edition manifest.
- vendor/: PreTeXt's SCORM bridge (ptx_scorm_events.js) and SPLICE resize
  handler (lti_iframe_resizer.js), plus lz-string; see vendor/VENDORED.md.

Local modifications to the vendored bridge (all marked VENDOR-MOD), a
candidate to contribute upstream to PreTeXt:
- Capture Doenet state in a Runestone-free package (upstream only saved it
  when RunestoneBase was present).
- Persist that state (compressed, LZ-string) into cmi.suspend_data so it
  restores across LMS launches, since localStorage does not survive them;
  4th Edition is declared for the 64,000-char suspend_data limit, with a
  size guard that degrades gracefully on overflow.

Verified end-to-end on Canvas: score + state save, persist, and restore.
The vendored files are Prettier-ignored to keep them verbatim. A separate
DocViewer state-restore race in @doenet/standalone is filed against the
DoenetML repo and is not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ring it

lz-string is an unmodified, published package, so npm is a better fit than
a committed minified blob: version + integrity now live in package.json and
the lockfile, and `npm audit`/Renovate can track it.

- Pin `lz-string@1.5.0` in package.json (exact — these bytes ship into
  student-facing packages, so bumps should be deliberate).
- build.mjs resolves the minified build from node_modules via
  import.meta.resolve and copies it into the package; the two PreTeXt files
  (locally modified) stay vendored.
- Remove vendor/lz-string.min.js; update VENDORED.md and README.

The packaged bytes are identical to the previously-vendored copy; runtime
behavior is unchanged. README also corrected for the 4th-Edition +
compressed-suspend_data persistence (was still describing the old
localStorage-only prototype).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Move the state-blob / suspend_data diagnostic out of the always-shipped
index.html template into debug/size-probe.html. build.mjs inlines it into
index.html only when run with --debug; a normal package contains no trace
of it (file count unchanged, since it is inlined rather than added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Extend debug/size-probe.html to log each lti.frameResize: the height
Doenet reported, the height the resizer applied to the activity iframe,
and whether index.html overflows the box the LMS gave it. The overflow
clause distinguishes an inner scrollbar (ours to fix) from the LMS
player's outer frame (not controllable from the SCO). The vendored
lti_iframe_resizer.js is left byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pure reformatting from the repo-wide prettier run; no semantic change.
The vendored files under vendor/ are excluded via .prettierignore and
remain byte-for-byte upstream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Makes the scorm-export package usable from the website. Building a package
is pure string substitution plus a zip, so it now runs client-side behind
the button rather than through a new API endpoint.

Package refactor:

- Move the builder into src/index.js with no fs, child_process, or DOM, so
  it runs unchanged in Node and the browser. Callers pass the six constant
  package files in as strings, which keeps vendor/ the single verbatim copy
  of the GPL sources instead of duplicating them into a generated module.
- build.mjs shrinks to a CLI adapter (same flags, same output); the web
  adapter loads the same files through Vite's ?raw imports.
- Zip with fflate instead of shelling out to zip(1), so both callers share
  one path and the build no longer needs the binary. Entries carry a fixed
  timestamp, so re-exporting an unchanged activity yields identical bytes.
- Inject the DoenetML into activity.html from a JS string literal with "<"
  escaped, rather than writing it inline into a raw-text element. A source
  containing "</script>" no longer has to be rejected.

Button:

- Shown only for single documents: a package is one SCO wrapping one source
  in one iframe, so problem sets and sequences would need a different
  manifest and shell.
- Passes the activity's contentId as the SCORM id, so renaming an activity
  no longer orphans the student's saved score and state, and pins the
  activity's own doenetmlVersion instead of tracking latest.

Verified by loading a built package in a browser: the injected source lands
in the applet container and the viewer renders it. Not yet tested against a
real LMS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Doenet state grows monotonically as a student works, so an activity that
crosses the 60,000-char suspend_data budget stays across it. Dropping the
compressed state blob was correct, but the caller then wrote the blob-less
payload over cmi.suspend_data — erasing the last snapshot that HAD fit, with
no way back. Measured against the real buildSuspendData(): a 20 KB state saved
fine, a 200 KB state then replaced it with nothing.

Remember the last blob known to fit (seeded on restore, updated on each
successful write) and re-attach it instead of clearing the field, so the
failure mode is "saved state stops updating" rather than "saved state is
destroyed". Grading fields stay current, and the gradebook value is unaffected
either way since it lives in cmi.score.*.

Also warn once per divId when saveDoenetState()'s localStorage write fails.
That catch was empty, hiding the second half of a total-loss case: localStorage
is the documented fallback once the blob is dropped, and an LMS runs the SCO as
a cross-site iframe where storage may be partitioned away or over quota.

Both are VENDOR-MOD changes to the vendored PreTeXt bridge, recorded in
vendor/VENDORED.md as candidates to contribute upstream.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The runtime looked like untestable integration glue, but the bridge only talks
to two interfaces: the SCORM API object and SPLICE postMessages from the
activity iframe. Faking both turns the state-size edge cases into ordinary
deterministic tests — we pick the exact size of the state blob instead of
hoping a real activity produces one. Each case builds a real package and boots
it in a fresh JSDOM, so the templates and vendored bridge under test are the
ones that ship.

Covers the builder (substitution, escaping, zip layout, determinism), package
coherence (manifest edition, every referenced file present), and the runtime:
session start, score and state writes, cross-launch restore, and the four
state-size failure paths — the SPM invariant across 1 KB to 1 MB of state,
growth past the budget, nothing-ever-fit, a truncating player, and a browser
refusing localStorage. Both recent vendor fixes were checked by mutation: each
regression test fails when its fix is removed, and only that test.

Two things the tests found:

- JSON.stringify emits U+2028/U+2029 raw, but JavaScript treats them as line
  terminators, so a source containing one would have broken the embedded
  literal. Escaped alongside "<".
- lz-string manages only ~1.2x on low-repetition state, so the 60,000-char
  budget is gone by ~72 KB of raw state — well inside the 10-100 KB range the
  bridge cites for Doenet state, and much lower than repetitive synthetic
  filler suggests. Recorded as an explicit test.

test/README.md records what this does not prove: real LMS behaviour, whether
real activities reach these sizes, the CDN viewer's message shape, and the
ways JSDOM is not a browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces the hand-written fake LMS with scorm-again (MIT), a real SCORM 2004
runtime — the same kind of API adapter LMSes and course authors ship. A mock
only tests the parts of the spec whoever wrote it remembered, and it agrees
with the content under test by construction. scorm-again enforces the actual
data model: vocabularies, read-only and undefined elements, value ranges, and
the string SPMs, returning the spec's error codes.

helpers/lms.js is now a thin wrapper adding only a call recorder and a
truncating-player variant, since scorm-again correctly REFUSES an over-SPM
suspend_data write (406, previous value stands) while some real players
silently truncate. Both behaviours are worth testing against; neither is worth
reimplementing.

It found a real deviation on the first run. The bridge keys each interaction
record by the exercise's div id, and a single-document package has exactly one,
so the second answer writes an id already used at index 0 — refused with 351,
and the dependent fields then fail with 408. The gradebook is unaffected (the
score lives in cmi.score.* and still gets through); what is lost on a strict
player is per-attempt interaction detail after the first answer. Left as-is
deliberately: changing the id scheme changes what the records mean, and it is a
question for upstream PreTeXt, whose multi-exercise pages do not hit it. A test
pins the behaviour so a fix is visible rather than silent.

Adds a conformance check that no write outside the interactions collection is
refused, at both ordinary and over-budget state sizes. Both vendor fixes were
re-verified by mutation against the new runtime.

test/README.md now also records the heavier options — the ADL conformance test
suite, SCORM Cloud's API, and a real LMS in Docker — and why each sits outside
PR CI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
build.mjs rolled its own argv loop, which accepted anything: an unknown flag
was silently stored and ignored, and a flag in last position ("--title" with
no value) set the option to undefined and fell through to a default. parseArgs
has been stable in node:util for years, the repo pins Node 24, and it costs no
dependency — it rejects both cases with a clear message and the usage line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
imsmanifest.xml was hand-written and only checked by regex for its edition
string. Now it is validated against ADL's own SCORM 2004 4th Edition XSDs with
libxml2 via xmllint-wasm — WASM, so no native build, no Java, and no binary on
PATH, which is what makes this viable in CI at all.

This checks the manifest against the spec's definition rather than against the
assertions we thought to write: element order, required attributes, identifiers
matching their declared type. It doubles as an end-to-end check of XML escaping,
since the activity title is interpolated into the document — a title containing
"&", "<", "]]>" or an XML declaration is not even well-formed if the escaping is
wrong, and all four are now tested.

Three deliberately broken manifests are asserted to fail. Without them a
validator that could not resolve its schemas would report everything as valid
and the rest of the file would prove nothing.

The 15 schema files are vendored in schemas/ with provenance and the licensing
situation recorded (IMS/ADL copyright, no accompanying licence text; published
for exactly this use and redistributed by every SCORM tool). They are added to
.prettierignore alongside vendor/ so they stay verbatim.

Noted as an open question in schemas/VENDORED.md: the manifest names these
schemas in xsi:schemaLocation but the zip does not ship them. Players resolve
the namespaces themselves so it has not caused trouble, but a content package
traditionally carries its control documents, and a strict validator may expect
them — ~88 KB per package if we decide to include them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Git normalized the XSDs' CRLF endings on commit, which defeats the point of
vendoring them verbatim: a future re-copy from upstream would show every line
as changed and hide the real diff. Mark them -text and restore the published
bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
build.mjs and the test helper each had their own copy of "read the templates,
the vendored bridge and lz-string off disk" — the helper's comment literally
said it did it "the same way build.mjs does", which is the sort of duplication
that drifts. Both now call loadNodeAssets().

It lives in src/node-assets.js rather than src/index.js because the browser
bundles index.js, which therefore has to stay free of fs. It uses createRequire
rather than import.meta.resolve so it works both under plain Node and under
Vite's SSR transform, which the tests run through.

Output is unchanged: a package built after this refactor is byte-identical to
one built before it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A Doenet activity has no single end-of-page submission. The viewer sends
SPLICE.reportScoreAndState after every answer carrying the whole document's
score — confirmed against @doenet/standalone, whose payload is exactly
{score, state, subject, activity_id, doc_id, message_id} with no completion
flag — and recordInteraction() writes and commits it immediately. The grade is
therefore already correct at all times.

Measured before removing it, the button contributed exactly two things beyond
what the SPLICE reports had already stored:

  after SPLICE report 0.6   scaled="0.6000"  completion="incomplete"
  after clicking Submit     scaled="0.6000"  completion="completed"  success="passed"

The score write it made was a duplicate of the value already there. Also note
_totalQuestions is 1 in a single-doc package (one [data-component="doenet"]),
so the bridge's per-question accounting is a passthrough of Doenet's own
aggregate score — all of PreTeXt's multi-exercise machinery is inert here.

So handlePageExit() now sets completion_status instead, gated on having
received at least one score (opening an activity and closing it is not
completing it) and keeping exit="suspend" so the attempt stays resumable.
Setting completion mid-session is what made Blackboard lock the attempt, and a
test pins that it stays "incomplete" while the student works.

Nothing writes success_status any more. The button set it to "passed"
unconditionally — a student who scored 0 was recorded as passed — which
contradicts the per-answer path's own comment that pass/fail belongs to the LMS
and its mastery score. submitSession() is now unreachable but keeps a defused
version of that write so re-enabling it cannot reintroduce the bug.

addSubmitButton() and submitSession() are left in place, uncalled, so an
upstream re-copy stays a small diff. This is the least likely of our vendor
mods to be wanted upstream: PreTeXt's multi-exercise pages are a case an
explicit submit genuinely fits.

Page exit now writes completion_status alongside exit="suspend", which adds an
element to the unload batch that Blackboard testing validated. Legal in SCORM
2004 (completion and exit are orthogonal) but unverified on a real Blackboard —
added to the manual pre-release checklist in test/README.md.

Also fixes the project's PostToolUse prettier hook, which reformatted the whole
vendored file twice this session: it ran with the shell's cwd, so from a
subdirectory prettier never found the repo-root .prettierignore. It now cds to
CLAUDE_PROJECT_DIR, falling back to the file's git root.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
imsmanifest.xml names five schemas in its xsi:schemaLocation by relative path,
but the zip did not contain them, so that attribute pointed at nothing and a
validator without network access could not check the manifest. A SCORM content
package traditionally carries its control documents at the root; now ours does.

All 15 go in, not just the five named: the rest are reached through those
files' imports and the ADL extension namespaces, so a partial set resolves to
nothing useful. A built package goes from 6 entries to 21, and from 47 KB to
62 KB — the 88 KB of XSD compresses well.

test/manifest.test.js now validates against the schemas as read out of the
built package rather than out of schemas/, so it covers the copies that
actually ship next to the manifest. A new test walks schemaLocation's
namespace/location pairs and asserts each location is present in the zip, so
the manifest cannot advertise a file the package omits.

On the browser side the schemas are inlined via ?raw like the other assets,
listed one by one in a separate module because import.meta.glob does not
reliably reach into a workspace package. That is ~88 KB of string on top of the
vendored bridge, so ActivityViewer now imports the exporter dynamically — the
whole SCORM payload became its own 218 KB (56 KB gzipped) chunk that only
downloads when someone actually clicks the button, instead of riding along in
the viewer chunk for every reader.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two directories of third-party files with the same "copied verbatim, never
edit, re-copy from upstream" rule were sitting under different conventions —
vendor/ for PreTeXt's GPL bridge, schemas/ for the ADL/IMS control documents,
each with its own VENDORED.md and its own .prettierignore entry.

Now vendor/ holds one directory per origin:

  vendor/pretext/   ptx_scorm_events.js, lti_iframe_resizer.js   (GPL)
  vendor/scorm/     15 SCORM 2004 4th Ed. schemas                (ADL/IMS)

with a vendor/VENDORED.md indexing them and stating the shared rule once, and
the per-origin VENDORED.md files keeping provenance, licensing and VENDOR-MOD
records next to what they describe.

Naming stays split along the two axes it always used: the directory says where
a file came from and what governs it, while the constants in src/index.js
(TEMPLATE_FILES / STATIC_FILES / SCHEMA_FILES) say what the builder does with
it. Those cross-cut — lz-string.min.js is a STATIC_FILE that is not vendored,
being read from node_modules — so collapsing them onto one name would not hold.

.prettierignore now names the vendored files rather than the whole tree, so
vendor/VENDORED.md is formatted like any other doc while the .js and .xsd stay
byte-for-byte. The .gitattributes -text rule that preserves the schemas' CRLF
endings moved with them; verified still binding at the new path.

All 22 files moved with git mv, so history follows them. Output is unchanged: a
package built after this is byte-identical to one built before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@cqnykamp
cqnykamp marked this pull request as draft September 3, 2026 14:35
@cqnykamp

cqnykamp commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

The main piece blocking this PR is the size limit on activity state for SCORM 2004. The size limit is quite tight, and Doenet activities I've tested reach it.

Since Doenet state grows (indefinitely?), we can't say for sure that a student's state will fit. If we merge this PR, we would be introducing a scenario where sometimes a student's state would save, and sometimes it wouldn't. That sounds like a major footgun we should avoid.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant