Read mdto spec guided-narration or the adjacent SPEC.md before authoring.
Outside a checkout, read the published spec.
The spec and fixtures define conformance; the editorial advice here adapts to the user's brief.
Write speech against the source
Set markdownto: guided-narration@0.1 and source to a relative Markdown path or
HTTPS URL. Read or capture the source through an available tool before choosing targets.
The source reference itself never instructs the renderer or validator to fetch a website.
For a third-party webpage, supply a capture or imported source in the reader; ownership
or permission to modify the original website is not required for this separate manuscript.
Each paragraph is one spoken beat of 1–1200 characters, with exactly one target field.
Use ## for chapter navigation; headings are not spoken. Write transitions in the speech.
Do not put lists, tables, HTML, blockquotes, or stage directions in the manuscript.
Do not invent timestamps, CSS selectors, block IDs, or automatic summarization instructions.
Prefer [target-quote:: exact passage] copied from the source. Matching normalizes
whitespace but remains case-sensitive. If the quote appears in multiple blocks, choose a
more distinctive quote or add [target-prefix:: previous block ending] and/or[target-suffix:: next block beginning]. These refer to neighboring blocks, not context
inside the matched block. Use [target:: #existing-id] only for a known captured block ID.
Keep the field in the spoken paragraph; a blank line starts another beat.
To explain an image, target its captured caption or alt text, or its existing captured
block ID. Say what to look at and explain what it means. Describe only details the source
actually shows. The embedded reader supports bundled demo illustrations and embedded raster
images; a remote image URL is not automatically fetched. Inspect the rendered capture to
confirm that the intended picture is present, not merely its caption.
Make the tour worth hearing
Choose an audience and a purpose from the brief. Explain relationships and consequences
rather than reading every visible sentence aloud. Keep each beat tied to one useful visual
stop. Give the listener enough context to understand a diagram without pretending a conceptual
illustration is historical evidence. Distinguish sourced facts from analogies and interpretation.
Use shorter paragraphs when a long explanation moves attention across several elements.
Resolve before delivery
Validate the manuscript with mdto validate tour.md. Read mdto guided-narration --help
for the current inspect and resolve commands. Resolution takes an explicit JSON capture:{ "source": "./article.md", "blocks": [{ "id": "intro", "text": "Exact source passage" }] }.
The capture's source string must match the manuscript exactly; IDs must be unique and nonempty.
Resolve every target against that capture and fix missing or ambiguous matches before playback.
The reader's Download capture for agents action exports its actual blocks for this purpose.
Record it, once the targets resolve
mdto guided-narration estimate counts the tour offline: characters, estimated duration, and
estimated cost per beat, plus what producing right now would actually buy against the cache.
Nothing is spent and no provider is called, so run it freely and read the incremental figure —
it is the one that answers "what does this edit cost".
mdto guided-narration produce records one track per beat. It prints the estimate every time and
stops until --yes acknowledges it, so there is no accidental spend; --dry-run shows the plan
and stops regardless. Choose the voice with --voice and the speed with --pace: the manuscript
carries neither, on purpose. Use --provider mock --voice mock-narrator to exercise the whole
path offline before spending anything, and --beat N or --beats 2-5 to re-record a passage you
rewrote. Identical beats are bought once, and an interrupted run resumes from its cache.
Output lands in guided-narration/ beside the manuscript, in the version-addressed shape the
agentsfs Hub reads: a joined MP3, a receipt, an audio index naming each beat's recording, the
per-beat files, and a pointer written last. The paths inside the pointer and receipt are
repository-relative. If the manuscript lives inside an embedded agentsfs, that directory's prefix
is stripped when it publishes, so pass --source-path <path-relative-to-the-agentsfs-root> —
otherwise the Hub refuses the pointer silently and simply shows no recording.
A host that renders the reader can hand it that recording instead of a speech service. Put arecording inside guided-restore's saved object — {version: 1, voice, audio: [{text,, one entry per beat, keyed by the beat's
audioBase64 | url, mimeType, durationMs}]}
whitespace-collapsed narration — and the reader plays it and asks no provider for anything.<basename>.audio.json (guided-narration-audio@0.1) is the natural source: its text is
already collapsed and each beat's file becomes an entry's url. Prefer url overaudioBase64 for anything long, and let the reader page's own connect-src reach the host so
it can fetch those files. A beat the recording misses is read by the computer voice, so a
partial recording still helps; an invalid one is refused with guided-recording-refused and
the reader carries on without it.
Render and inspect the tour, including any image stops. Confirm the actual speech, target
order, and chapter navigation. The runtime speaks authored paragraphs without rewriting them.
The hosted playground uses Hub-authenticated Gemini when available and computer voice otherwise;
validation and rendering do not generate audio. Do not claim cloud playback was verified from
conformance checks alone. For a local checkout, use node site/tools/preview.mjs 4382 so
authenticated audio can reach the existing service through the fixed local relay.