Generate and publish websites.
For motivation and design choices, see Site Publisher.
This file is about how to use the generator and the syntax of each construct in each markup.
- Usage
- Summary
- Opinionated
- Markup
- Frontmatter
- Categories
- Page Names
- Entities
- Directories
- TEI dates
- TEI
gap - Facsimiles
- Blog Posts
- Paging
- SEO
- Graph
- External links
- Chunking
- Table of contents
- Tables
- Task Lists
- Footnotes
- Glossary
- Bibliography
- Code Blocks
- Callouts
- Admonitions
- Asides
- Quotes
- Strikethrough
- Figures
- PDF embeds
- Video
The source directory must contain _site_config.yml.
Pass that directory as the only positional argument:
./gradlew run --args="/path/to/source-directory"
--serve generates and leaves the HTTP server running (same as IntelliJ @main def generate()).
--pretty-print rewrites authored TEI and DocBook XML in place and does not generate a site.
It skips a file whose bytes would not change.
It does not reformat Markdown, AsciiDoc, or HTML.
--pretty-print together with --serve is an error.
--log-level defaults to INFO.
--treat-errors-as-warnings, --include-drafts, --production, and --target-directory-name are the other flags.
SITE_PUBLISHER_* environment variables are the same names with hyphens turned into underscores.
./gradlew build compiles and tests; ./gradlew test runs tests only.
Maven coordinates: org.podval.tools:org.podval.tools.publisher:0.4.0
A site with Gradle applies plugin id org.podval.tools.site-publisher (same version as the library).
The plugin adds a detached sitePublisher configuration and two JavaExec tasks that run
org.podval.tools.publish.site.Site.
Do not put the library on implementation: Scala version and Playwright stay off the project compile classpath, and off
the Gradle daemon.
plugins {
id 'org.podval.tools.site-publisher' version '0.4.0'
}
site {
treatErrorsAsWarnings = true
logLevel = 'INFO'
}
pluginManagement.repositories must include mavenCentral() (the plugin marker is published there) as well as
gradlePluginPortal() for other plugins such as Foojay.
./gradlew generateSite writes the target directory.
./gradlew serveSite is the same with --serve.
./gradlew prettyPrintSite passes --pretty-print and does not publish _site.
generateSite and serveSite pass --pretty-print=false.
Do not make generateSite or prettyPrintSite a dependency of build; Pages CI calls generate explicitly.
OpenTorah :docs keeps tasks.named('generateSite') { dependsOn generateTables }.
site { } maps to CLI flags (defaults match SiteOptions):
| Property | Default | CLI |
|---|---|---|
| sourceDirectory | project directory | positional source path |
| targetDirectoryName | _site | --target-directory-name (absolute path used as-is) |
| treatErrorsAsWarnings | false | --treat-errors-as-warnings |
| logLevel | INFO | --log-level |
| includeDrafts | false | --include-drafts |
| production | false | --production |
Both tasks set --enable-native-access=ALL-UNNAMED, --sun-misc-unsafe-memory-access=allow,
PLAYWRIGHT_SKIP_VALIDATE_HOST_REQUIREMENTS=1, and PLAYWRIGHT_BROWSERS_PATH=$GRADLE_USER_HOME/ms-playwright.
generateSite launches Java 25.
Its inputs are the source tree minus the target directory and Gradle build/; its output is the target directory, so
Gradle up-to-date can skip a run when sources are unchanged.
The generator still deletes and rewrites the target when the task runs.
The plugin adds the library as a default dependency of the sitePublisher configuration, at the same version as the
plugin (org.podval.tools:org.podval.tools.publisher:0.4.0).
Gradle defaultDependencies run only when that configuration is still empty, so a consumer can pick another published
library version by declaring sitePublisher themselves (there is no site.publisherVersion).
Do not add a second coordinate alongside the default — a non-empty configuration replaces the default entirely:
dependencies {
sitePublisher "org.podval.tools:org.podval.tools.publisher:${providers.gradleProperty('sitePublisherVersion').get()}"
}
The supported pair is the same version as the plugin.
An older library may not accept flags the plugin always passes; a newer one is usually fine.
A local settings-body includeBuild of this repo substitutes the library module for whatever version string was
requested.
Site workflows: checkout, setup-java 25, setup-gradle, cache ~/.gradle/ms-playwright, ./gradlew generateSite (or
:docs:generateSite), upload _site / docs/_site, deploy Pages.
Pin the plugin version in build.gradle.
Known sites apply the plugin and call ./gradlew generateSite (or :docs:generateSite).
GitHub Actions can still call dubinsky/site-publisher/.github/actions/generate@v0.1.0 for repos that do not apply the
plugin (pin the action tag; optional version input overrides that action checkout’s gradle.properties).
Sites using this publisher:
This repo already includeBuild`s sibling `../xml and ../../OpenTorah/opentorah.org when those checkouts exist
(-PxmlDir=
/ -PopentorahDir=).
A site that applies org.podval.tools.site-publisher substitutes a local checkout with two includeBuild`s of the
same directory: `pluginManagement (plugin id; pass a String, not a File) and the settings body (library child
org.podval.tools:org.podval.tools.publisher).
pluginManagement alone does not substitute the library now that it is not the included-build root.
Do not use mavenLocal().
pluginManagement {
repositories {
mavenCentral()
gradlePluginPortal()
}
final String publisherDir =
providers.gradleProperty('sitePublisherDir').getOrElse('../site-publisher')
if (file(publisherDir).isDirectory() && file("${publisherDir}/settings.gradle").isFile()) {
includeBuild publisherDir
}
}
// after plugins { }:
final File publisherDir = file(
providers.gradleProperty('sitePublisherDir').getOrElse('../site-publisher')
)
if (publisherDir.isDirectory() && new File(publisherDir, 'settings.gradle').isFile()) {
includeBuild(publisherDir)
}
Override with -PsitePublisherDir=.
OpenTorah / MathWorlds defaults stay ../../Podval/site-publisher.
When the directory is missing (CI), Gradle resolves the plugin and the library from Maven Central.
Without a site-side plugin, generate from this repo: ./gradlew run --args="/path/to/source-directory".
./gradlew nmcpZipAggregation ./gradlew publishAggregationToCentralPortal
The zip includes the library (org.podval.tools:org.podval.tools.publisher), the plugin implementation
(org.podval.tools:site-publisher-plugin),
and the plugin marker (org.podval.tools.site-publisher:org.podval.tools.site-publisher.gradle.plugin).
Same version.
Site-header links are listed in _site_config.yml in display order (source path or site path, same resolution as
home):
header-pages: - notes/index.md - posts - tags - graph
Omitted or empty means no header links (up/prev/next and format icons still appear). Unknown entries are recorded on the Errors page.
Optional facsimiles-url is the base URL of a facsimile JPEG site (trailing slash as in the collector
tei@facsimilesUrl).
Omitted means none.
See Facsimiles.
facsimiles-url: https://storage.googleapis.com/facsimiles.alter-rebbe.org/
Optional tei-default-calendar is julian or gregorian (default gregorian).
TEI date/@when numbers are in that calendar unless the element has calendar="#julian".
See TEI dates.
tei-default-calendar: julian
Optional named-windows (default off) reuses four collector window names so collection, names, transcription, and
facsimile can sit in separate tabs.
Internal links target the destination viewer; each page sets window.name.
Other sites should leave it off.
named-windows: true
This here is a static site generator. It:
-
is Opinionated;
-
recognizes various Markup languages;
-
interprets Frontmatter;
-
resolves various Page Names;
-
recognizes TEI Entities;
-
lists Directories;
-
supports Blog Posts;
-
supports Chunking;
-
supports Paging;
-
supports PDF generation;
-
supports TEI Facsimiles;
-
writes SEO tags in each page
<head>; -
can emit an interactive Graph of pages;
-
checks External links when
check-linksis set;
Supported markup constructs:
Since I am writing this site generator for my own use, it is going to be opinionated:
-
everything is written in Scala
-
no plugins
-
no SCSS
-
no template languages
-
no layouts
TODO expound
Recognized by extension:
-
Markdown (
.md) -
AsciiDoc (
.adoc) -
HTML (
.html) -
TEI (
.xmlwith a TEI root:TEI,store,collection,person,place,org) -
DocBook (
.xmlwith a DocBook root:article,book,chapter,appendix,part,set,preface,refentry,topic)
Internal (YAML between --- lines in the markup file) or external (same basename, .yaml / .yml).
Do not use both.
Obsidian treats internal YAML as Properties; it stores extra keys but does not
interpret this publisher’s fields (chunk, bibliography, csl, …).
Set title in front matter only when it differs from the file name.
Both the file name and the title are used for link resolution.
If the document also has a title (h1, AsciiDoc =, TEI/DocBook title), it must match the front matter title; a
mismatch is recorded on the Errors page.
Useful fields:
-
chunk/chunk-depth— see Chunking -
bibliography/csl/lang— see Bibliography -
toc-depth -
description,author,lang— page<head>(SEO); site values in_site_config.ymlare the fallback. Sitelangalso selects store and selector display names (publisherSelector.xml, TEI<name lang>); omitted isen. -
tags,categories(Categories),aliases -
icon,permalink,post,date -
pdf- see PDF -
asset— copy the markup file as an asset (no dialect processing, no site chrome). Standalone sidecar only:name.yml/name.yamlnext toname.html(or.md,.adoc,.xml, …) withasset: true. The sidecar is not published. Internal---asset: trueis an error (the file stays markup). A directoryindexso marked stays a directory page (parent/children) butwritecopies the file (no listing, no chrome).
permalink (absolute, e.g. /short) and aliases (relative or absolute) each create a Refresh page at that URL.
A permalink that does not start with / is an error (/errors.html#permalink) and creates no alias.
When the page is a directory (an index or a TEI store / collection beside its folder), /permalink/child and wiki
[[permalink/child]] resolve to child under that directory — the short name is a path prefix, not only a leaf
bookmark.
A remainder under a non-directory alias does not resolve.
A last-dot suffix in a href that is all digits is part of the name (255.2), not a file extension.
categories is a list of links to authored pages.
A value is a wiki link ([[Books]] or [[Books|label]]) or a bare name (Books).
It resolves the same way as a wiki link in the body.
The target is a normal page: its prose stays, and the generator appends the pages that name it, sorted by title.
A value that resolves to nothing is an unresolved link on the Errors page.
No category page is created for it.
![[Note.base]] and ![[Note.base#View]] are Obsidian queries and are omitted, with no error.
The member page header lists each category beside its tags.
A wiki link in the body is a backlink.
It does not put the page in that list.
A page may name several categories.
Jekyll relies on page titles being given explicitly in the front matter and ignores the file names; I had to use a TODO
Obsidian plugin to set the titles from the file names to communicate them to Jekyll; as a result, pages where the title
is different from the file name (e.g., index pages) had to be marked with TODO in their front matter to disable this
plugin from setting the front matter title.
My plugin uses both the front matter title and the file name for link resolution, so I can remove this Obsidian plugin and redundant front matter titles which are the same as the file name, and set the title in the front matter only when it is different from the file name.
Obsidian uses file names exclusively and ignores the front matter title; for link resolution to use front matter titles also, TODO Obsidian plugin needs to be installed.
Wiki links ([[…]], including [[note|text]] and [[#id]]) are native Obsidian.
Block ids: ^id at the end of a paragraph, or on its own line after a list, table, quote, or code block (blank line
before the id).
Link with [[note#^id]].
Obsidian embeds of pages (![[note]], ![[note#heading]], ![[note#^block]]) copy that authored
region into the host (any markup on the target).
The copy is a boxed aside.transclusion whose header is a link to the source (Notes Title or
Notes Title § heading).
The gear Settings checkbox Seamless transclusions hides the box and header (html.transclusion-clean).
A missing heading or block is an error (not a silent whole-page embed).
Footnotes inside a transcluded region are merged and numbered with the host.
Media embeds (![[file.pdf]], images, audio, video) are unchanged.
AsciiDoc include:: and TEI/DocBook xi:include are not this feature.
Wiki links in Obsidian:
-
are they case-sensitive?
-
do they take into account file name, title property, document title?
-
how do they look in the file if name reference is ambiguous?
-
can they be split between the lines?
-
do they support any kind of agglutination?
A TEI file whose root is person, place, or org is an entity; the kind is that root.
The entity id is the file name without .xml (name/alter-rebbe.xml → alter-rebbe).
The file may sit anywhere in the tree.
persName, placeName, and orgName with a non-empty @ref resolve to an entity when the kinds match and @ref is
exactly that id (not a path, not the displayed name, not a title).
A kind mismatch or unknown id is an unresolved internal link and is recorded on the Errors page.
Without @ref, the name is left as-is (not a link) and is recorded on the Errors page (name without @ref,
/errors.html#no-ref).
That check runs on the first parse of TEI documents and of store/collection title / abstract / body — not on the
name elements that define an entity file.
TEI unclear is /errors.html#unclear.
An entity file whose name is not the first name with spaces turned into underscores is /errors.html#misnamed-entity.
Each kind on the Errors page is a fragment (AmbiguousTitle → #ambiguous-title).
When more than one kind is present, a TOC links to those fragments.
<persName ref="alter-rebbe">Алтер Ребе</persName>
в <placeName ref="Вильна">Вильне</placeName>
<orgName ref="Виленский_кагал">виленский кагал</orgName>Wiki [[alter-rebbe]] still finds the entity page by file name.
Entity refs do not title-walk, so a markdown page titled alter-rebbe does not steal persName ref="alter-rebbe".
A TEI entityLists file is a catalog of buckets, not a store include list.
Beside a directory (name.xml next to name/) it is that directory’s index (/name/index.html).
Each listPerson, listPlace, or listOrg is a bucket:
-
the list kind is the element (
listPerson→personfiles); -
@nis the list id (/name/{n}.htmlwhen the catalog has two or more non-empty lists); -
@rolepresent: entities of that kind whose root@rolematches, anywhere in the site; omitted: entities of that kind with no role; -
<title>on the list is the list page title; a direct child<title>ofentityListsis the catalog page title.
Empty lists are omitted.
Members are every matching entity on the site, sorted by file name; the link text is the first persName / placeName
/ orgName.
The entity page <h1> / <title> is that first name, not the file name.
Document backlinks on an entity page still come from @ref in other documents (grouped by TEI collection when the
source sits under one); the catalog itself is not a backlink source.
When two or more lists are non-empty, the catalog page is a TOC of those lists (links to the list pages; members are not
inlined).
A catalog with a single non-empty list is that list (no {n} subpage).
Several entityLists files are independent catalogs.
The catalog directory is an identity prefix in collection-aliases.json (/name → /name/index.html, /name/{id} →
/name/{id}.html), the same Worker mechanism as collection aliases.
There is no synthetic /name.html dump.
<entityLists>
<title>Имена</title>
<listPerson n="jews" role="jew"><title>Жиды (они же Евреи)</title></listPerson>
<listPerson n="unknown"><title>Неизвестно кто</title></listPerson>
<listPlace n="places"><title>Места</title></listPlace>
</entityLists>Root / is index.md / index.adoc, or a synthesized listing of that folder.
To open at a specific page instead, set home in _site_config.yml to an absolute site path.
If the document is chunked, use the TOC (P/index.html):
home: /book/book/index.html
/index.html is then a meta Refresh to that page (not an HTTP 301).
Do not also author index.md / index.adoc.
Other root files stay on the site; they are just not listed at /.
Collection short names (collector site.xml <alias n="rgada" to="…"/>) are an optional alias on the TEI store /
collection root, not site config and not front-matter permalink:
<collection n="3140" directory="3140" alias="rgada">
alias is the URL prefix (/rgada, /rgada/003).
The publisher does not write a Refresh file at /rgada.html; Pages.find, emitted hrefs (including canonical/sitemap),
and serve() rewrite /rgada/003 to the written …/3140/003.html.
Inbound old collector facsimile URLs /rgada/facsimile/003 rewrite the same way (see Facsimiles).
Front-matter permalink / aliases remain for leaf Refresh pages.
Generate also writes collection-aliases.json (segment from → collection directory to) for a Cloudflare Worker.
The Worker is an internal rewrite (not a 301): longest slash-delimited prefix, then .html unless the request already
has an extension (keep .xml); inbound /alias/facsimile/P becomes {to}/P/facsimile.html.
Entity-lists catalogs are an identity prefix (/name → /name/index.html).
Prefix routes (hostname/rgada*, hostname/name*, …) so assets skip the Worker.
Deploy with .github/actions/deploy-alias-worker.
Canonical entity hrefs stay the file path (/name/{id}.html when the file lives under name/).
A directory permalink (an index or TEI store/collection) is also a prefix: /short/child resolves under that
directory, and emitted hrefs use the short path.
A TEI store file in the source root (archive.xml next to archive/) also gets two synthetic indexes, named from the
file: {name}-collections.html (nested tree of child stores and collections, collector /collections) and
{name}-index.html (flat list of descendant collections, collector /).
Titles come from this project’s Selector.xml catalog in the site lang (omitted lang is en): the store’s
by/@selector title for the tree (archive → Архивы when lang is ru) and the case selector title for the flat
list (Дела).
When there is exactly one such root store, /collections rewrites to the tree page (no Refresh file; Worker table
includes it).
Point header-pages at {name}-collections and home at /{name}-index.html when those should be chrome / the
landing page.
A TEI store or collection file beside a directory (dir.xml next to dir/) is that directory’s index.
A collection index sets class="wide" on <html> so the content .wrapper is 1800px; header and footer stay the
default width.
Store indexes and documents under the collection do not.
xi:include/@href (relative to the store file) is an ordered child reference, not XInclude: the target is not inlined.
Store children are listed in include order (name: title).
A collection body is table.collection-index (Описание, Дата, Кто, Кому, Язык, Документ, Страницы, Расшифровка), not
a page list.
Optional <part from="000"> title rows split the table; pageType is manuscript (default: 000 / 000об) or book
(numeric).
A file named {base}-{xx} (two-letter language, dash at length 3) is a translation of {base}: not a table row; Язык
is he ru with ru linking to the translation; the original document’s site header has [ru]; prev/next skip
translations.
text/@xml:lang must match the suffix.
The page header (.store-header) holds ancestor path lines and this node’s selector + name + title as <l>`s, then
`abstract/body and the by/@selector label.
A TEI document under a collection also gets a document-header table from teiHeader (Описание, Дата, Кто, Кому,
Расшифровка).
The Дата cell (and in-text dates) get calendar hover tables; see TEI dates.
Selector labels (category → разряд when site lang is ru) come from this project’s Selector.xml
(page.Selectors).
Documents under a store or collection use the same header.
A scanned file under that directory that is not named in the includes is an error (translations are not).
Selector hops in the href (book/ in books/book/derzhavin.xml) are URL segments only — they are not pages, and up
skips them.
Missing pb@missing photos are listed under the table.
Collection Страницы links to pb ids in the transcription (#p{n}); in-text pb links to the
facsimile viewer when facsimiles-url is set.
-
directory structure
-
navigation
-
header pages
-
icons
A TEI <date> with @when, @from, @to, @notBefore, or @notAfter gets a hover table.
.. in any of those attributes is InvalidDate, and the element is left as authored.
Each value is YYYY, YYYY-MM, or YYYY-MM-DD.
Month names follow the site lang.
The column headers stay English.
A full day on @when is one column, Date.
A year or a month on @when is both ends of that year or month.
A year or a month on @from or @notBefore expands to the first day of that end only.
A year or a month on @to or @notAfter expands to the last day of that end only.
@from equal to @to, or @notBefore equal to @notAfter, stays two columns, not Date.
| What you write | Hover columns |
|---|---|
| @when on a full day (YYYY-MM-DD) | Date |
| @when on a year or a month | From, To |
| @from and @to | From, To |
| @notBefore and @notAfter | Not before, Not after |
| @from only | From |
| @to only | To |
| @notBefore only | Not before |
| @notAfter only | Not after |
| @from and @notAfter | From, Not after |
| @notBefore and @to | Not before, To |
These combinations are InvalidDate, and the element is left as authored:
-
@whentogether with@notBefore,@notAfter,@from, or@to -
@fromtogether with@notBefore -
@totogether with@notAfter
@from with @notAfter, and @notBefore with @to, are allowed.
The numbers are Gregorian unless the element has calendar="#julian", or _site_config.yml has
tei-default-calendar: julian and @calendar is omitted.
Any other @calendar value is Gregorian.
This disagrees with the TEI Guidelines: the attributes are not rewritten to Gregorian; @calendar selects how the
numbers are read, and it still describes the prose.
The collection index and document-header Дата cell show @when as the visible text when @when is present.
Otherwise the cell shows the element’s own text, still inside <date>, so the tip can attach.
In-text dates always keep their text.
A <date> with none of @when, @from, @to, @notBefore, and @notAfter is unchanged.
A TEI <gap reason="…"> with a non-empty @reason gets a hover tip with that text (same family as dates).
Missing or blank @reason leaves the element as-is.
Existing tei.css rules for gap (lost / illegible brackets) still apply.
When _site_config.yml has facsimiles-url, each TEI document with at least one non-missing pb gets a viewer at
P/facsimile.html (published /alias/P/facsimile.html when the collection has alias).
Photos scroll in a pane under the header (the window is the size; there is no inner resize box).
Old collector bookmarks /alias/facsimile/P (and /alias/facsimile/P.html) still resolve to that viewer, including
translation /alias/facsimile/P-xx (shared original).
#p{n} fragments are unchanged.
New pages emit only the new URL.
The JPEG for pb@n is {facsimiles-url}{directory of the TEI file}/{n}.jpg.
A trailing slash on the config value is optional.
pb@facs overrides that URL.
Missing pb@missing pages are omitted from the scroller.
In the transcription, <pb n="000-1"/> is an images icon (a.pb, id="p000-1") linking to the viewer fragment,
target="facsimile" (facsimileViewer when named-windows is on).
The site-header format icon does the same without a fragment.
Collection Страницы is page numbers on the transcription #p{n} only.
Each photo in the viewer links back to that pb in the transcription (target="text" / textViewer).
{base}-{xx} translations share the original’s viewer.
Omitted facsimiles-url: no viewer pages; pb keeps its id (collection Страницы still works) and has no href.
A page is a blog post in one of three ways:
-
A file in
_posts/or_drafts/namedYYYY-MM-DD-titleis published at/YYYY/MM/DD/title.html. -
Front matter
post: truekeeps the file path and adds a Refresh page at/YYYY/MM/DD/{file name}.html.dateis required.post-titlereplaces the file name in that path. -
Front matter
permalink: /YYYY/MM/DD/titledoes the same for any source file name. The Refresh page is that permalink.
/posts lists each real post once.
The Refresh file is not a second entry.
Front matter description is the post teaser: Atom <summary> and, when present, a p.post-excerpt on every /posts
batch.
Not auto-cut from the body.
Site description is not used here (SEO still falls back to it).
The /posts listing only.
In _site_config.yml:
paginate-posts: 10
Omitted or less than 1 is off.
Page 1 stays /posts.html; further batches are /posts/2.html, /posts/3.html.
A nav.pagination sits under the list.
No extra config.
Each HTML page <head> gets a document title (Page | Site title; home is just the site title),
meta name="generator" pointing at this publisher, description, author, canonical URL, Open Graph, a Twitter summary
card, and JSON-LD (WebSite / BlogPosting / WebPage).
Taken from existing fields: site title, description, url, author, lang, social.twitter; page title,
description, author, lang, date (and git/modified_time for modified).
No image, Facebook, or webmaster keys.
Off by default.
When _site_config.yml has graph.enabled: true, generate writes /graph.json and /graph.html.
/graph.html is a pan/zoom map of canonical authored pages and the resolved internal links between them (Cytoscape.js,
loaded only on that page).
Click a vertex to open the page (named-windows targets apply when that option is on).
Vertices are authored pages with a source file: notes, posts, TEI documents, entities, stores, and entity-list catalogs.
Aliases, chunks, facsimile viewers, PDFs, assets, and synthetics (/errors, /tags, /posts, …) are not vertices;
links to those views count as links to the owning page.
Edges come from existing backlinks (kind: link, including TEI @ref) and, unless turned off, transclusions
(kind: transclude).
include-transclusions defaults to true; omit it or set it explicitly.
Unresolved links stay on the Errors page, not as graph nodes.
Do not name a collection @alias="graph": that path is reserved for the graph page.
graph:
enabled: true
include-transclusions: true
exclude-path-prefixes:
- days
exclude-path-prefixes match the source path (days/2020-10-23.md, archive/lvia/…), not the published URL.
Daily notes publish as /YYYY/MM/DD/index.html; TEI collection aliases publish as /lvia1799-2/….
A notes vault typically enables the graph and excludes the Obsidian daily-notes folder (days when that is the
configured folder).
A large TEI archive should leave the graph off, or exclude archive to keep only names and notes.
Add graph to header-pages if you want a header link (there is no automatic icon).
Pipeline notes belong in the Site Publisher design note, not here.
Off by default.
When _site_config.yml has check-links: true, generate requests each author-written http and https URL and lists
failures on the Errors page (#broken-link).
check-links: true
Checked: links (a@href) and media (img / video / audio / source @src, PDF object@data), including links
inside footnotes and bibliography entries.
An absolute URL whose host is the site url is a spurious self-link and is not requested.
mailto:, tel:, javascript:, and data: are not requested.
URLs inside code, and URLs only reached by transcluding another page, are checked on the page that authors them.
The same URL is requested once per generate.
Each page that contains it gets one error.
The request is HEAD.
On 403, 405, or 501 it is retried once with GET.
Redirects are followed.
A status of 400 or higher, or a timeout or connection failure, is a broken link.
The error text includes the status or the failure message.
Footer social links, the license link, and facsimile JPEGs are not checked. Generation still finishes. The Errors page is the report, as for unresolved links. Timeouts are 5 seconds to connect and 10 seconds per request.
Split a document into per-section HTML pages (any markup, not only AsciiDoc).
chunk: true chunk-depth: 2
P.html is the unsplit document (wiki lands there).
The split view is P/index.html (P/ on a static host): preamble plus TOC, then P/Section.html for each chunk.
File hrefs to P/index.html stay there; they are not rewritten to P.html.
No directory listing is synthesized for P/.
When both views exist, the site header has icon links between them (list for chunked, file for one-page). Chunked TOC shows ancestors plus the current subtree; sibling branches are titles only.
A page with chunk: true or toc-depth gets a generated TOC at the top if none is marked in the body.
A placeholder in the body is replaced with that TOC (the first one only) even without those keys; toc-depth still
controls how deep it goes (default 2).
- Markdown
-
Kramdown
{:toc}on a list. The list (seed item included) is replaced. Ordered or unordered:* this list is replaced with the TOC {:toc} ## Alpha
or
1. seed/{:toc}for a numbered outline.
A paragraph whose only text is [TOC] (any case) is the same placeholder (Typora, GitLab).
It must be a top-level paragraph: not a link ([TOC]: /url then [TOC]), not in a list, quote, or code span.
- HTML
-
<div class="toc-placeholder"></div>.
AsciiDoc, TEI, and DocBook have no in-body placeholder; use chunk / toc-depth.
Front matter pdf: true prints P.html to P.pdf (Chromium).
The site header on P.html (and on chunked pages, if any) gets a PDF icon linking to it.
The site header is hidden in print, so the PDF itself does not include those icons.
Markdown tables via FlexMark Tables Extension (GFM; Obsidian supports them natively).
- Markdown
-
| A | B | |---|---| | 1 | 2 |
AsciiDoc tables via Asciidoctor.
- AsciiDoc
-
|=== | A | B | 1 | 2 |===
- TEI
-
<table> <row role="label"><cell>A</cell><cell>B</cell></row> <row><cell>1</cell><cell>2</cell></row> </table>
- DocBook
-
CALS
row/entry(optionaltgroup);morerowsbecomesrowspan.<table> <tgroup cols="2"> <thead><row><entry>A</entry><entry>B</entry></row></thead> <tbody><row><entry>1</entry><entry>2</entry></row></tbody> </tgroup> </table>
- HTML
-
<table> <thead><tr><th>A</th><th>B</th></tr></thead> <tbody><tr><td>1</td><td>2</td></tr></tbody> </table>
Unchecked [ ] and checked [x] items in a list.
The static site shows disabled checkboxes.
- Markdown
-
GFM / FlexMark Task List Extension (Obsidian supports them natively):
- [ ] open - [x] done
- AsciiDoc
-
native checklists:
* [ ] open * [x] done
DocBook has no native task-list markup.
- HTML
-
the IR shape (already a task list; no conversion):
<ul class="task-list"> <li class="task-list-item"> <input type="checkbox" class="task-list-item-checkbox" disabled> open </li> <li class="task-list-item"> <input type="checkbox" class="task-list-item-checkbox" disabled checked> done </li> </ul>
Hovering a footnote number shows the note; the numbered list is at the end of the (chunk of the) page. A note written with no space after the preceding word or element stays adjacent in the HTML.
A marker in a table cell is published as a lowercase Latin letter (a, b, …, z, aa) with the list immediately
under that table (div.table-with-notes).
Markdown still defines those bodies at the end of the source file; the published page moves them.
AsciiDoc [1] is written at the call site (including in a cell); the published page still lists those notes
under the table.
Collector collection-index and document-header tables stay in the page arabic series.
A footnote inside another footnote is one level of letters a, b at the end of the parent note.
There is no nested hover tip: the outer tooltip shows inner letter markers, not a nested list.
Markdown and AsciiDoc get nested notes only if the processor emits an inner stub.
TEI note place="end" inside another place="end" does.
TEI still converts only place="end".
place="foot" and unplaced note are unchanged.
@n, when non-blank, is the marker at the call site and in the list.
Order and fragment ids stay on the generated series.
The same marker may appear on more than one note.
Obsidian inline ^[…] footnotes are unsupported.
- AsciiDoc
-
inline note, or a named definition reused with an empty macro: A
[2]in a table cell is written at the call site; the published page lists it under the table.See this footnote:[A note.] and again footnote:fn[Same note.] later footnote:fn[]. - Markdown
-
FlexMark: use
[^id]in the text; define[^id]:later. Obsidian supports that form natively (Help: Footnotes). Obsidian also has inline^[…]footnotes; this publisher does not. A[^id]in a table cell uses the same[^id]:body at the file end; the published page moves it under the table.See this [^note] and again [^note]. [^note]: A note.
- TEI
-
definition and use are the same element at the point of reference (
@n, when non-blank, is the marker):<p>See this<note place="end">A note.</note>.</p>
A non-blank
@nis that marker. The next note with no@nis still numbered in series.<p>See this<note place="end" n="*">A note.</note> and <note place="end">the next.</note>.</p>
A
place="end"inside anotherplace="end"is a nested letter on the parent note:<p>See this<note place="end">Outer<note place="end">Inner.</note></note>.</p>
- DocBook
-
<footnote>at the call site (xml:idoptional; non-blanklabelis the same marker). Reuse with<footnoteref linkend="id"/>. A<footnote>inside another footnote’s<para>is nested the same way if the converter emits an inner stub.<para>See this<footnote><para>A note.</para></footnote>.</para>
- HTML
-
IR stubs (same
footnote-correlation-idon the call site and the body):<p>See this<span class="footnote-link" footnote-correlation-id="note"></span>.</p> <span class="footnote" footnote-correlation-id="note">A note.</span>
A glossary is a description list marked as such. Term links get hover definitions (print inlines them in parentheses).
- AsciiDoc
-
[glossary]on the description list; term ids from. Reference a term with[id].See <<posuk>>. [glossary] [[posuk]]posuk:: verse - Markdown
-
PHP Extra definition lists (https://michelf.ca/projects/php-markdown/extra/#def-list) via FlexMark Definition List Extension. Mark the list with
{:.glossary}after it (Kramdown IAL, same family as Table of contents). Without that marker the list is not a glossary (no term ids, no tooltips). The marker also separates adjacent lists.
Obsidian does not render PHP Extra definition lists natively (Markdown
Guide: Obsidian).
Community plugin Pandoc Extended Markdown renders Pandoc/PHP
Extra term / : definition lists.
{:.glossary} is Kramdown IAL, not Obsidian; it shows as a paragraph in the note.
[[#id]] heading links are native Obsidian.
+
See [[#posuk]].
posuk
: verse
{:.glossary}Term id is the id on <dt> if present, otherwise the term text with spaces turned into hyphens (Alter Rebbe →
Alter-Rebbe).
Reference a term with [[#id]] or [text](#id).
A lone : after a Markdown term is not a reliable empty definition: FlexMark can take the next term as that definition.
Write a real definition line, or omit the term.
- TEI
-
<list type="gloss">(ortype="glossary") of<label>/<item>. Term id fromxml:idon the label, else on the item, else the label text with spaces turned into hyphens. Reference a term with<ref target="#id">or<term ref="#id">.<p>See <ref target="#posuk">posuk</ref>.</p> <list type="gloss"> <label xml:id="posuk">posuk</label> <item>verse</item> </list>
- DocBook
-
<glosslist>(or<glossary>) of<glossentry>/<glossterm>/<glossdef>. Term id fromxml:idon the entry, else the term, else the term text with spaces turned into hyphens. Reference a term with<link linkend="id">.<variablelist>is a description list, not a glossary.<para>See <link linkend="posuk">posuk</link>.</para> <glosslist> <glossentry xml:id="posuk"> <glossterm>posuk</glossterm> <glossdef><para>verse</para></glossdef> </glossentry> </glosslist>
- HTML
-
the IR shape (
class="glossary"; each term isclass="glossary-item"with anid):<p>See <a href="#posuk">posuk</a>.</p> <div class="glossary"> <div class="glossary-item" id="posuk"> <dt>posuk</dt> <dd>verse</dd> </div> </div>
Two kinds, per document (no site-wide file or style). They can appear in the same document.
Front matter:
-
bibliography— path to a.bibfile, relative to this source file (.bibfiles are not copied to the site; un-ignore with_site_ignoreif you want one published) -
csl— CSL style name from the bundled styles (for exampleapa,ieee) -
lang— CSL locale; falls back to the sitelang, thenen-US
Both bibliography and csl are required.
Only works cited on the page appear in the list.
In-text cites become links to #bibl-{key} on the matching list entry (first key if a cite names several).
Unknown keys are highlighted and recorded as a page error.
- AsciiDoc
-
no
asciidoctor-bibliographygem. Alsocite:key[]/cite:key[p. 12](standard inline-macro form).- `cite:[knuth79]`, `cite:[knuth79, p. 12]`, `cite:[knuth79, lamport94]` — parenthetical - `cite:knuth79[]`, `cite:knuth79[p. 12]` — same, standard inline-macro form - `citenp:[lamport94]` — narrative (author in the sentence) - `bibliography::[]` — list placeholder (otherwise the list is appended)
- Markdown
-
Pandoc citation syntax (no extra FlexMark extension).
me@host.comis left as an email. Obsidian does not format[@key]natively. Community plugin Pandoc Reference List recognizes Pandoc citekeys in the note (sidebar; optional inline). Citations inserts[@key]from a.bibfile.- `[@knuth79]`, `[@knuth79, p. 12]`, `[@knuth79; @lamport94]` — parenthetical - `@knuth79` — narrative (`me@host.com` is left as an email) - `[-@lamport94]` — suppress author - `:::bibliography` — list placeholder See [@knuth79] and [@knuth79, p. 12] and [@knuth79; @lamport94]
- DocBook
-
<citation>key</citation>(optional locator after a comma).biblioref@linkendthat is not a native entry id is the same. Empty<bibliography/>is the list placeholder (otherwise the list is appended).<para>See <citation>knuth79</citation> and <citation>knuth79, p. 12</citation>.</para> <bibliography/>
- HTML
-
the IR shape:
<p>See <span class="citation" data-mode="parenthetical"><span class="citation-item" data-key="knuth79"></span></span> and <span class="citation" data-mode="parenthetical"><span class="citation-item" data-key="knuth79" data-locator="p. 12"></span></span>.</p> <div class="bibliography"></div>
- TEI
-
@cRefonref/ptris the citeproc key.@nis an optional locator. Empty<div type="bibliography"/>is the list placeholder (otherwise the list is appended).<p>See <ref cRef="knuth79"/> and <ptr cRef="lamport94" n="p. 12"/>.</p> <div type="bibliography"/>
Authored entries stay on the page (not cited-only).
Links to an entry id get a hover tip.
Entry ids are as written (#knuth79); they do not use the #bibl- prefix of the CSL list, so the same key can be cited
both ways in one document.
- AsciiDoc
-
[bibliography]list with[[[id]]]/[[[id,xreftext]]]anchors; cite with[id].See <<knuth79>> and <<lamport94>>. [bibliography] * [[[knuth79]]] Knuth, Donald E. _TeX and Metafont_. 1979. * [[[lamport94,Lamport 94]]] Lamport, Leslie. _LaTeX_. 1994.
- TEI
-
ref/ptr@targetthexml:idof abibl/biblStructinlistBibl(not inteiHeader). Emptyptrshows@nor the id.citwith aquoteis a quote, not this. A bare@targetthat matches alistBiblid is treated as#id.<p>See <ref target="#knuth79">Knuth 1979</ref> and <ptr target="#lamport94"/>.</p> <listBibl> <bibl xml:id="knuth79">Knuth, Donald E. <title>TeX and Metafont</title>. 1979.</bibl> <biblStruct xml:id="lamport94"> <monogr><author>Lamport, Leslie</author><title>LaTeX</title></monogr> </biblStruct> </listBibl>
- DocBook
-
<bibliography>withbiblioentry/bibliomixed; cite with<link linkend="id">or<biblioref linkend="id"/>(emptybibliorefshows@xreflabelor the id).<para>See <link linkend="knuth79">Knuth 1979</link> and <biblioref linkend="lamport94"/>.</para> <bibliography> <biblioentry xml:id="knuth79">Knuth, Donald E. TeX and Metafont. 1979.</biblioentry> <bibliomixed xml:id="lamport94">Lamport, Leslie. LaTeX. 1994.</bibliomixed> </bibliography>
Markdown has no native in-document bibliography list.
Inline code and fenced blocks.
A language class on <code> (language-scala, …) loads highlight.js; language-mermaid loads Mermaid instead.
- Markdown
-
CommonMark inline
and fenced blocks. Obsidian supports both natively, including a language identifier (Help: Code). Fencedcode```mermaiddiagrams are native in Obsidian too (Help: Diagram); this publisher loads Mermaid forlanguage-mermaid.Use `map` in ```scala xs.map(f) ```
- AsciiDoc
-
inline
;code[source]listing (optional language).Use `map` in [source,scala] ---- xs.map(f) ----
- TEI
-
<code>(tagdocs module) with@langfor the formal language (notxml:lang). A newline in the content is treated as a block and wrapped in<pre>.<eg>/<egXML>are not converted.<p>Use <code lang="scala">map</code> in</p> <code lang="scala">xs.map(f) </code>
- DocBook
-
<code>/<literal>inline;<programlisting>/<screen>/<literallayout>as a block.@language(notxml:lang) is the highlight.js tag.<para>Use <code language="scala">map</code> in</para> <programlisting language="scala">xs.map(f) </programlisting>
- HTML
-
<p>Use <code>map</code> in</p> <pre><code class="language-scala">xs.map(f) </code></pre>
Numbered annotations on lines of verbatim (listing, literal, or source), plus a list of those annotations. Not a general footnote mechanism.
- AsciiDoc
-
<1>(or<.>for automatic numbering) at the end of a line; repeat the numbers in a callout list below. Hide a marker in the raw source behind a line comment (// <1>,# <1>,<!--1-→).[source,ruby] ---- require 'sinatra' (1) ---- <1> Library import
- DocBook
-
<co>in a listing;<calloutlist>/<callout arearefs>below.@labeloncois the marker number.<programlisting>require 'sinatra'<co xml:id="co1" label="1"/></programlisting> <calloutlist> <callout arearefs="co1"><para>Library import</para></callout> </calloutlist>
- HTML
-
the IR shape:
<pre><code>require 'sinatra' <span class="callout" data-value="1">1</span></code></pre> <ol class="callout-list"> <li>Library import</li> </ol>
Typed notices (note, tip, warning, …). Not listing Callouts and not asides/sidebars.
- Markdown
-
Obsidian core callouts (
> [!type]; optional title;+/-to fold). No plugin.> [!tip] Save time > Use the shortcut. > [!faq]- Hidden > Secret.
- AsciiDoc
-
NOTE:/TIP:/IMPORTANT:/CAUTION:/WARNING:(paragraph or[NOTE]====block; optional.Title).NOTE: Auxiliary information. .Watch out [WARNING] ==== Multi-paragraph body. ====
- DocBook
-
note,tip,warning,caution,important(optional childtitle).<note><title>Save time</title><para>Use the shortcut.</para></note>
- HTML
-
the IR shape (
detailswhen folded):<div class="admonition" data-type="tip"> <div class="admonition-title">Save time</div> <p>Use the shortcut.</p> </div>
Untyped auxiliary content next to the main flow. Not an admonition.
- AsciiDoc
-
[sidebar]or a**block (optional.Title)..Optional Title **** Auxiliary content. ****
- Markdown
-
no native syntax; Obsidian has none. Write an HTML
block.<aside> <p>From Markdown.</p> </aside>
- DocBook
-
<sidebar>(optional childtitle).<sidebar> <title>Optional Title</title> <para>Auxiliary content.</para> </sidebar>
- HTML
-
<aside>(IR addsclass="aside"if missing); optional title:<aside class="aside"> <div class="aside-title">Optional Title</div> <p>Auxiliary content.</p> </aside>
Block quotations.
Not Admonitions (> [!tip]) and not Asides.
- AsciiDoc
-
[quote](or a__block); optional.Title; attribution and citation title as[quote, Author, Source]..A title [quote, Jefferson, Papers] ____ A little rebellion now and then is a good thing. ____
- Markdown
-
a CommonMark blockquote (
>). No attribution syntax (an em-dash line is still quote body). Obsidian> [!type]is an admonition, not a quote.> A Markdown quotation. - DocBook
-
<blockquote>or<epigraph>; optionaltitleandattribution. Inline<quote>stays<q>.<blockquote> <title>A title</title> <para>A little rebellion now and then is a good thing.</para> <attribution>Jefferson</attribution> </blockquote>
- HTML
-
<blockquote>(IR addsclass="quote"if missing); optional title and<footer>attribution (citekept as-is):<blockquote class="quote"> <div class="quote-title">A title</div> <p>A little rebellion now and then is a good thing.</p> <footer class="quote-attribution">— Jefferson<br/><cite>Papers</cite></footer> </blockquote>
- TEI
-
<quote>;<cit>groups a quote with<bibl>(attribution). A<bibl>directly in<quote>is attribution too.<q>stays inline.<cit> <quote>A TEI quotation.</quote> <bibl>Jefferson, <title>Papers</title></bibl> </cit>
Deleted or superseded text.
HTML <del> (browser line-through).
- Markdown
-
GFM
~text~(Obsidian native; FlexMark strikethrough extension).See ~~struck out~~ here.
- AsciiDoc
-
text.See [line-through]#obsolete phrase# here. - TEI
-
<del>.<p>See <del>TEI struck</del>.</p>
- DocBook
-
<emphasis role="strikethrough">(orline-through). Default<emphasis>is<em>;role="bold"is<strong>.<para>See <emphasis role="strikethrough">DocBook struck</emphasis>.</para>
- HTML
-
<del>(IR);<s>andspan/markwithclass="line-through"become<del>:<p>See <del>HTML struck</del>.</p>
A block image (or other graphic) with an optional caption.
Inline images stay <img>.
Local src values are resolved to a site path (/pixel.svg); a missing file is an error (missing asset) and the
<img> gets unresolved-asset.
Wiki ![[image]] also finds a unique file of that name anywhere in the source tree.
- Markdown
-
a
on its own line; optional"title"is the caption. FlexMark does not turn Obsidian|WIDTH/|WIDTHxHEIGHTin the alt intoimgwidth/height(the size stays inalt). Obsidian![[image]]embeds stay<img>(resolved later);|WIDTHor|WIDTHxHEIGHTon that wiki embed do becomewidth/height. ![[pixel.svg|320]] ![[pixel.svg|320x240]]
- AsciiDoc
-
block
image::(optional.Title). Inlineis not a figure. Width and height (image::file.png[alt, 320, 240]orwidth=320, height=240) are kept on the<img>..A figure caption image::pixel.svg[A square]
- TEI
-
<figure>with<graphic url>and optional<head>(caption).<graphic>becomes<img>.<figure> <head>A TEI figure</head> <graphic url="pixel.svg"/> </figure>
- DocBook
-
<figure>(or<informalfigure>) with<imagedata fileref>and optional<title>(caption).<figure> <title>A DocBook figure</title> <mediaobject> <imageobject> <imagedata fileref="pixel.svg"/> </imageobject> </mediaobject> </figure>
- HTML
-
<figure>(IR addsclass="figure"if missing); optional<figcaption>:<figure class="figure"> <img src="pixel.svg" alt="A square"/> <figcaption class="figure-caption">HTML figure</figcaption> </figure>
A PDF file shown on the page. Not PDF generation of a page.
- Markdown
-
Obsidian
![[file.pdf]]; optional#page=Nand&height=H(pixels if H is digits). Alias![[file.pdf|Title]]is the link label.![[sample.pdf]]
- AsciiDoc
-
no native block; write the HTML IR (passthrough
).
DocBook has no native PDF-embed markup.
- HTML
-
<object type="application/pdf">(IR wraps it) or the IR:<div class="pdf-embed"> <object data="sample.pdf" type="application/pdf" aria-label="sample.pdf"> <a href="sample.pdf">Open PDF: sample.pdf</a> </object> <p class="pdf-embed-link"><a href="sample.pdf">Open PDF: sample.pdf</a></p> </div>
The sibling link is always visible and is what remains in print (the inner viewer is hidden).
data (and the fallback href) are resolved like other local assets; #page=N is kept.
A local video file, or a YouTube/Vimeo player.
Local files use HTML <video> (like <audio>).
Local src is resolved like images; YouTube/Vimeo <iframe> URLs are left alone.
- Markdown
-
Obsidian
![[file.mp4]](alsowebm,ogv,m4v). Alias![[file.mp4|Title]]is the fallback link label.![[clip.mp4]]
- AsciiDoc
-
video::file.mp4[](optional.Titlebecomes a figure caption). YouTube/Vimeo:video::id[youtube]/video::id[vimeo].video::clip.mp4[] - DocBook
-
<videodata fileref>. A YouTube/Vimeo URL becomes an iframe.<mediaobject> <videoobject> <videodata fileref="clip.mp4"/> </videoobject> </mediaobject>
- HTML
-
<video src>(IR addsclass="video"andcontrols); a YouTube/Vimeo<iframe>getsclass="video-embed":<video class="video" src="clip.mp4" controls="controls"> <a href="clip.mp4">Open video: clip.mp4</a> </video>