Skip to content

Repository files navigation

Static Site Publisher

Usage

The source directory must contain _site_config.yml. Pass that directory as the only positional argument:

./gradlew run --args="/path/to/source-directory"

--serve generates and leaves the HTTP server running (same as IntelliJ @main def generate()). --pretty-print rewrites authored TEI and DocBook XML in place and does not generate a site. It skips a file whose bytes would not change. It does not reformat Markdown, AsciiDoc, or HTML. --pretty-print together with --serve is an error. --log-level defaults to INFO. --treat-errors-as-warnings, --include-drafts, --production, and --target-directory-name are the other flags. SITE_PUBLISHER_* environment variables are the same names with hyphens turned into underscores.

./gradlew build compiles and tests; ./gradlew test runs tests only.

Maven coordinates: org.podval.tools:org.podval.tools.publisher:0.4.0

Gradle plugin

A site with Gradle applies plugin id org.podval.tools.site-publisher (same version as the library). The plugin adds a detached sitePublisher configuration and two JavaExec tasks that run org.podval.tools.publish.site.Site. Do not put the library on implementation: Scala version and Playwright stay off the project compile classpath, and off the Gradle daemon.

plugins {
  id 'org.podval.tools.site-publisher' version '0.4.0'
}
site {
  treatErrorsAsWarnings = true
  logLevel = 'INFO'
}

pluginManagement.repositories must include mavenCentral() (the plugin marker is published there) as well as gradlePluginPortal() for other plugins such as Foojay.

./gradlew generateSite writes the target directory. ./gradlew serveSite is the same with --serve. ./gradlew prettyPrintSite passes --pretty-print and does not publish _site. generateSite and serveSite pass --pretty-print=false. Do not make generateSite or prettyPrintSite a dependency of build; Pages CI calls generate explicitly. OpenTorah :docs keeps tasks.named('generateSite') { dependsOn generateTables }.

site { } maps to CLI flags (defaults match SiteOptions):

| Property | Default | CLI | |---|---|---| | sourceDirectory | project directory | positional source path | | targetDirectoryName | _site | --target-directory-name (absolute path used as-is) | | treatErrorsAsWarnings | false | --treat-errors-as-warnings | | logLevel | INFO | --log-level | | includeDrafts | false | --include-drafts | | production | false | --production |

Both tasks set --enable-native-access=ALL-UNNAMED, --sun-misc-unsafe-memory-access=allow, PLAYWRIGHT_SKIP_VALIDATE_HOST_REQUIREMENTS=1, and PLAYWRIGHT_BROWSERS_PATH=$GRADLE_USER_HOME/ms-playwright. generateSite launches Java 25. Its inputs are the source tree minus the target directory and Gradle build/; its output is the target directory, so Gradle up-to-date can skip a run when sources are unchanged. The generator still deletes and rewrites the target when the task runs.

The plugin adds the library as a default dependency of the sitePublisher configuration, at the same version as the plugin (org.podval.tools:org.podval.tools.publisher:0.4.0). Gradle defaultDependencies run only when that configuration is still empty, so a consumer can pick another published library version by declaring sitePublisher themselves (there is no site.publisherVersion). Do not add a second coordinate alongside the default — a non-empty configuration replaces the default entirely:

dependencies {
  sitePublisher "org.podval.tools:org.podval.tools.publisher:${providers.gradleProperty('sitePublisherVersion').get()}"
}

The supported pair is the same version as the plugin. An older library may not accept flags the plugin always passes; a newer one is usually fine. A local settings-body includeBuild of this repo substitutes the library module for whatever version string was requested.

Site workflows: checkout, setup-java 25, setup-gradle, cache ~/.gradle/ms-playwright, ./gradlew generateSite (or :docs:generateSite), upload _site / docs/_site, deploy Pages. Pin the plugin version in build.gradle. Known sites apply the plugin and call ./gradlew generateSite (or :docs:generateSite). GitHub Actions can still call dubinsky/site-publisher/.github/actions/generate@v0.1.0 for repos that do not apply the plugin (pin the action tag; optional version input overrides that action checkout’s gradle.properties).

Sites using this publisher:

Dogfooding an unreleased publisher

This repo already includeBuild`s sibling `../xml and ../../OpenTorah/opentorah.org when those checkouts exist (-PxmlDir= / -PopentorahDir=).

A site that applies org.podval.tools.site-publisher substitutes a local checkout with two includeBuild`s of the same directory: `pluginManagement (plugin id; pass a String, not a File) and the settings body (library child org.podval.tools:org.podval.tools.publisher). pluginManagement alone does not substitute the library now that it is not the included-build root. Do not use mavenLocal().

pluginManagement {
  repositories {
    mavenCentral()
    gradlePluginPortal()
  }
  final String publisherDir =
    providers.gradleProperty('sitePublisherDir').getOrElse('../site-publisher')
  if (file(publisherDir).isDirectory() && file("${publisherDir}/settings.gradle").isFile()) {
    includeBuild publisherDir
  }
}

// after plugins { }:
final File publisherDir = file(
  providers.gradleProperty('sitePublisherDir').getOrElse('../site-publisher')
)
if (publisherDir.isDirectory() && new File(publisherDir, 'settings.gradle').isFile()) {
  includeBuild(publisherDir)
}

Override with -PsitePublisherDir=. OpenTorah / MathWorlds defaults stay ../../Podval/site-publisher. When the directory is missing (CI), Gradle resolves the plugin and the library from Maven Central.

Without a site-side plugin, generate from this repo: ./gradlew run --args="/path/to/source-directory".

Publish

./gradlew nmcpZipAggregation
./gradlew publishAggregationToCentralPortal

The zip includes the library (org.podval.tools:org.podval.tools.publisher), the plugin implementation (org.podval.tools:site-publisher-plugin), and the plugin marker (org.podval.tools.site-publisher:org.podval.tools.site-publisher.gradle.plugin). Same version.

Site-header links are listed in _site_config.yml in display order (source path or site path, same resolution as home):

header-pages:
  - notes/index.md
  - posts
  - tags
  - graph

Omitted or empty means no header links (up/prev/next and format icons still appear). Unknown entries are recorded on the Errors page.

Optional facsimiles-url is the base URL of a facsimile JPEG site (trailing slash as in the collector tei@facsimilesUrl). Omitted means none. See Facsimiles.

facsimiles-url: https://storage.googleapis.com/facsimiles.alter-rebbe.org/

Optional tei-default-calendar is julian or gregorian (default gregorian). TEI date/@when numbers are in that calendar unless the element has calendar="#julian". See TEI dates.

tei-default-calendar: julian

Optional named-windows (default off) reuses four collector window names so collection, names, transcription, and facsimile can sit in separate tabs. Internal links target the destination viewer; each page sets window.name. Other sites should leave it off.

named-windows: true

Summary

This here is a static site generator. It:

Supported markup constructs:

Opinionated

Since I am writing this site generator for my own use, it is going to be opinionated:

  • everything is written in Scala

  • no plugins

  • no SCSS

  • no template languages

  • no layouts

TODO expound

Markup

Recognized by extension:

  • Markdown (.md)

  • AsciiDoc (.adoc)

  • HTML (.html)

  • TEI (.xml with a TEI root: TEI, store, collection, person, place, org)

  • DocBook (.xml with a DocBook root: article, book, chapter, appendix, part, set, preface, refentry, topic)

Frontmatter

Internal (YAML between --- lines in the markup file) or external (same basename, .yaml / .yml). Do not use both.

Obsidian treats internal YAML as Properties; it stores extra keys but does not interpret this publisher’s fields (chunk, bibliography, csl, …).

Set title in front matter only when it differs from the file name. Both the file name and the title are used for link resolution. If the document also has a title (h1, AsciiDoc =, TEI/DocBook title), it must match the front matter title; a mismatch is recorded on the Errors page.

Useful fields:

  • chunk / chunk-depth — see Chunking

  • bibliography / csl / lang — see Bibliography

  • toc-depth

  • description, author, lang — page <head> (SEO); site values in _site_config.yml are the fallback. Site lang also selects store and selector display names (publisher Selector.xml, TEI <name lang>); omitted is en.

  • tags, categories (Categories), aliases

  • icon, permalink, post, date

  • pdf - see PDF

  • asset — copy the markup file as an asset (no dialect processing, no site chrome). Standalone sidecar only: name.yml / name.yaml next to name.html (or .md, .adoc, .xml, …) with asset: true. The sidecar is not published. Internal --- asset: true is an error (the file stays markup). A directory index so marked stays a directory page (parent/children) but write copies the file (no listing, no chrome).

permalink (absolute, e.g. /short) and aliases (relative or absolute) each create a Refresh page at that URL. A permalink that does not start with / is an error (/errors.html#permalink) and creates no alias. When the page is a directory (an index or a TEI store / collection beside its folder), /permalink/child and wiki [[permalink/child]] resolve to child under that directory — the short name is a path prefix, not only a leaf bookmark. A remainder under a non-directory alias does not resolve. A last-dot suffix in a href that is all digits is part of the name (255.2), not a file extension.

Categories

categories is a list of links to authored pages. A value is a wiki link ([[Books]] or [[Books|label]]) or a bare name (Books). It resolves the same way as a wiki link in the body. The target is a normal page: its prose stays, and the generator appends the pages that name it, sorted by title. A value that resolves to nothing is an unresolved link on the Errors page. No category page is created for it. ![[Note.base]] and ![[Note.base#View]] are Obsidian queries and are omitted, with no error. The member page header lists each category beside its tags. A wiki link in the body is a backlink. It does not put the page in that list. A page may name several categories.

Page Names

Jekyll relies on page titles being given explicitly in the front matter and ignores the file names; I had to use a TODO Obsidian plugin to set the titles from the file names to communicate them to Jekyll; as a result, pages where the title is different from the file name (e.g., index pages) had to be marked with TODO in their front matter to disable this plugin from setting the front matter title.

My plugin uses both the front matter title and the file name for link resolution, so I can remove this Obsidian plugin and redundant front matter titles which are the same as the file name, and set the title in the front matter only when it is different from the file name.

Obsidian uses file names exclusively and ignores the front matter title; for link resolution to use front matter titles also, TODO Obsidian plugin needs to be installed.

Wiki links ([[…]], including [[note|text]] and [[#id]]) are native Obsidian. Block ids: ^id at the end of a paragraph, or on its own line after a list, table, quote, or code block (blank line before the id). Link with [[note#^id]].

Obsidian embeds of pages (![[note]], ![[note#heading]], ![[note#^block]]) copy that authored region into the host (any markup on the target). The copy is a boxed aside.transclusion whose header is a link to the source (Notes Title or Notes Title § heading). The gear Settings checkbox Seamless transclusions hides the box and header (html.transclusion-clean). A missing heading or block is an error (not a silent whole-page embed). Footnotes inside a transcluded region are merged and numbered with the host. Media embeds (![[file.pdf]], images, audio, video) are unchanged. AsciiDoc include:: and TEI/DocBook xi:include are not this feature.

Wiki links in Obsidian:

  • are they case-sensitive?

  • do they take into account file name, title property, document title?

  • how do they look in the file if name reference is ambiguous?

  • can they be split between the lines?

  • do they support any kind of agglutination?

Entities

A TEI file whose root is person, place, or org is an entity; the kind is that root. The entity id is the file name without .xml (name/alter-rebbe.xml → alter-rebbe). The file may sit anywhere in the tree.

persName, placeName, and orgName with a non-empty @ref resolve to an entity when the kinds match and @ref is exactly that id (not a path, not the displayed name, not a title). A kind mismatch or unknown id is an unresolved internal link and is recorded on the Errors page. Without @ref, the name is left as-is (not a link) and is recorded on the Errors page (name without @ref, /errors.html#no-ref). That check runs on the first parse of TEI documents and of store/collection title / abstract / body — not on the name elements that define an entity file. TEI unclear is /errors.html#unclear. An entity file whose name is not the first name with spaces turned into underscores is /errors.html#misnamed-entity. Each kind on the Errors page is a fragment (AmbiguousTitle → #ambiguous-title). When more than one kind is present, a TOC links to those fragments.

<persName ref="alter-rebbe">Алтер Ребе</persName>
в <placeName ref="Вильна">Вильне</placeName>
<orgName ref="Виленский_кагал">виленский кагал</orgName>

Wiki [[alter-rebbe]] still finds the entity page by file name. Entity refs do not title-walk, so a markdown page titled alter-rebbe does not steal persName ref="alter-rebbe".

A TEI entityLists file is a catalog of buckets, not a store include list. Beside a directory (name.xml next to name/) it is that directory’s index (/name/index.html). Each listPerson, listPlace, or listOrg is a bucket:

  • the list kind is the element (listPerson → person files);

  • @n is the list id (/name/{n}.html when the catalog has two or more non-empty lists);

  • @role present: entities of that kind whose root @role matches, anywhere in the site; omitted: entities of that kind with no role;

  • <title> on the list is the list page title; a direct child <title> of entityLists is the catalog page title.

Empty lists are omitted. Members are every matching entity on the site, sorted by file name; the link text is the first persName / placeName / orgName. The entity page <h1> / <title> is that first name, not the file name. Document backlinks on an entity page still come from @ref in other documents (grouped by TEI collection when the source sits under one); the catalog itself is not a backlink source.

When two or more lists are non-empty, the catalog page is a TOC of those lists (links to the list pages; members are not inlined). A catalog with a single non-empty list is that list (no {n} subpage). Several entityLists files are independent catalogs.

The catalog directory is an identity prefix in collection-aliases.json (/name → /name/index.html, /name/{id} → /name/{id}.html), the same Worker mechanism as collection aliases. There is no synthetic /name.html dump.

<entityLists>
  <title>Имена</title>
  <listPerson n="jews" role="jew"><title>Жиды (они же Евреи)</title></listPerson>
  <listPerson n="unknown"><title>Неизвестно кто</title></listPerson>
  <listPlace n="places"><title>Места</title></listPlace>
</entityLists>

Directories

Root / is index.md / index.adoc, or a synthesized listing of that folder.

To open at a specific page instead, set home in _site_config.yml to an absolute site path. If the document is chunked, use the TOC (P/index.html):

home: /book/book/index.html

/index.html is then a meta Refresh to that page (not an HTTP 301). Do not also author index.md / index.adoc. Other root files stay on the site; they are just not listed at /.

Collection short names (collector site.xml <alias n="rgada" to="…"/>) are an optional alias on the TEI store / collection root, not site config and not front-matter permalink:

<collection n="3140" directory="3140" alias="rgada">

alias is the URL prefix (/rgada, /rgada/003). The publisher does not write a Refresh file at /rgada.html; Pages.find, emitted hrefs (including canonical/sitemap), and serve() rewrite /rgada/003 to the written …/3140/003.html. Inbound old collector facsimile URLs /rgada/facsimile/003 rewrite the same way (see Facsimiles). Front-matter permalink / aliases remain for leaf Refresh pages.

Generate also writes collection-aliases.json (segment from → collection directory to) for a Cloudflare Worker. The Worker is an internal rewrite (not a 301): longest slash-delimited prefix, then .html unless the request already has an extension (keep .xml); inbound /alias/facsimile/P becomes {to}/P/facsimile.html. Entity-lists catalogs are an identity prefix (/name → /name/index.html). Prefix routes (hostname/rgada*, hostname/name*, …) so assets skip the Worker. Deploy with .github/actions/deploy-alias-worker. Canonical entity hrefs stay the file path (/name/{id}.html when the file lives under name/).

A directory permalink (an index or TEI store/collection) is also a prefix: /short/child resolves under that directory, and emitted hrefs use the short path.

A TEI store file in the source root (archive.xml next to archive/) also gets two synthetic indexes, named from the file: {name}-collections.html (nested tree of child stores and collections, collector /collections) and {name}-index.html (flat list of descendant collections, collector /). Titles come from this project’s Selector.xml catalog in the site lang (omitted lang is en): the store’s by/@selector title for the tree (archive → Архивы when lang is ru) and the case selector title for the flat list (Дела). When there is exactly one such root store, /collections rewrites to the tree page (no Refresh file; Worker table includes it). Point header-pages at {name}-collections and home at /{name}-index.html when those should be chrome / the landing page.

A TEI store or collection file beside a directory (dir.xml next to dir/) is that directory’s index. A collection index sets class="wide" on <html> so the content .wrapper is 1800px; header and footer stay the default width. Store indexes and documents under the collection do not. xi:include/@href (relative to the store file) is an ordered child reference, not XInclude: the target is not inlined. Store children are listed in include order (name: title). A collection body is table.collection-index (Описание, Дата, Кто, Кому, Язык, Документ, Страницы, Расшифровка), not a page list. Optional <part from="000"> title rows split the table; pageType is manuscript (default: 000 / 000об) or book (numeric). A file named {base}-{xx} (two-letter language, dash at length 3) is a translation of {base}: not a table row; Язык is he ru with ru linking to the translation; the original document’s site header has [ru]; prev/next skip translations. text/@xml:lang must match the suffix. The page header (.store-header) holds ancestor path lines and this node’s selector + name + title as <l>`s, then `abstract/body and the by/@selector label. A TEI document under a collection also gets a document-header table from teiHeader (Описание, Дата, Кто, Кому, Расшифровка). The Дата cell (and in-text dates) get calendar hover tables; see TEI dates. Selector labels (category → разряд when site lang is ru) come from this project’s Selector.xml (page.Selectors). Documents under a store or collection use the same header. A scanned file under that directory that is not named in the includes is an error (translations are not). Selector hops in the href (book/ in books/book/derzhavin.xml) are URL segments only — they are not pages, and up skips them. Missing pb@missing photos are listed under the table. Collection Страницы links to pb ids in the transcription (#p{n}); in-text pb links to the facsimile viewer when facsimiles-url is set.

  • directory structure

  • navigation

  • header pages

  • icons

TEI dates

A TEI <date> with @when, @from, @to, @notBefore, or @notAfter gets a hover table. .. in any of those attributes is InvalidDate, and the element is left as authored. Each value is YYYY, YYYY-MM, or YYYY-MM-DD. Month names follow the site lang. The column headers stay English.

A full day on @when is one column, Date. A year or a month on @when is both ends of that year or month. A year or a month on @from or @notBefore expands to the first day of that end only. A year or a month on @to or @notAfter expands to the last day of that end only. @from equal to @to, or @notBefore equal to @notAfter, stays two columns, not Date.

| What you write | Hover columns | |---|---| | @when on a full day (YYYY-MM-DD) | Date | | @when on a year or a month | From, To | | @from and @to | From, To | | @notBefore and @notAfter | Not before, Not after | | @from only | From | | @to only | To | | @notBefore only | Not before | | @notAfter only | Not after | | @from and @notAfter | From, Not after | | @notBefore and @to | Not before, To |

These combinations are InvalidDate, and the element is left as authored:

  • @when together with @notBefore, @notAfter, @from, or @to

  • @from together with @notBefore

  • @to together with @notAfter

@from with @notAfter, and @notBefore with @to, are allowed.

The numbers are Gregorian unless the element has calendar="#julian", or _site_config.yml has tei-default-calendar: julian and @calendar is omitted. Any other @calendar value is Gregorian. This disagrees with the TEI Guidelines: the attributes are not rewritten to Gregorian; @calendar selects how the numbers are read, and it still describes the prose.

The collection index and document-header Дата cell show @when as the visible text when @when is present. Otherwise the cell shows the element’s own text, still inside <date>, so the tip can attach. In-text dates always keep their text. A <date> with none of @when, @from, @to, @notBefore, and @notAfter is unchanged.

TEI gap

A TEI <gap reason="…"> with a non-empty @reason gets a hover tip with that text (same family as dates). Missing or blank @reason leaves the element as-is. Existing tei.css rules for gap (lost / illegible brackets) still apply.

Facsimiles

When _site_config.yml has facsimiles-url, each TEI document with at least one non-missing pb gets a viewer at P/facsimile.html (published /alias/P/facsimile.html when the collection has alias). Photos scroll in a pane under the header (the window is the size; there is no inner resize box). Old collector bookmarks /alias/facsimile/P (and /alias/facsimile/P.html) still resolve to that viewer, including translation /alias/facsimile/P-xx (shared original). #p{n} fragments are unchanged. New pages emit only the new URL.

The JPEG for pb@n is {facsimiles-url}{directory of the TEI file}/{n}.jpg. A trailing slash on the config value is optional. pb@facs overrides that URL. Missing pb@missing pages are omitted from the scroller.

In the transcription, <pb n="000-1"/> is an images icon (a.pb, id="p000-1") linking to the viewer fragment, target="facsimile" (facsimileViewer when named-windows is on). The site-header format icon does the same without a fragment. Collection Страницы is page numbers on the transcription #p{n} only. Each photo in the viewer links back to that pb in the transcription (target="text" / textViewer). {base}-{xx} translations share the original’s viewer.

Omitted facsimiles-url: no viewer pages; pb keeps its id (collection Страницы still works) and has no href.

Blog Posts

A page is a blog post in one of three ways:

  • A file in _posts/ or _drafts/ named YYYY-MM-DD-title is published at /YYYY/MM/DD/title.html.

  • Front matter post: true keeps the file path and adds a Refresh page at /YYYY/MM/DD/{file name}.html. date is required. post-title replaces the file name in that path.

  • Front matter permalink: /YYYY/MM/DD/title does the same for any source file name. The Refresh page is that permalink.

/posts lists each real post once. The Refresh file is not a second entry.

Front matter description is the post teaser: Atom <summary> and, when present, a p.post-excerpt on every /posts batch. Not auto-cut from the body. Site description is not used here (SEO still falls back to it).

Paging

The /posts listing only. In _site_config.yml:

paginate-posts: 10

Omitted or less than 1 is off. Page 1 stays /posts.html; further batches are /posts/2.html, /posts/3.html. A nav.pagination sits under the list.

SEO

No extra config. Each HTML page <head> gets a document title (Page | Site title; home is just the site title), meta name="generator" pointing at this publisher, description, author, canonical URL, Open Graph, a Twitter summary card, and JSON-LD (WebSite / BlogPosting / WebPage).

Taken from existing fields: site title, description, url, author, lang, social.twitter; page title, description, author, lang, date (and git/modified_time for modified). No image, Facebook, or webmaster keys.

Graph

Off by default. When _site_config.yml has graph.enabled: true, generate writes /graph.json and /graph.html. /graph.html is a pan/zoom map of canonical authored pages and the resolved internal links between them (Cytoscape.js, loaded only on that page). Click a vertex to open the page (named-windows targets apply when that option is on).

Vertices are authored pages with a source file: notes, posts, TEI documents, entities, stores, and entity-list catalogs. Aliases, chunks, facsimile viewers, PDFs, assets, and synthetics (/errors, /tags, /posts, …) are not vertices; links to those views count as links to the owning page. Edges come from existing backlinks (kind: link, including TEI @ref) and, unless turned off, transclusions (kind: transclude). include-transclusions defaults to true; omit it or set it explicitly. Unresolved links stay on the Errors page, not as graph nodes.

Do not name a collection @alias="graph": that path is reserved for the graph page.

graph:
  enabled: true
  include-transclusions: true
  exclude-path-prefixes:
    - days

exclude-path-prefixes match the source path (days/2020-10-23.md, archive/lvia/…), not the published URL. Daily notes publish as /YYYY/MM/DD/index.html; TEI collection aliases publish as /lvia1799-2/…. A notes vault typically enables the graph and excludes the Obsidian daily-notes folder (days when that is the configured folder). A large TEI archive should leave the graph off, or exclude archive to keep only names and notes.

Add graph to header-pages if you want a header link (there is no automatic icon). Pipeline notes belong in the Site Publisher design note, not here.

Off by default. When _site_config.yml has check-links: true, generate requests each author-written http and https URL and lists failures on the Errors page (#broken-link).

check-links: true

Checked: links (a@href) and media (img / video / audio / source @src, PDF object@data), including links inside footnotes and bibliography entries. An absolute URL whose host is the site url is a spurious self-link and is not requested. mailto:, tel:, javascript:, and data: are not requested. URLs inside code, and URLs only reached by transcluding another page, are checked on the page that authors them. The same URL is requested once per generate. Each page that contains it gets one error.

The request is HEAD. On 403, 405, or 501 it is retried once with GET. Redirects are followed. A status of 400 or higher, or a timeout or connection failure, is a broken link. The error text includes the status or the failure message.

Footer social links, the license link, and facsimile JPEGs are not checked. Generation still finishes. The Errors page is the report, as for unresolved links. Timeouts are 5 seconds to connect and 10 seconds per request.

Chunking

Split a document into per-section HTML pages (any markup, not only AsciiDoc).

chunk: true
chunk-depth: 2

P.html is the unsplit document (wiki lands there). The split view is P/index.html (P/ on a static host): preamble plus TOC, then P/Section.html for each chunk. File hrefs to P/index.html stay there; they are not rewritten to P.html. No directory listing is synthesized for P/.

When both views exist, the site header has icon links between them (list for chunked, file for one-page). Chunked TOC shows ancestors plus the current subtree; sibling branches are titles only.

Table of contents

A page with chunk: true or toc-depth gets a generated TOC at the top if none is marked in the body. A placeholder in the body is replaced with that TOC (the first one only) even without those keys; toc-depth still controls how deep it goes (default 2).

Markdown

Kramdown {:toc} on a list. The list (seed item included) is replaced. Ordered or unordered:

* this list is replaced with the TOC
{:toc}

## Alpha

or 1. seed / {:toc} for a numbered outline.

A paragraph whose only text is [TOC] (any case) is the same placeholder (Typora, GitLab). It must be a top-level paragraph: not a link ([TOC]: /url then [TOC]), not in a list, quote, or code span.

HTML

<div class="toc-placeholder"></div>.

AsciiDoc, TEI, and DocBook have no in-body placeholder; use chunk / toc-depth.

PDF

Front matter pdf: true prints P.html to P.pdf (Chromium). The site header on P.html (and on chunked pages, if any) gets a PDF icon linking to it. The site header is hidden in print, so the PDF itself does not include those icons.

Tables

Markdown tables via FlexMark Tables Extension (GFM; Obsidian supports them natively).

Markdown
| A | B |
|---|---|
| 1 | 2 |

AsciiDoc tables via Asciidoctor.

AsciiDoc
|===
| A | B

| 1 | 2
|===
TEI
<table>
  <row role="label"><cell>A</cell><cell>B</cell></row>
  <row><cell>1</cell><cell>2</cell></row>
</table>
DocBook

CALS row / entry (optional tgroup); morerows becomes rowspan.

<table>
  <tgroup cols="2">
    <thead><row><entry>A</entry><entry>B</entry></row></thead>
    <tbody><row><entry>1</entry><entry>2</entry></row></tbody>
  </tgroup>
</table>
HTML
<table>
  <thead><tr><th>A</th><th>B</th></tr></thead>
  <tbody><tr><td>1</td><td>2</td></tr></tbody>
</table>

Task Lists

Unchecked [ ] and checked [x] items in a list. The static site shows disabled checkboxes.

Markdown

GFM / FlexMark Task List Extension (Obsidian supports them natively):

- [ ] open
- [x] done
AsciiDoc

native checklists:

* [ ] open
* [x] done

DocBook has no native task-list markup.

HTML

the IR shape (already a task list; no conversion):

<ul class="task-list">
  <li class="task-list-item">
    <input type="checkbox" class="task-list-item-checkbox" disabled> open
  </li>
  <li class="task-list-item">
    <input type="checkbox" class="task-list-item-checkbox" disabled checked> done
  </li>
</ul>

Footnotes

Hovering a footnote number shows the note; the numbered list is at the end of the (chunk of the) page. A note written with no space after the preceding word or element stays adjacent in the HTML.

A marker in a table cell is published as a lowercase Latin letter (a, b, …, z, aa) with the list immediately under that table (div.table-with-notes). Markdown still defines those bodies at the end of the source file; the published page moves them. AsciiDoc [1] is written at the call site (including in a cell); the published page still lists those notes under the table. Collector collection-index and document-header tables stay in the page arabic series.

A footnote inside another footnote is one level of letters a, b at the end of the parent note. There is no nested hover tip: the outer tooltip shows inner letter markers, not a nested list. Markdown and AsciiDoc get nested notes only if the processor emits an inner stub. TEI note place="end" inside another place="end" does.

TEI still converts only place="end". place="foot" and unplaced note are unchanged. @n, when non-blank, is the marker at the call site and in the list. Order and fragment ids stay on the generated series. The same marker may appear on more than one note. Obsidian inline ^[…] footnotes are unsupported.

AsciiDoc

inline note, or a named definition reused with an empty macro: A [2] in a table cell is written at the call site; the published page lists it under the table.

See this footnote:[A note.]
and again footnote:fn[Same note.] later footnote:fn[].
Markdown

FlexMark: use [^id] in the text; define [^id]: later. Obsidian supports that form natively (Help: Footnotes). Obsidian also has inline ^[…] footnotes; this publisher does not. A [^id] in a table cell uses the same [^id]: body at the file end; the published page moves it under the table.

See this [^note] and again [^note].

[^note]: A note.
TEI

definition and use are the same element at the point of reference (@n, when non-blank, is the marker):

<p>See this<note place="end">A note.</note>.</p>

A non-blank @n is that marker. The next note with no @n is still numbered in series.

<p>See this<note place="end" n="*">A note.</note> and <note place="end">the next.</note>.</p>

A place="end" inside another place="end" is a nested letter on the parent note:

<p>See this<note place="end">Outer<note place="end">Inner.</note></note>.</p>
DocBook

<footnote> at the call site (xml:id optional; non-blank label is the same marker). Reuse with <footnoteref linkend="id"/>. A <footnote> inside another footnote’s <para> is nested the same way if the converter emits an inner stub.

<para>See this<footnote><para>A note.</para></footnote>.</para>
HTML

IR stubs (same footnote-correlation-id on the call site and the body):

<p>See this<span class="footnote-link" footnote-correlation-id="note"></span>.</p>
<span class="footnote" footnote-correlation-id="note">A note.</span>

Glossary

A glossary is a description list marked as such. Term links get hover definitions (print inlines them in parentheses).

AsciiDoc

[glossary] on the description list; term ids from . Reference a term with [id].

See <<posuk>>.

[glossary]
[[posuk]]posuk:: verse
Markdown

PHP Extra definition lists (https://michelf.ca/projects/php-markdown/extra/#def-list) via FlexMark Definition List Extension. Mark the list with {:.glossary} after it (Kramdown IAL, same family as Table of contents). Without that marker the list is not a glossary (no term ids, no tooltips). The marker also separates adjacent lists.

Obsidian does not render PHP Extra definition lists natively (Markdown Guide: Obsidian). Community plugin Pandoc Extended Markdown renders Pandoc/PHP Extra term / : definition lists. {:.glossary} is Kramdown IAL, not Obsidian; it shows as a paragraph in the note. [[#id]] heading links are native Obsidian.

+

See [[#posuk]].

posuk
: verse

{:.glossary}

Term id is the id on <dt> if present, otherwise the term text with spaces turned into hyphens (Alter Rebbe → Alter-Rebbe). Reference a term with [[#id]] or [text](#id).

A lone : after a Markdown term is not a reliable empty definition: FlexMark can take the next term as that definition. Write a real definition line, or omit the term.

TEI

<list type="gloss"> (or type="glossary") of <label> / <item>. Term id from xml:id on the label, else on the item, else the label text with spaces turned into hyphens. Reference a term with <ref target="#id"> or <term ref="#id">.

<p>See <ref target="#posuk">posuk</ref>.</p>
<list type="gloss">
  <label xml:id="posuk">posuk</label>
  <item>verse</item>
</list>
DocBook

<glosslist> (or <glossary>) of <glossentry> / <glossterm> / <glossdef>. Term id from xml:id on the entry, else the term, else the term text with spaces turned into hyphens. Reference a term with <link linkend="id">. <variablelist> is a description list, not a glossary.

<para>See <link linkend="posuk">posuk</link>.</para>
<glosslist>
  <glossentry xml:id="posuk">
    <glossterm>posuk</glossterm>
    <glossdef><para>verse</para></glossdef>
  </glossentry>
</glosslist>
HTML

the IR shape (class="glossary"; each term is class="glossary-item" with an id):

<p>See <a href="#posuk">posuk</a>.</p>
<div class="glossary">
  <div class="glossary-item" id="posuk">
    <dt>posuk</dt>
    <dd>verse</dd>
  </div>
</div>

Bibliography

Two kinds, per document (no site-wide file or style). They can appear in the same document.

External (.bib + CSL)

Front matter:

  • bibliography — path to a .bib file, relative to this source file (.bib files are not copied to the site; un-ignore with _site_ignore if you want one published)

  • csl — CSL style name from the bundled styles (for example apa, ieee)

  • lang — CSL locale; falls back to the site lang, then en-US

Both bibliography and csl are required. Only works cited on the page appear in the list. In-text cites become links to #bibl-{key} on the matching list entry (first key if a cite names several). Unknown keys are highlighted and recorded as a page error.

AsciiDoc

no asciidoctor-bibliography gem. Also cite:key[] / cite:key[p. 12] (standard inline-macro form).

- `cite:[knuth79]`, `cite:[knuth79, p. 12]`, `cite:[knuth79, lamport94]` — parenthetical
- `cite:knuth79[]`, `cite:knuth79[p. 12]` — same, standard inline-macro form
- `citenp:[lamport94]` — narrative (author in the sentence)
- `bibliography::[]` — list placeholder (otherwise the list is appended)
Markdown

Pandoc citation syntax (no extra FlexMark extension). me@host.com is left as an email. Obsidian does not format [@key] natively. Community plugin Pandoc Reference List recognizes Pandoc citekeys in the note (sidebar; optional inline). Citations inserts [@key] from a .bib file.

- `[@knuth79]`, `[@knuth79, p. 12]`, `[@knuth79; @lamport94]` — parenthetical
- `@knuth79` — narrative (`me@host.com` is left as an email)
- `[-@lamport94]` — suppress author
- `:::bibliography` — list placeholder
See [@knuth79] and [@knuth79, p. 12] and [@knuth79; @lamport94]
DocBook

<citation>key</citation> (optional locator after a comma). biblioref @linkend that is not a native entry id is the same. Empty <bibliography/> is the list placeholder (otherwise the list is appended).

<para>See <citation>knuth79</citation> and <citation>knuth79, p. 12</citation>.</para>
<bibliography/>
HTML

the IR shape:

<p>See
<span class="citation" data-mode="parenthetical"><span class="citation-item" data-key="knuth79"></span></span>
and
<span class="citation" data-mode="parenthetical"><span class="citation-item" data-key="knuth79" data-locator="p. 12"></span></span>.</p>
<div class="bibliography"></div>
TEI

@cRef on ref / ptr is the citeproc key. @n is an optional locator. Empty <div type="bibliography"/> is the list placeholder (otherwise the list is appended).

<p>See <ref cRef="knuth79"/> and <ptr cRef="lamport94" n="p. 12"/>.</p>
<div type="bibliography"/>

Internal (in-document list)

Authored entries stay on the page (not cited-only). Links to an entry id get a hover tip. Entry ids are as written (#knuth79); they do not use the #bibl- prefix of the CSL list, so the same key can be cited both ways in one document.

AsciiDoc

[bibliography] list with [[[id]]] / [[[id,xreftext]]] anchors; cite with [id].

See <<knuth79>> and <<lamport94>>.

[bibliography]
* [[[knuth79]]] Knuth, Donald E. _TeX and Metafont_. 1979.
* [[[lamport94,Lamport 94]]] Lamport, Leslie. _LaTeX_. 1994.
TEI

ref / ptr @target the xml:id of a bibl / biblStruct in listBibl (not in teiHeader). Empty ptr shows @n or the id. cit with a quote is a quote, not this. A bare @target that matches a listBibl id is treated as #id.

<p>See <ref target="#knuth79">Knuth 1979</ref>
and <ptr target="#lamport94"/>.</p>
<listBibl>
  <bibl xml:id="knuth79">Knuth, Donald E. <title>TeX and Metafont</title>. 1979.</bibl>
  <biblStruct xml:id="lamport94">
    <monogr><author>Lamport, Leslie</author><title>LaTeX</title></monogr>
  </biblStruct>
</listBibl>
DocBook

<bibliography> with biblioentry / bibliomixed; cite with <link linkend="id"> or <biblioref linkend="id"/> (empty biblioref shows @xreflabel or the id).

<para>See <link linkend="knuth79">Knuth 1979</link>
and <biblioref linkend="lamport94"/>.</para>
<bibliography>
  <biblioentry xml:id="knuth79">Knuth, Donald E. TeX and Metafont. 1979.</biblioentry>
  <bibliomixed xml:id="lamport94">Lamport, Leslie. LaTeX. 1994.</bibliomixed>
</bibliography>

Markdown has no native in-document bibliography list.

Code Blocks

Inline code and fenced blocks. A language class on <code> (language-scala, …) loads highlight.js; language-mermaid loads Mermaid instead.

Markdown

CommonMark inline code and fenced blocks. Obsidian supports both natively, including a language identifier (Help: Code). Fenced ```mermaid diagrams are native in Obsidian too (Help: Diagram); this publisher loads Mermaid for language-mermaid.

Use `map` in

```scala
xs.map(f)
```
AsciiDoc

inline code ; [source] listing (optional language).

Use `map` in

[source,scala]
----
xs.map(f)
----
TEI

<code> (tagdocs module) with @lang for the formal language (not xml:lang). A newline in the content is treated as a block and wrapped in <pre>. <eg> / <egXML> are not converted.

<p>Use <code lang="scala">map</code> in</p>
<code lang="scala">xs.map(f)
</code>
DocBook

<code> / <literal> inline; <programlisting> / <screen> / <literallayout> as a block. @language (not xml:lang) is the highlight.js tag.

<para>Use <code language="scala">map</code> in</para>
<programlisting language="scala">xs.map(f)
</programlisting>
HTML
<p>Use <code>map</code> in</p>
<pre><code class="language-scala">xs.map(f)
</code></pre>

Callouts

Numbered annotations on lines of verbatim (listing, literal, or source), plus a list of those annotations. Not a general footnote mechanism.

AsciiDoc

<1> (or <.> for automatic numbering) at the end of a line; repeat the numbers in a callout list below. Hide a marker in the raw source behind a line comment (// <1>, # <1>, <!--1-→).

[source,ruby]
----
require 'sinatra' (1)
----
<1> Library import
DocBook

<co> in a listing; <calloutlist> / <callout arearefs> below. @label on co is the marker number.

<programlisting>require 'sinatra'<co xml:id="co1" label="1"/></programlisting>
<calloutlist>
  <callout arearefs="co1"><para>Library import</para></callout>
</calloutlist>
HTML

the IR shape:

<pre><code>require 'sinatra' <span class="callout" data-value="1">1</span></code></pre>
<ol class="callout-list">
  <li>Library import</li>
</ol>

Admonitions

Typed notices (note, tip, warning, …). Not listing Callouts and not asides/sidebars.

Markdown

Obsidian core callouts (> [!type]; optional title; +/- to fold). No plugin.

> [!tip] Save time
> Use the shortcut.

> [!faq]- Hidden
> Secret.
AsciiDoc

NOTE: / TIP: / IMPORTANT: / CAUTION: / WARNING: (paragraph or [NOTE]==== block; optional .Title).

NOTE: Auxiliary information.

.Watch out
[WARNING]
====
Multi-paragraph body.
====
DocBook

note, tip, warning, caution, important (optional child title).

<note><title>Save time</title><para>Use the shortcut.</para></note>
HTML

the IR shape (details when folded):

<div class="admonition" data-type="tip">
  <div class="admonition-title">Save time</div>
  <p>Use the shortcut.</p>
</div>

Asides

Untyped auxiliary content next to the main flow. Not an admonition.

AsciiDoc

[sidebar] or a ** block (optional .Title).

.Optional Title
****
Auxiliary content.
****
Markdown

no native syntax; Obsidian has none. Write an HTML

block.

<aside>
<p>From Markdown.</p>
</aside>
DocBook

<sidebar> (optional child title).

<sidebar>
  <title>Optional Title</title>
  <para>Auxiliary content.</para>
</sidebar>
HTML

<aside> (IR adds class="aside" if missing); optional title:

<aside class="aside">
  <div class="aside-title">Optional Title</div>
  <p>Auxiliary content.</p>
</aside>

Quotes

Block quotations. Not Admonitions (> [!tip]) and not Asides.

AsciiDoc

[quote] (or a __ block); optional .Title; attribution and citation title as [quote, Author, Source].

.A title
[quote, Jefferson, Papers]
____
A little rebellion now and then is a good thing.
____
Markdown

a CommonMark blockquote (>). No attribution syntax (an em-dash line is still quote body). Obsidian > [!type] is an admonition, not a quote.

> A Markdown quotation.
DocBook

<blockquote> or <epigraph>; optional title and attribution. Inline <quote> stays <q>.

<blockquote>
  <title>A title</title>
  <para>A little rebellion now and then is a good thing.</para>
  <attribution>Jefferson</attribution>
</blockquote>
HTML

<blockquote> (IR adds class="quote" if missing); optional title and <footer> attribution (cite kept as-is):

<blockquote class="quote">
  <div class="quote-title">A title</div>
  <p>A little rebellion now and then is a good thing.</p>
  <footer class="quote-attribution">— Jefferson<br/><cite>Papers</cite></footer>
</blockquote>
TEI

<quote>; <cit> groups a quote with <bibl> (attribution). A <bibl> directly in <quote> is attribution too. <q> stays inline.

<cit>
  <quote>A TEI quotation.</quote>
  <bibl>Jefferson, <title>Papers</title></bibl>
</cit>

Strikethrough

Deleted or superseded text. HTML <del> (browser line-through).

Markdown

GFM ~text~ (Obsidian native; FlexMark strikethrough extension).

See ~~struck out~~ here.
AsciiDoc

text.

See [line-through]#obsolete phrase# here.
TEI

<del>.

<p>See <del>TEI struck</del>.</p>
DocBook

<emphasis role="strikethrough"> (or line-through). Default <emphasis> is <em>; role="bold" is <strong>.

<para>See <emphasis role="strikethrough">DocBook struck</emphasis>.</para>
HTML

<del> (IR); <s> and span/mark with class="line-through" become <del>:

<p>See <del>HTML struck</del>.</p>

Figures

A block image (or other graphic) with an optional caption. Inline images stay <img>. Local src values are resolved to a site path (/pixel.svg); a missing file is an error (missing asset) and the <img> gets unresolved-asset. Wiki ![[image]] also finds a unique file of that name anywhere in the source tree.

Markdown

a ![alt](src) on its own line; optional "title" is the caption. FlexMark does not turn Obsidian |WIDTH / |WIDTHxHEIGHT in the alt into img width/height (the size stays in alt). Obsidian ![[image]] embeds stay <img> (resolved later); |WIDTH or |WIDTHxHEIGHT on that wiki embed do become width/height.

![A square](pixel.svg "A Markdown figure")
![[pixel.svg|320]]
![[pixel.svg|320x240]]
AsciiDoc

block image:: (optional .Title). Inline src is not a figure. Width and height (image::file.png[alt, 320, 240] or width=320, height=240) are kept on the <img>.

.A figure caption
image::pixel.svg[A square]
TEI

<figure> with <graphic url> and optional <head> (caption). <graphic> becomes <img>.

<figure>
  <head>A TEI figure</head>
  <graphic url="pixel.svg"/>
</figure>
DocBook

<figure> (or <informalfigure>) with <imagedata fileref> and optional <title> (caption).

<figure>
  <title>A DocBook figure</title>
  <mediaobject>
    <imageobject>
      <imagedata fileref="pixel.svg"/>
    </imageobject>
  </mediaobject>
</figure>
HTML

<figure> (IR adds class="figure" if missing); optional <figcaption>:

<figure class="figure">
  <img src="pixel.svg" alt="A square"/>
  <figcaption class="figure-caption">HTML figure</figcaption>
</figure>

PDF embeds

A PDF file shown on the page. Not PDF generation of a page.

Markdown

Obsidian ![[file.pdf]]; optional #page=N and &height=H (pixels if H is digits). Alias ![[file.pdf|Title]] is the link label.

![[sample.pdf]]
AsciiDoc

no native block; write the HTML IR (passthrough ).

DocBook has no native PDF-embed markup.

HTML

<object type="application/pdf"> (IR wraps it) or the IR:

<div class="pdf-embed">
  <object data="sample.pdf" type="application/pdf" aria-label="sample.pdf">
    <a href="sample.pdf">Open PDF: sample.pdf</a>
  </object>
  <p class="pdf-embed-link"><a href="sample.pdf">Open PDF: sample.pdf</a></p>
</div>

The sibling link is always visible and is what remains in print (the inner viewer is hidden). data (and the fallback href) are resolved like other local assets; #page=N is kept.

Video

A local video file, or a YouTube/Vimeo player. Local files use HTML <video> (like <audio>). Local src is resolved like images; YouTube/Vimeo <iframe> URLs are left alone.

Markdown

Obsidian ![[file.mp4]] (also webm, ogv, m4v). Alias ![[file.mp4|Title]] is the fallback link label.

![[clip.mp4]]
AsciiDoc

video::file.mp4[] (optional .Title becomes a figure caption). YouTube/Vimeo: video::id[youtube] / video::id[vimeo].

video::clip.mp4[]
DocBook

<videodata fileref>. A YouTube/Vimeo URL becomes an iframe.

<mediaobject>
  <videoobject>
    <videodata fileref="clip.mp4"/>
  </videoobject>
</mediaobject>
HTML

<video src> (IR adds class="video" and controls); a YouTube/Vimeo <iframe> gets class="video-embed":

<video class="video" src="clip.mp4" controls="controls">
  <a href="clip.mp4">Open video: clip.mp4</a>
</video>

1. …
2. …

About

Generate, serve and publish websites

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages