Skip to content

[#24] SEO: declare the sitemap in robots.txt, add description and OpenGraph tags - #29

Merged
vharseko merged 2 commits into
OpenIdentityPlatform:masterfrom
vharseko:issue-24-seo-meta
Oct 1, 2026
Merged

vharseko merged 2 commits into
OpenIdentityPlatform:masterfrom
vharseko:issue-24-seo-meta

Conversation

@vharseko

@vharseko vharseko commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Fixes #24

Changes

  • antora-playbook.yml: site.robots is now the text of robots.txt instead of allow, so it also declares the sitemap index:

    User-agent: *
    Allow: /
    Sitemap: https://doc.openidentityplatform.org/sitemap.xml
    

    Antora writes this value verbatim, so a comment next to it asks to keep the Sitemap: host in sync with site.url. The file also gets the CDDL license header (Copyright 2024-2026 3A Systems, LLC.: the playbook has been written by 3A Systems since 2024).

  • supplemental-ui/partials/head-meta.hbs:

    • <meta name="description"> for pages without :description:. The default UI emits it only when the page sets :description:, and no page does today;
    • OpenGraph: og:type, og:site_name, og:title, og:description, og:url (the canonical URL; absent on the 404 page), og:image;
    • twitter:card = summary.

    og:image is the GitHub organization avatar, the same image www.openidentityplatform.org uses, so no image is added to the repository.

  • supplemental-ui/partials/meta-description.hbs (new): the generated description:

    • product pages: <page title> - <component title> documentation, Open Identity Platform, e.g. "Installing OpenAM Core Services - OpenAM: Access Management documentation, Open Identity Platform";
    • start page and 404: a description of the platform naming the four products.

Verification

Local Antora build of this branch:

  • robots.txt has the Sitemap: line;
  • all 285 pages have every tag exactly once; og:url is missing only on 404.html;
  • titles with special characters (Release Levels & Interface Stability, Standards, RFCs, & Internet-Drafts) are escaped correctly; no tags or unresolved entities are left in the attributes;
  • a page given a temporary :description: with quotes, & and <b> gets a single description tag (emitted by the UI) and an identical og:description, escaped the same way;
  • no new Antora errors.

@vharseko vharseko added documentation Improvements or additions to documentation enhancement New feature or request labels Sep 30, 2026

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The new tags are placed so that they never collide with what the default UI already emits.

  • The generated description is wrapped in {{#unless page.description}}, so a page that sets :description: keeps the single tag the UI writes (supplemental-ui/partials/head-meta.hbs:22-24).
  • og:url sits inside {{#with page.canonicalUrl}}, so the 404 page gets no og:url tag rather than an empty one (supplemental-ui/partials/head-meta.hbs:29-31).
  • Titles pass through {{{detag … attribute=true}}} before they go into attributes (supplemental-ui/partials/head-meta.hbs:27, supplemental-ui/partials/meta-description.hbs:18).

suggestion (non-blocking): No automated check covers the new head tags or robots.txt.

.github/workflows/publish.yml:3

publish.yml is the only workflow. It runs on push to master and on workflow_dispatch, never on pull_request, and package.json has no test script. The checks in the description (285 pages, each tag exactly once, og:url missing only on 404) were done once in a local build. A later template edit that drops or duplicates a tag would still build green and deploy, and so would a HEAD UI-bundle change (snapshot: true) that stops detag … attribute=true from escaping.

# .github/workflows/check.yml
name: Check site build
on: pull_request
jobs:
  build:
    runs-on: ubuntu-latest
    steps:
    - uses: actions/checkout@v4
    - uses: actions/setup-node@v4
      with:
        node-version: '24'
    - run: npm i antora
    - run: npm run build
    - name: Check robots.txt and head tags
      run: |
        grep -qx 'Sitemap: https://doc.openidentityplatform.org/sitemap.xml' build/site/robots.txt
        fail=0
        while IFS= read -r f; do
          grep -q 'http-equiv="refresh"' "$f" && continue
          for tag in '<meta name="description"' '<meta property="og:title"' '<meta property="og:description"' '<meta property="og:image"' '<meta name="twitter:card"'; do
            n=$(grep -cF "$tag" "$f" || true)
            [ "$n" -eq 1 ] || { echo "$f: $n x $tag"; fail=1; }
          done
        done < <(find build/site -name '*.html')
        exit $fail

Pin: deleting any of the new tags from head-meta.hbs, or the Sitemap: line from the playbook, turns this job red.


suggestion (non-blocking): The Sitemap: line in robots.txt repeats the host instead of following site.url.

antora-playbook.yml:22, :18

Antora writes a custom site.robots value verbatim (only trimEnd()), while sitemap.xml is built from the effective site URL. So npx antora --url https://staging.example antora-playbook.yml (or URL set in the environment, or a later edit of line 18 alone) produces a staging sitemap.xml with a robots.txt that still points to the production one. The output is correct today. Antora cannot generate the line itself, so a comment is the cheapest guard.

  url: https://doc.openidentityplatform.org
  # Antora writes robots verbatim: keep the Sitemap host in sync with site.url above.
  robots: |
    User-agent: *
    Allow: /
    Sitemap: https://doc.openidentityplatform.org/sitemap.xml

nitpick (non-blocking): The new header in antora-playbook.yml says "Portions Copyright"; it should say "Copyright".

antora-playbook.yml:13

At base this file had no license header and no copyright line, so 3A Systems is the sole copyright holder. "Portions copyright" is for a contribution to a file that someone else holds, as in head-meta.hbs (ForgeRock 2017). The new meta-description.hbs already gets this right.

# Copyright 2026 3A Systems, LLC.

nitpick (non-blocking): The new CDDL headers point to legal/CDDLv1.0.txt, which this repository does not have.

antora-playbook.yml:5, :9, supplemental-ui/partials/meta-description.hbs:6, :10

The head tree has no legal/ directory and no license file. The only legal/ paths are the JDK notices under openam/apidocs/legal/. The same boilerplate was already in head-meta.hbs at base, so this copies a precedent.


note (non-blocking): A change that the description does not mention.

  • antora-playbook.yml:1-13 — a new CDDL license header.

… tags

- antora-playbook.yml: robots.txt declares
  Sitemap: https://doc.openidentityplatform.org/sitemap.xml
- head-meta.hbs: generated <meta name="description"> for pages without
  :description:, OpenGraph (og:type, og:site_name, og:title,
  og:description, og:url, og:image) and twitter:card tags
- meta-description.hbs: the generated description

Fixes OpenIdentityPlatform#24
@vharseko

vharseko commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

Re: no automated check of the head tags and robots.txt — #28 adds the pull request check this asks for: .github/workflows/build.yml runs on pull_request to master, with npm ci and Node 24 action versions. It also replaces the UI bundle from GitLab HEAD (snapshot: true) with ui/ui-bundle.zip kept in the repository, so a UI change can no longer silently stop detag … attribute=true from escaping. What #28 does not do is the tag check itself; I will add it as a step of that build.yml once #28 and this PR are merged, rather than as a second workflow here: tracked in #37.

Re: Sitemap: host repeated — added the comment above robots:.

Re: "Portions Copyright" — changed to Copyright 2024-2026 3A Systems, LLC.: the playbook has been written by 3A Systems since 2024, so the range starts there. #28 adds the same header to this file (and to publish.yml), so I changed it there in the same way; the two PRs still merge without a conflict.

Re: legal/CDDLv1.0.txt — no change here: the same header is in hundreds of files on master, the documentation pages included. Adding the license file belongs to the whole repository, not to this PR.

Re: the header not mentioned in the description — added to the description.

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: Both round-1 points are fixed as asked, and every page still gets each head tag once.

  • antora-playbook.yml:13 now reads Copyright 2024-2026 3A Systems, LLC.. The file had no earlier holder and dates from 2024.
  • antora-playbook.yml:19 ties the hard-coded Sitemap: host to site.url.
  • The description fallback is gated on {{#unless page.description}} (supplemental-ui/partials/head-meta.hbs:22-24), which exactly complements the default UI's own tag. og:url sits inside {{#with page.canonicalUrl}} (:29-31).

@vharseko
vharseko merged commit 8b40321 into OpenIdentityPlatform:master Oct 1, 2026
vharseko added a commit to vharseko/doc.openidentityplatform.org that referenced this pull request Oct 1, 2026
…A copyright in header-less files

- lib/versions-from-tags.js: a failed tag fetch no longer claims the tag is
  missing on origin; the cause (missing tag, no network) is left to git's
  error on the next line
- lib/add-pdf-link.js, <product>/antora.yml: "Copyright 2024-2026 3A Systems,
  LLC." instead of "Portions Copyright" - these files had no earlier
  copyright holder, as for the playbook in OpenIdentityPlatform#29
vharseko added a commit that referenced this pull request Oct 2, 2026
…wler (#33)

Fixes #25, together with #28 — this PR does item 5, the migration to the
current DocSearch; items 1–4 are #28. The front end moves now, on the
existing index; the crawler stays the legacy scraper until Algolia
Crawler access is set up, for which this PR prepares the configuration.

## Changes
- **Front end: `@docsearch/js` 3.9.0** instead of `docsearch.js` 2
(vendored as before, `supplemental-ui/js/vendor/docsearch.min.js` and
`supplemental-ui/css/vendor/docsearch.min.css`):
- the header shows the DocSearch button, which opens the search modal;
shortcuts `Ctrl/Cmd+K` and `/`;
- same application, key and index (`doc_openidentityplatform`): the
records of the legacy scraper carry the fields the v3 front end reads
(`hierarchy.lvl0–6`, `content`, `type`, `url`), so no reindexing is
needed;
- `transformItems` drops the `#cookie-notice` anchor that the legacy
scraper gives to page titles and to the text before the first section
(the first id on the page, the cookie notice);
- rendered only when the build has both `ALGOLIA_SEARCH_API_KEY` and
`ALGOLIA_APP_ID` (the v3 client requires `appId`); otherwise the
`SITE_SEARCH_PROVIDER` branch (plain search input), which is kept, or no
search.

3.9.0 is the last 3.x; 4.x and 5.x add Ask AI and a side panel, which
this site does not use, and are 4–5 times larger (510 KB and 619 KB of
JS against 133 KB).
- **`docsearch/crawler-config.js`** (new): the same crawl for the
[Algolia Crawler](https://www.algolia.com/doc/tools/crawler/) — records
only from the product pages, as the `start_urls` of `config.json` (the
site home page is crawled for links only), the same exclusions (API
docs, generated references, and the older versions of #31) and the same
selectors for the hierarchy and the content. Not carried over, because
none of it changes the results: the version in lvl0 and the
`desc(version)` ranking (every component has `version: ~`), the
`component` attribute (nothing reads it) and `min_indexed_level`. It
writes a new index, `doc_openidentityplatform_v3`, so the live search is
not affected while it is checked. **Not run yet**: it needs Crawler
access on the Algolia application (for an open-source site, through the
DocSearch program).
- **`docsearch/readme.md`**: how the search is set up, and the steps to
move to the Crawler. The link it held before (how the index was set up
for the legacy scraper) is kept.

`publish.yml` and `config.json` are unchanged: the legacy scraper keeps
reindexing after every deployment.

## Verification
Local build with the public search key of the site, checked in a browser
(Playwright):
- the DocSearch button is rendered in the header, no console errors;
- `Ctrl+K` opens the modal; "session" and "replication" return results
from the live index, grouped by product; page titles link to the page
with no anchor, and no result URL ends with `#cookie-notice` (in the
live index every page-title record does);
- a build with only `ALGOLIA_SEARCH_API_KEY` has neither the button nor
the DocSearch bundle;
- the `pathsToMatch` pattern of `crawler-config.js`, checked with
micromatch, matches `/openam` and `/openam/` but not `/`, `/openicf/…`
or `/commons/…`;
- rebased onto `master` (with #29); merges with #30 and #32 without
conflicts (`git merge-tree`).
vharseko added a commit to vharseko/doc.openidentityplatform.org that referenced this pull request Oct 2, 2026
…arlier holder exists

index.adoc and openig/antora.yml had no copyright line before, so
"Portions" named a holder that does not exist. Use
"Copyright 2024-2026 3A Systems, LLC.", as antora-playbook.yml (OpenIdentityPlatform#29)
and OpenIdentityPlatform#31 do; the openig/antora.yml header is now identical to OpenIdentityPlatform#31's,
which removes the conflict between the two pull requests.
vharseko added a commit that referenced this pull request Oct 2, 2026
…leaseversion (#32)

Fixes #23

## Changes
- **`ROOT/modules/ROOT/pages/index.adoc`** (start page): under the
description of each product, a line with
- a shields.io badge of the latest release (e.g. *latest release
v16.1.3* for OpenAM), linking to `releases/latest`;
- *Releases and downloads* →
`github.com/OpenIdentityPlatform/<Product>/releases`;
  - *Docker images* → `hub.docker.com/r/openidentityplatform/<product>`;
- *Security advisories* → `…/security/advisories` for OpenAM, OpenDJ and
OpenIDM (OpenIG has no published advisories, so no link).
- **`supplemental-ui/partials/header-content.hbs`**: a **Downloads**
menu next to Projects, with the release pages of the four products, the
Docker Hub organization and the security advisories of OpenDJ, OpenAM
and OpenIDM.
- **`openig/antora.yml`**: `releaseversion: 5.2.4` removed — no page or
UI template uses it (checked in the sources and in the generated site).
- **CDDL headers** added to `index.adoc` and `openig/antora.yml` with
`Copyright 2024-2026 3A Systems, LLC.`: neither file had a copyright
holder before, so this is the same line as `antora-playbook.yml` (#29),
and the `openig/antora.yml` header is identical to the one in #31.
`header-content.hbs` gets none: it overrides an antora-ui-default
partial (MPL-2.0).

No version number is written as text: the badge always shows the current
release, and the version menus from #31 show the current line (16.x,
5.x, 7.x, 6.x) and the older versions.

## Verification
Local Antora build of this branch:
- all external links on the start page and in the header return 200; the
badge is served as `image/svg+xml`;
- the Downloads menu is present on every page (checked on the start page
and `openam/admin-guide`);
- no new Antora errors;
- rebased onto `master` after #29 and #30; the start page keeps the *API
Reference (Javadoc)* items from #30 next to the new links. No conflicts
with the other open pull requests (`git merge-tree`):
`openig/antora.yml` with #31, `header-content.hbs` with #33.
vharseko added a commit that referenced this pull request Oct 2, 2026
…r version (#31)

Fixes #26

Publishes the last release of each previous major line as an older
version; `master` stays the default version, so no URL changes.

| Product | `master` (unchanged URLs) | Older version |
|---|---|---|
| OpenAM | `/openam/…`, shown as **16.x** | `/openam/15.2/…`, **15.2.2**
(tag `OpenAM-15.2.2`) |
| OpenDJ | `/opendj/…`, **5.x** | `/opendj/4.10/…`, **4.10.2** |
| OpenIDM | `/openidm/…`, **7.x** | `/openidm/6.3/…`, **6.3.0** |
| OpenIG | `/openig/…`, **6.x** | `/openig/5.3/…`, **5.3.2** |

The version menus of the default UI (component list and the per-page
version menu) show the older versions as soon as a component has more
than one.

## Why an extension (`lib/versions-from-tags.js`)
Listing the tags as content sources in the playbook is not enough:
1. the `antora.yml` of every tag declares `version: ~`, and Antora gives
the version in `antora.yml` precedence over the version of the content
source, so a tag would be merged into the unversioned component version
built from `master`;
2. Antora reads **every** file under the start path of a git tree, and
at a tag that includes the API docs of that time (18,672 files for
`OpenAM-15.2.2`, against 330 pages): the build took 4m14s instead of
43s.

The extension replaces `aggregateContent`: the branch sources are
aggregated as before; for each tag source it extracts only
`<product>/antora.yml` and `<product>/modules` with `git archive` into a
temporary repository under `build/antora-tags/<tag>` (a fixed path, so
that the errors Antora logs for tag content can be listed in the
baseline of #28), aggregates it separately and sets the version from the
tag name (`OpenAM-15.2.2` → version `15.2`, display `15.2.2`). The build
now takes about 1m16s locally. If the tag is missing — as in the shallow
`actions/checkout` of CI — the extension fetches it from `origin` (with
`--depth=1` in a shallow repository), so no workflow change is needed.
If the fetch fails (e.g. a fork created before the tags were pushed, or
no network), the build fails with the `git fetch` command that gets the
tag from the upstream, followed by git's own error. The temporary commit
runs with `core.hooksPath=/dev/null`, so global git hooks of a developer
cannot abort the build.

## Other changes
- `<product>/antora.yml`: `display_version` of `master` (16.x, 5.x, 7.x,
6.x).
- `lib/add-pdf-link.js`: the "Download PDF" link only on the latest
version — the wiki PDFs are built from the current sources.
- `lib/add-pdf-link.js`, `<product>/antora.yml`: CDDL header with
`Copyright 2024-2026 3A Systems, LLC.` — these files had no copyright
line; the playbook got the same line in #29.
- Tag sources have `edit_url: false`: no "Edit this Page" on older
versions.
- `docsearch/config.json`: older versions
(`/<product>/<major>.<minor>/`) added to `stop_urls`, so search does not
mix them with the current docs.

## Verification
Local builds:
- the 285 `master` pages differ from the current build only by the older
versions added to the version menus; the start page still links to
`master`;
- the canonical URL of an older page points to the same page in `master`
(e.g. `…/openam/admin-guide/`); no PDF or edit link on older pages;
- CI scenario: a `--depth 1` clone without tags — the extension fetched
the 4 tags and the build succeeded;
- rebased onto `master` with #27 and #29 merged: the only conflict, the
copyright line of `antora-playbook.yml`, resolved to `master`; merges
with #28 without conflicts;
- a `--depth 1 --no-tags` clone whose `origin` is unreachable: the build
fails with `cannot fetch tag OpenDJ-4.10.2 from the origin of …; if the
tag is missing there, run git fetch …`, then git's `Couldn't connect to
server`;
- the OpenIG-5.3.2 copy of the known `sec-release-levels.adoc` error is
logged as
`build/antora-tags/OpenIG-5.3.2/openig/modules/ROOT/partials/sec-release-levels.adoc`
with refname `OpenIG-5.3.2`, the same on every run;
- site size with API docs: **877 MB** (GitHub Pages limit: 1 GB);
- `lychee --offline`: 106 broken links instead of 45. The new ones are
the same legacy ForgeRock-era links inside the frozen tags, plus 8 links
from `openam/15.2` pages to `openam/15.2/apidocs` — the API docs are
published for the current version only.

## Merge order with #28
#28 adds a gate that fails a pull request on an Antora error or a broken
link missing from `.github/build-baseline`. This PR adds one Antora
error (the OpenIG-5.3.2 copy above) and the new broken links of the
older versions. Whichever of the two lands second runs
`.github/build-baseline/check.sh --update` on its build and commits both
lists.

## Not in this PR
- When a new major line is released (e.g. OpenAM 17), the tag in the
playbook and `display_version` are updated by hand.
- The stray tag `${TAG_NAME}` (`0abcd92e4f`) mentioned in #26 is not
deleted yet; deleting a tag on `origin` is done separately.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SEO: add Sitemap to robots.txt, meta description and OpenGraph tags

2 participants