Skip to content

[#25] Move the site search to DocSearch 3 and prepare the Algolia Crawler - #33

Merged
vharseko merged 3 commits into
OpenIdentityPlatform:masterfrom
vharseko:issue-25-docsearch-v3
Oct 2, 2026
Merged

vharseko merged 3 commits into
OpenIdentityPlatform:masterfrom
vharseko:issue-25-docsearch-v3

Conversation

@vharseko

@vharseko vharseko commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Fixes #25, together with #28 — this PR does item 5, the migration to the current DocSearch; items 1–4 are #28. The front end moves now, on the existing index; the crawler stays the legacy scraper until Algolia Crawler access is set up, for which this PR prepares the configuration.

Changes

  • Front end: @docsearch/js 3.9.0 instead of docsearch.js 2 (vendored as before, supplemental-ui/js/vendor/docsearch.min.js and supplemental-ui/css/vendor/docsearch.min.css):

    • the header shows the DocSearch button, which opens the search modal; shortcuts Ctrl/Cmd+K and /;
    • same application, key and index (doc_openidentityplatform): the records of the legacy scraper carry the fields the v3 front end reads (hierarchy.lvl0–6, content, type, url), so no reindexing is needed;
    • transformItems drops the #cookie-notice anchor that the legacy scraper gives to page titles and to the text before the first section (the first id on the page, the cookie notice);
    • rendered only when the build has both ALGOLIA_SEARCH_API_KEY and ALGOLIA_APP_ID (the v3 client requires appId); otherwise the SITE_SEARCH_PROVIDER branch (plain search input), which is kept, or no search.

    3.9.0 is the last 3.x; 4.x and 5.x add Ask AI and a side panel, which this site does not use, and are 4–5 times larger (510 KB and 619 KB of JS against 133 KB).

  • docsearch/crawler-config.js (new): the same crawl for the Algolia Crawler — records only from the product pages, as the start_urls of config.json (the site home page is crawled for links only), the same exclusions (API docs, generated references, and the older versions of [#26] Publish the last release of each previous major line as an older version #31) and the same selectors for the hierarchy and the content. Not carried over, because none of it changes the results: the version in lvl0 and the desc(version) ranking (every component has version: ~), the component attribute (nothing reads it) and min_indexed_level. It writes a new index, doc_openidentityplatform_v3, so the live search is not affected while it is checked. Not run yet: it needs Crawler access on the Algolia application (for an open-source site, through the DocSearch program).

  • docsearch/readme.md: how the search is set up, and the steps to move to the Crawler. The link it held before (how the index was set up for the legacy scraper) is kept.

publish.yml and config.json are unchanged: the legacy scraper keeps reindexing after every deployment.

Verification

Local build with the public search key of the site, checked in a browser (Playwright):

@vharseko vharseko added documentation Improvements or additions to documentation enhancement New feature or request labels Sep 30, 2026
@vharseko vharseko added the ci Continuous integration, build and publish workflows label Sep 30, 2026

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The front end moves to DocSearch 3 on the live index without touching the scraper, and the crawler config is prepared so it cannot affect production.

  • transformItems in supplemental-ui/partials/footer-scripts.hbs:16 strips the legacy #topbar-nav anchor client-side, so no reindex is needed.
  • docsearch/crawler-config.js keeps the write key out of the repository (apiKey: '<crawler API key>', :26) and writes a separate index, doc_openidentityplatform_v3 (:46), so the live search is untouched while it is checked.
  • The SITE_SEARCH_PROVIDER branch stays intact (supplemental-ui/partials/header-content.hbs:23).

issue (non-blocking): The readme says the search renders only when both ALGOLIA_SEARCH_API_KEY and ALGOLIA_APP_ID are set, but the templates check only the search key.

docsearch/readme.md:23, supplemental-ui/partials/header-content.hbs:19, supplemental-ui/partials/footer-scripts.hbs:9-11

header-content.hbs, head-styles.hbs and footer-scripts.hbs check only env.ALGOLIA_SEARCH_API_KEY, and appId is emitted only inside {{#with env.ALGOLIA_APP_ID}}. A local or fork build that sets only the search key renders the DocSearch button. On click or Ctrl/Cmd+K, the v3 client throws Error("`appId` is missing.") (it is in the vendored bundle), so the modal never opens. docsearch.js 2 fell back to the shared DocSearch app, so a build like that already searched the wrong app before this PR; what is new is a sentence that says such a build hides the search. Production is unaffected because publish.yml sets both variables.

  It is rendered when the build has `ALGOLIA_SEARCH_API_KEY` in the environment; `ALGOLIA_APP_ID` must be set too,
  or the search modal fails to open.

Or: check both variables in the three partials, e.g. {{#if (and env.ALGOLIA_SEARCH_API_KEY env.ALGOLIA_APP_ID)}}.


question (non-blocking): Is it intended that crawler-config.js indexes the site home page, which config.json never crawled?

docsearch/crawler-config.js:29, :19-20

startUrls is the site root, discoveryPatterns/pathsToMatch are https://doc.openidentityplatform.org/**, and no exclusion matches /. The config.json start URLs are only /openam, /opendj, /openidm and /openig, and the live index has no record for the ROOT page. So doc_openidentityplatform_v3 will contain a page the live index lacks, which matters for the comparison in readme step 4, and the "same pages" in the header and in the PR description is not accurate. Minor either way: if this is intended, the header should say so; if not, exclude the page.

  exclusionPatterns: [
    // the site home page (ROOT component), not crawled by config.json
    'https://doc.openidentityplatform.org/',

Or: change the header to "same pages plus the site home page".


nitpick (non-blocking): The new readme header says "Portions Copyright", but the file has no earlier copyright holder.

docsearch/readme.md:14

At base the file was a three-line note with no header. With no earlier holder, the line is Copyright 2026 3A Systems, LLC., the form docsearch/crawler-config.js:14 uses in the same commit. "Portions" is for files that already carry another holder's line, such as supplemental-ui/partials/head-meta.hbs.

Copyright 2026 3A Systems, LLC.

nitpick (non-blocking): The comment refers to README.md, but the file is docsearch/readme.md.

docsearch/crawler-config.js:19

The repository has no README.md, so the name does not resolve on GitHub or on case-sensitive file systems.

// Crawler dashboard, which runs it; see readme.md. It mirrors config.json: same pages, same

note (non-blocking): Differences from what the description says, with no stated reason:

  • docsearch/crawler-config.js: the header and the description say "same selectors", but lvl0 drops .nav-panel-explore .context .version, and the component extra attribute, min_indexed_level and the desc(version) ranking are gone. None of this changes the result today: every antora.yml has version: ~, so the version span is empty; nothing reads component; live records have no version attribute.

@vharseko
vharseko force-pushed the issue-25-docsearch-v3 branch from 948ccc6 to 8518b3a Compare October 1, 2026 12:35
@vharseko

vharseko commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

Rebased onto master (no conflicts) and addressed in 8518b3a:

  • ALGOLIA_APP_ID: fixed in the templates rather than in the readme. header-content.hbs, head-styles.hbs and footer-scripts.hbs now check (and env.ALGOLIA_SEARCH_API_KEY env.ALGOLIA_APP_ID), and appId is emitted directly: the {{#with}} was left from docsearch.js 2, where appId was optional. A build with only the search key now falls through to the SITE_SEARCH_PROVIDER branch (or no search) instead of a button that cannot open; the readme says "both".
  • Home page: not intended. pathsToMatch is now https://doc.openidentityplatform.org/{openam,opendj,openidm,openig}/**, as the start_urls of config.json; the root stays in startUrls/discoveryPatterns, so it is crawled for links but gives no records. Checked with micromatch: the pattern matches /openam and /openam/, not /, /openicf/… or /commons/…. Readme step 2 adds a check with the dashboard's URL tester, since the crawler itself has not run yet.
  • README.md → readme.md in the crawler-config.js header.
  • Selectors: the header now says what is not carried over and why (version in lvl0 and desc(version), component, min_indexed_level), instead of "same selectors"; the PR description is updated to match.
  • "Portions Copyright" in docsearch/readme.md: kept. The file existed before this PR (559c113e41), and the convention for this repository is a Portions Copyright line in every modified existing file; a standalone Copyright 2026 would also misdate the 2025 text.

Verified again: built with both variables, Ctrl+K opens the modal and "session" returns results from OpenAM, OpenIDM and OpenIG with no #topbar-nav URL; built with only ALGOLIA_SEARCH_API_KEY, the page has neither the button nor the DocSearch bundle.

@vharseko
vharseko requested a review from maximthomas October 1, 2026 12:36

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The review fixes landed where the problems were, and the crawler config does what it says.

  • One gate, (and env.ALGOLIA_SEARCH_API_KEY env.ALGOLIA_APP_ID), in all three partials (footer-scripts.hbs:3, head-styles.hbs:3, header-content.hbs:19).
  • crawler-config.js writes doc_openidentityplatform_v3, so the live search is untouched until the switch. Under micromatch 4.0.8 its pathsToMatch takes /openam and /openam/ but not /, and its exclusions catch apidocs, /openam/15.2/… and the three generated references.
  • It merges cleanly with #30 and #32 (git merge-tree, both heads).

issue (non-blocking): transformItems strips #topbar-nav, but page-title records in the live index end in #cookie-notice.

supplemental-ui/partials/footer-scripts.hbs:13-18, supplemental-ui/partials/head-scripts.hbs:29

The legacy scraper anchors each page title to the first id on the page, and that id is <div id="cookie-notice">, which comes before the header. In doc_openidentityplatform, all 286 lvl1 records among the first 1000 hits of an empty query end in #cookie-notice, and none ends in #topbar-nav. The same holds for all 33 lvl1 hits for "session", "replication", "admin guide" and "opendj". So the rewrite never fires: page-title hits open …/chap-session-state#cookie-notice, and that URL is also stored in recent searches. The page still lands at the top, because site.js scrolls to 0 for an id outside the article. The check "no result URL ends with #topbar-nav" passes only because no URL ever has that anchor.

  // the legacy scraper gives page titles the first id on the page (#cookie-notice) as anchor
  transformItems: function (items) {
    return items.map(function (item) {
      return Object.assign({}, item, { url: item.url.replace(/#(?:cookie-notice|topbar-nav)$/, '') })
    })
  }

question (non-blocking): When a configuration with apiKey: '<crawler API key>' is pasted over the generated one, does the Crawler dashboard keep the apiKey it generated?

docsearch/readme.md:36, docsearch/crawler-config.js:28-29

The dashboard editor stores the configuration as code, with appId and apiKey as literal values. If it does not re-inject the key on save, step 2 ("paste crawler-config.js; the dashboard sets apiKey") replaces the write key with the placeholder, and the first run fails to authenticate. If it does re-inject the key, nothing needs to change. This was not run, because it needs Crawler access on X0ME9NKL6F. Wording that is safe either way:

2. Create a crawler in the Crawler dashboard and paste `crawler-config.js` over the generated configuration,
   keeping the `apiKey` line the dashboard generated.

suggestion (non-blocking): Nothing checks that the three copies of the search gate agree.

supplemental-ui/partials/footer-scripts.hbs:3, supplemental-ui/partials/head-styles.hbs:3, supplemental-ui/partials/header-content.hbs:19

The repo has no tests and no PR CI. Suppose one copy is edited on its own, for example footer-scripts.hbs:3 back to {{#if env.ALGOLIA_SEARCH_API_KEY}}. The site still builds, and a build that has only the key loads docsearch.min.js and calls docsearch({ container: '#docsearch' }) on a page with no #docsearch element. Two builds in the workflow that #28 adds would catch this:

ALGOLIA_APP_ID=X0ME9NKL6F ALGOLIA_SEARCH_API_KEY=test npx antora --clean antora-playbook.yml
grep -q '<div id="docsearch"></div>' build/site/openam/index.html
grep -q 'js/vendor/docsearch.min.js' build/site/openam/index.html
ALGOLIA_SEARCH_API_KEY=test npx antora --clean antora-playbook.yml
! grep -q 'docsearch' build/site/openam/index.html

Pin: reverting the gate in any one of the three partials to the key alone turns the last line red.


nitpick (non-blocking): The new header in docsearch/readme.md says "Portions Copyright", but the file has no earlier copyright holder.

docsearch/readme.md:14

The earlier two-line readme had no copyright notice, so no earlier holder shares the credit. crawler-config.js:14 in the same PR uses the plain form.

Copyright 2026 3A Systems, LLC.

nitpick (non-blocking): The PR description lists url_without_anchor among the fields the v3 front end reads.

supplemental-ui/js/vendor/docsearch.min.js:1

The vendored 3.9.0 bundle retrieves hierarchy.lvl0–lvl6, content, type and url. url_without_anchor does not occur in it. "No reindexing is needed" still holds. Drop the field from that list if the description becomes the squash commit message.

- @docsearch/js 3.9.0 (modal, Ctrl/Cmd+K) instead of docsearch.js 2,
  on the same index; results no longer point to #topbar-nav
- docsearch/crawler-config.js: the same crawl for the Algolia Crawler,
  into a new index; docsearch/readme.md: how to switch to it

Refs OpenIdentityPlatform#25
…oducts only

- the search is rendered only when both ALGOLIA_SEARCH_API_KEY and
  ALGOLIA_APP_ID are set; @docsearch/js 3 throws without appId
- crawler-config.js: records only from the product pages, as the
  start_urls of config.json; the header lists what is not carried over
- readme.md: fix the file name, check pages with the URL tester

Refs OpenIdentityPlatform#25
…awler key

- transformItems strips #cookie-notice, the anchor the legacy scraper
  gives to page titles and the text before the first section; no
  record has #topbar-nav
- readme.md: paste the crawler configuration keeping the apiKey line
  the dashboard generated; plain 3A Systems copyright, as the file had
  no earlier holder

Refs OpenIdentityPlatform#25
@vharseko
vharseko force-pushed the issue-25-docsearch-v3 branch from 8518b3a to f47c365 Compare October 1, 2026 14:26
@vharseko

vharseko commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

Rebased onto master (with #29, no conflicts) and addressed in f47c365:

  • #cookie-notice: confirmed and fixed. In the live index every page-title record, and the text before the first section, ends in #cookie-notice (for "session": 56 hits, 3 of them lvl1); none ends in #topbar-nav. transformItems now strips #cookie-notice from every item, and the comment says where the anchor comes from. My earlier check was vacuous, as you say: it looked for an anchor the index never has. Checked again in a browser: "session" and "replication" return the page titles …/chap-session-state, …/chap-session-failover and …/chap-replication with no anchor, and no result URL has #cookie-notice.
  • Crawler apiKey: readme step 2 now uses your wording (paste over the generated configuration, keeping the generated apiKey line), and the comment in crawler-config.js matches it.
  • url_without_anchor: dropped from the PR description; the 3.9.0 bundle retrieves hierarchy.lvl0–6, content, type and url only.
  • Check that the three gates agree: not in this PR. It has no CI to hook into; the PR check comes with [#25] CI: build pull requests and make the site build reproducible #28. Two more full Antora builds on every PR is a high price for three identical lines; if it is wanted, a grep for the gate in the three partials in [#25] CI: build pull requests and make the site build reproducible #28's workflow does the same without a build. I can add that as a follow-up once both are merged.
  • "Portions Copyright" in docsearch/readme.md: fixed to Copyright 2025-2026 3A Systems, LLC. The file was created in 2025 with no notice, so there is no earlier holder to share the credit with, and my answer in the previous round was wrong.

@vharseko
vharseko requested a review from maximthomas October 1, 2026 14:26

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: This push fixes each open round-2 finding where it was raised.

  • supplemental-ui/partials/footer-scripts.hbs:17 strips #cookie-notice. All 286 page-title records of the live doc_openidentityplatform index carry that anchor.
  • docsearch/readme.md:36-37 now says to keep the apiKey line the Crawler dashboard generated, and the comment at docsearch/crawler-config.js:28 matches.
  • docsearch/readme.md:14 reads Copyright 2025-2026 3A Systems, LLC.. There is no earlier holder, and the start year is the file's first commit (559c113).

@vharseko
vharseko merged commit 1c25aa7 into OpenIdentityPlatform:master Oct 2, 2026
@vharseko
vharseko deleted the issue-25-docsearch-v3 branch October 2, 2026 07:24
vharseko added a commit that referenced this pull request Oct 2, 2026
…leaseversion (#32)

Fixes #23

## Changes
- **`ROOT/modules/ROOT/pages/index.adoc`** (start page): under the
description of each product, a line with
- a shields.io badge of the latest release (e.g. *latest release
v16.1.3* for OpenAM), linking to `releases/latest`;
- *Releases and downloads* →
`github.com/OpenIdentityPlatform/<Product>/releases`;
  - *Docker images* → `hub.docker.com/r/openidentityplatform/<product>`;
- *Security advisories* → `…/security/advisories` for OpenAM, OpenDJ and
OpenIDM (OpenIG has no published advisories, so no link).
- **`supplemental-ui/partials/header-content.hbs`**: a **Downloads**
menu next to Projects, with the release pages of the four products, the
Docker Hub organization and the security advisories of OpenDJ, OpenAM
and OpenIDM.
- **`openig/antora.yml`**: `releaseversion: 5.2.4` removed — no page or
UI template uses it (checked in the sources and in the generated site).
- **CDDL headers** added to `index.adoc` and `openig/antora.yml` with
`Copyright 2024-2026 3A Systems, LLC.`: neither file had a copyright
holder before, so this is the same line as `antora-playbook.yml` (#29),
and the `openig/antora.yml` header is identical to the one in #31.
`header-content.hbs` gets none: it overrides an antora-ui-default
partial (MPL-2.0).

No version number is written as text: the badge always shows the current
release, and the version menus from #31 show the current line (16.x,
5.x, 7.x, 6.x) and the older versions.

## Verification
Local Antora build of this branch:
- all external links on the start page and in the header return 200; the
badge is served as `image/svg+xml`;
- the Downloads menu is present on every page (checked on the start page
and `openam/admin-guide`);
- no new Antora errors;
- rebased onto `master` after #29 and #30; the start page keeps the *API
Reference (Javadoc)* items from #30 next to the new links. No conflicts
with the other open pull requests (`git merge-tree`):
`openig/antora.yml` with #31, `header-content.hbs` with #33.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Continuous integration, build and publish workflows documentation Improvements or additions to documentation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CI: fail the build on broken xrefs/links and make the build reproducible

2 participants