[#25] Move the site search to DocSearch 3 and prepare the Algolia Crawler - #33
Conversation
maximthomas
left a comment
There was a problem hiding this comment.
praise: The front end moves to DocSearch 3 on the live index without touching the scraper, and the crawler config is prepared so it cannot affect production.
transformItemsinsupplemental-ui/partials/footer-scripts.hbs:16strips the legacy#topbar-navanchor client-side, so no reindex is needed.docsearch/crawler-config.jskeeps the write key out of the repository (apiKey: '<crawler API key>',:26) and writes a separate index,doc_openidentityplatform_v3(:46), so the live search is untouched while it is checked.- The
SITE_SEARCH_PROVIDERbranch stays intact (supplemental-ui/partials/header-content.hbs:23).
issue (non-blocking): The readme says the search renders only when both ALGOLIA_SEARCH_API_KEY and ALGOLIA_APP_ID are set, but the templates check only the search key.
docsearch/readme.md:23, supplemental-ui/partials/header-content.hbs:19, supplemental-ui/partials/footer-scripts.hbs:9-11
header-content.hbs, head-styles.hbs and footer-scripts.hbs check only env.ALGOLIA_SEARCH_API_KEY, and appId is emitted only inside {{#with env.ALGOLIA_APP_ID}}. A local or fork build that sets only the search key renders the DocSearch button. On click or Ctrl/Cmd+K, the v3 client throws Error("`appId` is missing.") (it is in the vendored bundle), so the modal never opens. docsearch.js 2 fell back to the shared DocSearch app, so a build like that already searched the wrong app before this PR; what is new is a sentence that says such a build hides the search. Production is unaffected because publish.yml sets both variables.
It is rendered when the build has `ALGOLIA_SEARCH_API_KEY` in the environment; `ALGOLIA_APP_ID` must be set too,
or the search modal fails to open.Or: check both variables in the three partials, e.g. {{#if (and env.ALGOLIA_SEARCH_API_KEY env.ALGOLIA_APP_ID)}}.
question (non-blocking): Is it intended that crawler-config.js indexes the site home page, which config.json never crawled?
docsearch/crawler-config.js:29, :19-20
startUrls is the site root, discoveryPatterns/pathsToMatch are https://doc.openidentityplatform.org/**, and no exclusion matches /. The config.json start URLs are only /openam, /opendj, /openidm and /openig, and the live index has no record for the ROOT page. So doc_openidentityplatform_v3 will contain a page the live index lacks, which matters for the comparison in readme step 4, and the "same pages" in the header and in the PR description is not accurate. Minor either way: if this is intended, the header should say so; if not, exclude the page.
exclusionPatterns: [
// the site home page (ROOT component), not crawled by config.json
'https://doc.openidentityplatform.org/',Or: change the header to "same pages plus the site home page".
nitpick (non-blocking): The new readme header says "Portions Copyright", but the file has no earlier copyright holder.
docsearch/readme.md:14
At base the file was a three-line note with no header. With no earlier holder, the line is Copyright 2026 3A Systems, LLC., the form docsearch/crawler-config.js:14 uses in the same commit. "Portions" is for files that already carry another holder's line, such as supplemental-ui/partials/head-meta.hbs.
Copyright 2026 3A Systems, LLC.nitpick (non-blocking): The comment refers to README.md, but the file is docsearch/readme.md.
docsearch/crawler-config.js:19
The repository has no README.md, so the name does not resolve on GitHub or on case-sensitive file systems.
// Crawler dashboard, which runs it; see readme.md. It mirrors config.json: same pages, samenote (non-blocking): Differences from what the description says, with no stated reason:
docsearch/crawler-config.js: the header and the description say "same selectors", but lvl0 drops.nav-panel-explore .context .version, and thecomponentextra attribute,min_indexed_leveland thedesc(version)ranking are gone. None of this changes the result today: everyantora.ymlhasversion: ~, so the version span is empty; nothing readscomponent; live records have noversionattribute.
948ccc6 to
8518b3a
Compare
|
Rebased onto
Verified again: built with both variables, |
maximthomas
left a comment
There was a problem hiding this comment.
praise: The review fixes landed where the problems were, and the crawler config does what it says.
- One gate,
(and env.ALGOLIA_SEARCH_API_KEY env.ALGOLIA_APP_ID), in all three partials (footer-scripts.hbs:3,head-styles.hbs:3,header-content.hbs:19). crawler-config.jswritesdoc_openidentityplatform_v3, so the live search is untouched until the switch. Under micromatch 4.0.8 itspathsToMatchtakes/openamand/openam/but not/, and its exclusions catchapidocs,/openam/15.2/…and the three generated references.- It merges cleanly with #30 and #32 (
git merge-tree, both heads).
issue (non-blocking): transformItems strips #topbar-nav, but page-title records in the live index end in #cookie-notice.
supplemental-ui/partials/footer-scripts.hbs:13-18, supplemental-ui/partials/head-scripts.hbs:29
The legacy scraper anchors each page title to the first id on the page, and that id is <div id="cookie-notice">, which comes before the header. In doc_openidentityplatform, all 286 lvl1 records among the first 1000 hits of an empty query end in #cookie-notice, and none ends in #topbar-nav. The same holds for all 33 lvl1 hits for "session", "replication", "admin guide" and "opendj". So the rewrite never fires: page-title hits open …/chap-session-state#cookie-notice, and that URL is also stored in recent searches. The page still lands at the top, because site.js scrolls to 0 for an id outside the article. The check "no result URL ends with #topbar-nav" passes only because no URL ever has that anchor.
// the legacy scraper gives page titles the first id on the page (#cookie-notice) as anchor
transformItems: function (items) {
return items.map(function (item) {
return Object.assign({}, item, { url: item.url.replace(/#(?:cookie-notice|topbar-nav)$/, '') })
})
}question (non-blocking): When a configuration with apiKey: '<crawler API key>' is pasted over the generated one, does the Crawler dashboard keep the apiKey it generated?
docsearch/readme.md:36, docsearch/crawler-config.js:28-29
The dashboard editor stores the configuration as code, with appId and apiKey as literal values. If it does not re-inject the key on save, step 2 ("paste crawler-config.js; the dashboard sets apiKey") replaces the write key with the placeholder, and the first run fails to authenticate. If it does re-inject the key, nothing needs to change. This was not run, because it needs Crawler access on X0ME9NKL6F. Wording that is safe either way:
2. Create a crawler in the Crawler dashboard and paste `crawler-config.js` over the generated configuration,
keeping the `apiKey` line the dashboard generated.suggestion (non-blocking): Nothing checks that the three copies of the search gate agree.
supplemental-ui/partials/footer-scripts.hbs:3, supplemental-ui/partials/head-styles.hbs:3, supplemental-ui/partials/header-content.hbs:19
The repo has no tests and no PR CI. Suppose one copy is edited on its own, for example footer-scripts.hbs:3 back to {{#if env.ALGOLIA_SEARCH_API_KEY}}. The site still builds, and a build that has only the key loads docsearch.min.js and calls docsearch({ container: '#docsearch' }) on a page with no #docsearch element. Two builds in the workflow that #28 adds would catch this:
ALGOLIA_APP_ID=X0ME9NKL6F ALGOLIA_SEARCH_API_KEY=test npx antora --clean antora-playbook.yml
grep -q '<div id="docsearch"></div>' build/site/openam/index.html
grep -q 'js/vendor/docsearch.min.js' build/site/openam/index.html
ALGOLIA_SEARCH_API_KEY=test npx antora --clean antora-playbook.yml
! grep -q 'docsearch' build/site/openam/index.htmlPin: reverting the gate in any one of the three partials to the key alone turns the last line red.
nitpick (non-blocking): The new header in docsearch/readme.md says "Portions Copyright", but the file has no earlier copyright holder.
docsearch/readme.md:14
The earlier two-line readme had no copyright notice, so no earlier holder shares the credit. crawler-config.js:14 in the same PR uses the plain form.
Copyright 2026 3A Systems, LLC.nitpick (non-blocking): The PR description lists url_without_anchor among the fields the v3 front end reads.
supplemental-ui/js/vendor/docsearch.min.js:1
The vendored 3.9.0 bundle retrieves hierarchy.lvl0–lvl6, content, type and url. url_without_anchor does not occur in it. "No reindexing is needed" still holds. Drop the field from that list if the description becomes the squash commit message.
- @docsearch/js 3.9.0 (modal, Ctrl/Cmd+K) instead of docsearch.js 2, on the same index; results no longer point to #topbar-nav - docsearch/crawler-config.js: the same crawl for the Algolia Crawler, into a new index; docsearch/readme.md: how to switch to it Refs OpenIdentityPlatform#25
…oducts only - the search is rendered only when both ALGOLIA_SEARCH_API_KEY and ALGOLIA_APP_ID are set; @docsearch/js 3 throws without appId - crawler-config.js: records only from the product pages, as the start_urls of config.json; the header lists what is not carried over - readme.md: fix the file name, check pages with the URL tester Refs OpenIdentityPlatform#25
…awler key - transformItems strips #cookie-notice, the anchor the legacy scraper gives to page titles and the text before the first section; no record has #topbar-nav - readme.md: paste the crawler configuration keeping the apiKey line the dashboard generated; plain 3A Systems copyright, as the file had no earlier holder Refs OpenIdentityPlatform#25
8518b3a to
f47c365
Compare
|
Rebased onto
|
maximthomas
left a comment
There was a problem hiding this comment.
praise: This push fixes each open round-2 finding where it was raised.
supplemental-ui/partials/footer-scripts.hbs:17strips#cookie-notice. All 286 page-title records of the livedoc_openidentityplatformindex carry that anchor.docsearch/readme.md:36-37now says to keep theapiKeyline the Crawler dashboard generated, and the comment atdocsearch/crawler-config.js:28matches.docsearch/readme.md:14readsCopyright 2025-2026 3A Systems, LLC.. There is no earlier holder, and the start year is the file's first commit (559c113).
…leaseversion (#32) Fixes #23 ## Changes - **`ROOT/modules/ROOT/pages/index.adoc`** (start page): under the description of each product, a line with - a shields.io badge of the latest release (e.g. *latest release v16.1.3* for OpenAM), linking to `releases/latest`; - *Releases and downloads* → `github.com/OpenIdentityPlatform/<Product>/releases`; - *Docker images* → `hub.docker.com/r/openidentityplatform/<product>`; - *Security advisories* → `…/security/advisories` for OpenAM, OpenDJ and OpenIDM (OpenIG has no published advisories, so no link). - **`supplemental-ui/partials/header-content.hbs`**: a **Downloads** menu next to Projects, with the release pages of the four products, the Docker Hub organization and the security advisories of OpenDJ, OpenAM and OpenIDM. - **`openig/antora.yml`**: `releaseversion: 5.2.4` removed — no page or UI template uses it (checked in the sources and in the generated site). - **CDDL headers** added to `index.adoc` and `openig/antora.yml` with `Copyright 2024-2026 3A Systems, LLC.`: neither file had a copyright holder before, so this is the same line as `antora-playbook.yml` (#29), and the `openig/antora.yml` header is identical to the one in #31. `header-content.hbs` gets none: it overrides an antora-ui-default partial (MPL-2.0). No version number is written as text: the badge always shows the current release, and the version menus from #31 show the current line (16.x, 5.x, 7.x, 6.x) and the older versions. ## Verification Local Antora build of this branch: - all external links on the start page and in the header return 200; the badge is served as `image/svg+xml`; - the Downloads menu is present on every page (checked on the start page and `openam/admin-guide`); - no new Antora errors; - rebased onto `master` after #29 and #30; the start page keeps the *API Reference (Javadoc)* items from #30 next to the new links. No conflicts with the other open pull requests (`git merge-tree`): `openig/antora.yml` with #31, `header-content.hbs` with #33.
Fixes #25, together with #28 — this PR does item 5, the migration to the current DocSearch; items 1–4 are #28. The front end moves now, on the existing index; the crawler stays the legacy scraper until Algolia Crawler access is set up, for which this PR prepares the configuration.
Changes
Front end:
@docsearch/js3.9.0 instead ofdocsearch.js2 (vendored as before,supplemental-ui/js/vendor/docsearch.min.jsandsupplemental-ui/css/vendor/docsearch.min.css):Ctrl/Cmd+Kand/;doc_openidentityplatform): the records of the legacy scraper carry the fields the v3 front end reads (hierarchy.lvl0–6,content,type,url), so no reindexing is needed;transformItemsdrops the#cookie-noticeanchor that the legacy scraper gives to page titles and to the text before the first section (the first id on the page, the cookie notice);ALGOLIA_SEARCH_API_KEYandALGOLIA_APP_ID(the v3 client requiresappId); otherwise theSITE_SEARCH_PROVIDERbranch (plain search input), which is kept, or no search.3.9.0 is the last 3.x; 4.x and 5.x add Ask AI and a side panel, which this site does not use, and are 4–5 times larger (510 KB and 619 KB of JS against 133 KB).
docsearch/crawler-config.js(new): the same crawl for the Algolia Crawler — records only from the product pages, as thestart_urlsofconfig.json(the site home page is crawled for links only), the same exclusions (API docs, generated references, and the older versions of [#26] Publish the last release of each previous major line as an older version #31) and the same selectors for the hierarchy and the content. Not carried over, because none of it changes the results: the version in lvl0 and thedesc(version)ranking (every component hasversion: ~), thecomponentattribute (nothing reads it) andmin_indexed_level. It writes a new index,doc_openidentityplatform_v3, so the live search is not affected while it is checked. Not run yet: it needs Crawler access on the Algolia application (for an open-source site, through the DocSearch program).docsearch/readme.md: how the search is set up, and the steps to move to the Crawler. The link it held before (how the index was set up for the legacy scraper) is kept.publish.ymlandconfig.jsonare unchanged: the legacy scraper keeps reindexing after every deployment.Verification
Local build with the public search key of the site, checked in a browser (Playwright):
Ctrl+Kopens the modal; "session" and "replication" return results from the live index, grouped by product; page titles link to the page with no anchor, and no result URL ends with#cookie-notice(in the live index every page-title record does);ALGOLIA_SEARCH_API_KEYhas neither the button nor the DocSearch bundle;pathsToMatchpattern ofcrawler-config.js, checked with micromatch, matches/openamand/openam/but not/,/openicf/…or/commons/…;master(with [#24] SEO: declare the sitemap in robots.txt, add description and OpenGraph tags #29); merges with [#22] Link API docs (Javadoc) from the start page and the Projects menu #30 and [#23] Add release, download, Docker and security links; drop stale releaseversion #32 without conflicts (git merge-tree).