diff --git a/WIP.Search.md b/WIP.Search.md new file mode 100644 index 00000000..308272a9 --- /dev/null +++ b/WIP.Search.md @@ -0,0 +1,1976 @@ +# twinBASIC Documentation — Site search design notes + +Why searching for a member such as `PaintPicture` does not find it, what was +measured, and the design that fixes it. Companion to +[builder/PLAN-6.md](builder/PLAN-6.md) §5.3, which describes the search-data +generator as ported from Jekyll. + +Like WIP.md, this file is not rendered through tbdocs, so literal dashes are fine here. + +## Resuming this work + +Everything needed to continue is in this file and in `eval/`; nothing +depends on the session that wrote it. + +**Where it stands.** Every numbered item below is done and committed on +branch `claude/paintpicture-docs-runtime-f3250d`, rebased onto +staging's `ea96e4c2`; the user handles the remotes. The working tree +is clean. **The arc is closed for now**: in the user's words it made +search at least semi-functional where it was quite broken before, and +further search work resumes later. No decision is open. The candidates +are under "Next", and the user picks. + +| commit | step | +|---|---| +| `091a1154` | this design doc | +| `6374edb8` | 1: the asterisk crash guard; the replica's tokenizer separator | +| `f91b8848` | 2: `eval/search_quality.mjs` and its baseline | +| `44615b92` | 3: h3 entries; `search.fold_headings` | +| `d267af6c` | 4: `names`/`qualified` fields; the smart dot split | +| `641f1613` | 5: stop words kept; dot runs split; lazy index build | +| `d42ced01` | intent 1: `search_quality.mjs` judges bare names by reader intent | +| `1e840741` | intent 3: exact-name and page-title fields, all words first, query tokens trimmed | +| `7f706e63` | intent 4: `primary` names, non-word characters kept in exact names, kind words | +| `38a4c41b` | pilot 1: prose queries can expect pages right behind (`behind`); ground truth intent-2 | +| `73ff7cb6` | pilot 2: hand-marked index entries, and the first five | +| `10d9c304` | qualified names: `qualified` at 500, reached only by qualified names and word pairs | +| `1a15dae3` | only a page's first heading can be its title (`Shape.Shape`, `Timer.Timer`) | +| `b814e260` | stem twins held whole in `qualified` (`Printer.Fonts`) | +| `18b023fc` | ground truth intent-3: a section of a symbol's page counts for it | +| `9210ce51` | lunr's token-set keys separated: queries with the word `a` threw | +| `2577a8d2` | the `New Functions` / `Form events` probes diagnosed and measured, not shipped | +| `05bd7ca0` | ground truth intent-4: page titles and page-plus-section queries in the eval | +| `2eab7b9d` | a query naming a whole title scores ×3 | +| `c36b5fa1` | a plural kind word doesn't make the other word a name | +| `cb407ad4` | HTML entities decoded per token in the index (`&H80004005`) | +| `85ba0a3a` | lunr's set unions add in place: `a page` 809 → 87 ms | +| `96b57071` | a kind word is required only while an entry found names the thing | +| `67653d2b` | ground truth intent-5: name-and-kind queries in the eval | +| `ec1850c1` | item 6: the approved prose queries, before their entries | +| `c892fa81` | item 6: index entries for the recommended terms | +| `6cd850e3` | item 6: the user's-choice prose queries, before their entries | +| `2007c92e` | item 6: their entries, and a Glossary entry for *namespace* | +| `dc6762f8` | item 6: `standard exe`, `create [an] ActiveX DLL` queries, before their entries | +| `25848eff` | item 6: their entries (New Project first); `pointers` as a secondary entry | +| `100213c9` | item 7: 41 prose queries from two surveys, before their fixes | +| `270f9dae` | item 7: content: ByRef/ByVal, optional and named arguments, returning more than one value | +| `c980c3ea` | item 7: index entries for what the new text still can't find | +| `74a436c6` | after the rebase: `search_quality.mjs` takes `REPO_ROOT` from `lib/repo-paths.mjs` | +| `e2ed3c1e` | `default property` also accepts the Glossary's definition, before it exists | +| `27fc0592` | Glossary *default member*, also called the default property | + +**The eval now** (`node eval/search_quality.mjs`, ground truth +`intent-5`, 10,327 queries): 98.2% at rank 1, 99.1% in the top 10. + +| category | hit@1 | n | what it types | +|---|---|---|---| +| bare names | 99.7% | 2,884 | `PaintPicture`, judged by reader intent (tiers) | +| qualified names | 100% | 5,108 | `Printer.Fonts` | +| name and kind | 91.1% | 1,785 | `MaxHeight property` | +| prose | 97.2% | 108 | hand-picked (`late binding`, `immediate window`, `ByVal`) | +| page titles | 97.2% | 142 | `Return Syntax` | +| page plus section | 96.7% | 300 | `DTPicker Properties` | + +No bare name is out of tier order, and every prose query's `behind` +page is within the top 3 (22 of 22). Hit@10 was 20.5% and MRR .182 when +this work began. The overall hit@1 fell from 99.7% when the name-and-kind +set joined; compare categories, not totals, across ground truths. + +**Known misses (185 at rank 1)**, each recorded where it was diagnosed: +- 158 name-and-kind queries, mostly properties; not diagnosed yet + ([Fixed: kind words](#fixed-kind-words)). +- 10 bare names: 7 enum constants, 2 members + ([second round](#what-shipped-second-round-tiers-in-the-index)) and + `MidB$` ([Same-page sections count](#same-page-sections-count)). +- 3 page titles and 10 page-plus-section queries, each at rank 2 or 3 + ([Fixed: whole titles](#fixed-whole-titles)). +- 1 page title, `Input #`, at 2 behind the `Input` function: the + query trims to `input`. Staging's quoting of the titles that end in + `#` brought it in: the titles had read `Input`, `Line Input` and + `Write`, and the title set skips one-word titles + ([After the rebase](#after-the-rebase-onto-ea96e4c2)). +- 3 prose queries, `declaration`, `comment` and `default property`, + whose Glossary definitions are 2nd, 5th and 5th. That is right by the user's ruling (a + Glossary definition counts within the top 5); an entry would reorder + the bare names `Declare` and `Comments` ([Item 6: shipped](#item-6-shipped)), + and one for `default property` would put it above CommandButton's + `Default` ([Item 7](#item-7-content-and-index-entries-from-two-surveys)). + +**Next**, for the user to choose from: +1. **The 158 name-and-kind misses.** Diagnose them as the whole-title + probes were: group by cause, measure a knob, ship only what makes + nothing worse. 93 aren't in the top 50 at all. +2. **Question-shaped queries.** `how do I register a com dll` misses the + index entry: the all-words pass requires `how`, `do` and `I`, and only + the FAQ holds them all. A fix would touch the all-words pass, so it + needs the spaced and kinds probes as well as the eval. +3. **The cost of one-letter words.** `a p` still takes about 250 ms and a + single letter 80–110 ms, lunr's own work now. Cheaper means changing + the query, and so the ranking, for instance not completing a + one-letter word ([Fixed: slow multi-word queries](#fixed-slow-multi-word-queries)). +4. **More content and index entries**, under the rules below. What item 7 + left: `COM interop` (no clear + target), and three content gaps: subclassing, IntelliSense and + conditional breakpoints (first find out whether the IDE has the last + two). See [Item 7](#item-7-content-and-index-entries-from-two-surveys). +5. **An extra word hides a page.** `memory window`, `history panel`, + `registry access`: the all-words pass finds a few pages holding both + words, and doesn't fall back, so the page named by the other word + never shows. Item 7 fixed 20 such queries with entries, but the cause + is the pass itself, as in item 2; a fix there needs the spaced and + kinds probes. + +6. **Operator-shaped page titles.** `Input #` (above) and `Mid =` + miss rank 1 because the query trims its non-word character. Exact + names already spell such characters as word characters (`#If`); + whether a whole title can too needs the title set and the spaced + probe. + +**Done, in order** (each section has the measurements): +1. The 9 qualified misses: [the title-heading fix](#fixed-a-member-heading-taken-for-the-page-title) + and [the stem twins](#fixed-stem-twins). +2. The same-page ground truth: [Same-page sections count](#same-page-sections-count). +3. Whole titles: [Fixed: whole titles](#fixed-whole-titles). +4. `&H80004005`: [Fixed: entities in the index](#fixed-entities-in-the-index). +5. Slow multi-word queries: [Fixed: slow multi-word queries](#fixed-slow-multi-word-queries). +6. The wider index pass: [Item 6: shipped](#item-6-shipped), after + [Item 6: candidates for approval](#item-6-candidates-for-approval); + on the way, [Fixed: kind words](#fixed-kind-words) and ground truth + `intent-5`. +7. Content and entries from two new surveys: + [Item 7](#item-7-content-and-index-entries-from-two-surveys). + +**The index pilot held up.** The user asked whether ranking tweaks are an +uphill battle, since a book's index is marked by hand. The conclusion, +which the user accepted: symbol lookups are not uphill (`names`, +`qualified` and `primary` are already a hand index, generated from +`tB/symbols.json`), but jargon is, because no ranking can find a page for +words it doesn't contain. So authors mark index entries by hand +([third round](#what-shipped-third-round-the-index-pilot)); item 6 added +about 60 terms. + +**Rules for index entries**, from the pilot and item 6: +- An entry names the page a reader wants for that term, not a summary of + the page. One main entry (`index`) per term across the site; the build + enforces it. Any number of secondary entries (`index_also`). +- Agents may draft terms and targets; **the user approves every target**. + The query goes into `eval/search_prose_queries.json` first, measured, + then the entry, measured again. +- Only a term the page's own words can't find gets an entry. Check first + where it lands (`node eval/site_search.mjs ""`). Where the page + lacks the content, write it first, measure, then mark what the new + text still can't find. +- A term matches only a query holding all its words, whole and in order, + alone or among others. Capitals and hyphens don't matter, word endings + don't (stems), but `type char` doesn't match `type character`, and + `register a dll` doesn't match `register dll`: each spelling a reader + types is its own term. +- A term whose stem is a bare name's reorders that name (`declaration` + and `Declare`). A secondary entry can overtake a page that ranks first + only on its own text, so that page takes a main entry for the term too. + A one-word term matches every query holding the word, page titles + included. The eval shows all three. + +**The user's criteria**, which govern every decision here: +- A reader either finds what they want or doesn't. A small regression is + still a miss, and being better than the old index is not the bar: that + index was nearly useless. +- Judge by what a reader typing the query wants. For a bare name the order + is type names and language elements first, then members, then enum + constants and similar, then prose. The order is a priority, not a filter: + lower tiers must still appear. +- Configuration belongs in `docs/_config.yml`, not code (for example + `search.fold_headings`). +- Ship in small steps, each committed on its own with its measured numbers. +- Existing links in the docs may be wrong or not the best. Treat them as + leads, never as evidence of the right page, and don't derive test + expectations from them. The glossary is a source of candidate *terms*, + not of targets. List doubtful links for the user; don't fix them in + passing, since which page is right is an editorial call. +- Fix what a tweak can fix first; use hand-marked index entries where the + right page can't be found from its text. A tweak that happens to fix a + handful of prose queries is overfitting, not a fix. +- A Glossary definition counts as a right answer for its term within + the top 5; it needn't be first. Where several pages answer a query, + they should all appear in the results, in the order the user chose. +- lunr's cost rules out extra query passes for now (the reason the + name-and-kind set leaves out `sub` and `member` rather than spelling a + Sub as `method`). +- Use Sonnet agents for mechanical and exploratory work. +- Review every agent's work before committing it. Agents have produced + false explanations (see [X1t](#rejected-tier-specific-exact-fields-x1t)), + a lookbehind regex that would break the whole client on Safari before + 16.4, and a loading message that blanked on the second keystroke. + +**Tools.** +- `node eval/search_quality.mjs --compare eval/search_baseline.json --worst 20` + measures a build against the saved baseline; `--save` updates it; + `--failures N` lists what misses rank 1, by category and tier. The + baseline records its ground truth (`intent-5`). +- The eval's symbol queries are one word each; its page-title and + page-plus-section sets are its only multi-word queries. So for any + change to the all-words pass, also rank every qualified symbol written + as two words (`FileListBox Name`) and every symbol as its name and kind + (`MaxHeight property`) with the committed replica and the candidate: + `eval/search-experiments/probes/spaced.mjs` and `kinds.mjs`, run once + against a copy of the committed `site_search.mjs`. +- `eval/site_search.mjs` is the replica of the client search. The site's + client and `builder/offline.mjs`'s `initSearch` must stay identical to it; + `test/search.test.mjs` fails if their fields or pipeline drift apart. +- Build first with `node builder/tbdocs.mjs --src docs --no-check --no-offline --no-pdf`. +- The research scripts and their records are in + [eval/search-experiments/](eval/search-experiments/README.md). +- Measure a candidate change with temporary knobs in the replica (an `EXP` + environment variable read by `eval/site_search.mjs`), then restore the + file from git and write the chosen version cleanly into all three + copies. +- Check anything in the client in a real browser. `.claude/launch.json` has + `docs-serve` (port 4001) and `docs-offline` (port 4002; build without + `--no-offline` first). The client runs `update()` on `keyup`, so browser + tools that insert text without key events don't trigger search; setting + the box's value and dispatching a `keyup` from script does, and is the + quick way to compare many queries with the replica. A hidden pane pauses `requestAnimationFrame`; + the client has a timer fallback for that. +- A twinBASIC sample on a page: mark it `check_build` and compile it + with `node scripts/check_examples.mjs --only "^Reference/Core/Sub\.md"` + (`check_run` only compiles, so far). To confirm what it prints, put the + code in a source tree with a `[RunAfterBuild]` Sub that starts with + `Debug.Cls`, and run `node scripts/tbrun.mjs `; `--keep` on + `check_examples` leaves a project whose `Settings` and `tbxMain.twin` + make a starting tree. No `MsgBox` in a probe: it hangs invisibly. +- Every title as a query: item 7 used a throwaway script that loads + `eval/site_search.mjs` with an absolute site path (a relative one + breaks `loadLunr`) and searches each distinct entry title; a result's + `ref` is its key in `ctx.docs`. Before item 7's entries and after: no + throws, 9 empty, 2 not in the top 50. +- `node eval/search_quality.mjs --save` needs the file name: + `--save eval/search_baseline.json`. +- `test.bat` stops at `check_axe_patch_equiv.mjs` in a worktree without + `node_modules`. Run `npm install` first for the full suite. +- In some agent shells `cmd` reports `test.bat` and `check.bat` as "not + recognized", even from the worktree, and from PowerShell too. Their + gates are plain `node` commands; run them in order, and rebuild the way + `build.bat` does first (`node builder/tbdocs.mjs --src docs + --check-audit-index`), since `check_tree_fresh.mjs` refuses a tree + older than any edited file. +- The browser pane's console keeps messages from earlier sessions' + pages. Check an error's line numbers against the served file before + chasing it. +- Rewriting WIP.Search.md with a script: pass replacement text as a + function, never as a string. In a string, `$` followed by a backtick + means "everything before the match": the `MidB$` sentence once pasted + the whole file's head into the resume section that way. + +**How lunr behaves here, learned the hard way:** +- The tokenizer tests one character at a time against `separator`, so a + multi-character alternative in the separator regex can never match. +- The index pipeline is trimmer, stop-word filter, stemmer; the search + pipeline is only the stemmer. The stop-word filter is now removed. +- BM25's length normalisation makes a short entry that mentions a term beat + a long entry about it. Tuning `b` doesn't help short of `b` = 0. +- A query of only `*` throws inside lunr. +- A clause's `boost` cancels out when it is the only clause on a field: + lunr divides each field's score by the query vector's magnitude *for + that field*. Weight such a field with its index-time field boost. +- A clause names every field unless it says otherwise, so a wildcard + clause also searches helper fields such as `exact`. +- A wildcard term is not stemmed usefully: `operator*` misses the index's + `oper`. Stem first, then add the wildcard, with `usePipeline: false`. +- The index trims non-word characters from token ends (`Date$` → `date`); + the query side didn't, until the intent step. A name that must keep them + (`#If`, `<>`) has to spell them as word characters. +- A REQUIRED clause names every field too, so it scores in helper fields + such as `exact` unless it is given the text fields. That leak once put + `With statement` first by accident, and pushed `error handling` down. +- BM25 compares a field's length with that field's average over *every* + entry, empty ones included. A field that is empty on nearly every entry + has an average near zero, so a match in it counts for almost nothing, + and counts for more as more entries fill it. The pilot's `index` field + pins its average at 1 (`pinIndexFieldLengths`). +- Every field costs a slot on every term of the whole index, not just on + the entries that use it: `add()` creates an empty object per field for + each new term. Three sparse fields took 24 MB more heap; one takes 5 MB. + Prefer encoding a variant inside one field (the pilot's secondary + entries carry one more `_`) to adding a field. +- A trailing wildcard completes a plain word to every token it begins, + including whole qualified names: `form*` reaches every `form.*` in + `qualified`. Harmless at a low field boost, it lifts a container by its + member count at a high one. +- A REQUIRED clause matches if the term is in *any* of its fields, and it + scores as well. To require a word without scoring it in some field, + give the REQUIRED clause boost 0 (its terms enter the query vector at + zero weight) and score with a second, optional clause. +- lunr 2.3.9's token set can invent words and lose real ones: its + minimisation keys nodes by `TokenSet#toString()`, which runs edge labels + into child ids (`{1 -> 656}` and `{1 -> 6, 5 -> 6}` are both `01656`). + Whether it happens depends on the whole term set, so any content or + field change can start it. Patched; see "Fixed: lunr invented words". + Check a candidate change with every title as a query, not only the + eval, since the eval never typed `a`. +- The search data must keep the page's HTML entities: the results panel + inserts titles and content with `innerHTML` and highlights by lunr's + character positions in that text. Decode inside the index instead, + per token, after the tokenizer splits (`Token#update` keeps the + position); see "Fixed: entities in the index". +- lunr 2.3.9's `Set#union` copies both sets, and `Index#query` unions + once per expanded term and field of a REQUIRED clause, so a short + wildcard word was quadratic. Patched (`accumulateSetUnions()`); see + "Fixed: slow multi-word queries". Profile before guessing: + `node --cpu-prof`. + +## The problem + +`PaintPicture` is documented on six classes (Form, PictureBox, Printer, +PropertyPage, Report, UserControl), each under `### PaintPicture` inside +`## Methods`. Searching for it returns 24 results, but the definitions come +13th to 19th, as entries titled just "Methods" that link to `#methods`. The +same happens to every member documented at `###`: `Line`, `Print`, `Cls`, the +Assert methods, and about 2,700 others. + +Three things combine: + +1. **Granularity.** [builder/search.mjs](builder/search.mjs) splits a page into + one entry per heading up to `search.heading_level`. `docs/_config.yml` does + not set it, so it is 2. A `###` member becomes part of the text of its + parent `##` entry, which is often 5 to 16 KB long. +2. **Length normalisation.** lunr scores with BM25 (`b` = 0.75). A term found + in a 10 KB "Methods" entry scores far lower than the same term in a short + entry that only mentions it, such as the glossary's "graphics method". +3. **Tokenisation.** The client splits on `/[\s\-/]+/`, not on `.`. So + `Form.PaintPicture` is the single token `form.paintpicture`, which occurs + nowhere. Qualified names find their target 0.4% of the time. + +The render step `headingLevelNormalizePlugin` ([builder/render.mjs](builder/render.mjs)) +matters here. On a page with `#` and `###` but no `##` (405 of 912 pages, +mostly `Reference/Core`), it promotes `###` to `##`, so those pages are +already split finer today. Counts below are after this step, which is what +search sees. + +## The corpus + +After normalisation: 909 h1, 2,549 h2, 3,820 h3 (on 222 pages), 31 h4, no h5 +or h6. Of the raw h3 headings, 2,761 are API members, 799 are generic +sections and 1,056 are prose. The generic ones are mostly `See Also` (412) +and `Example` (379). Nearly every h4 is an `Example` under a member. +Generic headings appear in two layouts: + +- next to the members, as on `ICustomControl`: `## Methods > ### Initialize, ### Destroy, ### See Also`; +- under a member, as on `NamedPipeServer`: `### PipeName > #### Example`. + +No page sets a heading id by hand. `uniqueSlug` handles the 33 pages that +repeat a heading, by adding `-1`, `-2` and so on. + +## Measurement + +A harness replays the client's search in Node: the same lunr build, the +same boosts, the same query construction and the same capped fuzzy fallback. +It runs against a built `search-data.json`. The ground truth is +`tB/symbols.json`, which gives 8,012 queries: + +- 2,884 bare names (`PaintPicture`). A result counts as correct if it is any + URL that documents a symbol of that name, since 18% of names exist on many + classes. +- 5,108 qualified names (`Form.PaintPicture`). Only that class's URL counts. +- 20 hand-picked prose queries ("late binding", "Option Explicit" and so on). + +It reports hit@1/5/10, MRR (capped at rank 20), results by symbol kind, cost +(size raw and gzipped, index build time in the browser, heap) and the +queries that changed rank between two configurations. The harness is +[eval/search_quality.mjs](eval/search_quality.mjs); today's numbers (the +"today (h2)" row below) are its baseline, saved as +[eval/search_baseline.json](eval/search_baseline.json). Run +`node eval/search_quality.mjs --compare eval/search_baseline.json` to measure +a change against it. + +### Results + +| configuration | entries | gzip | hit@1 | hit@10 | MRR | bare hit@10 | qualified hit@10 | prose hit@10 / MRR | +|---|---|---|---|---|---|---|---|---| +| today (h2) | 3,781 | 1,039 KB | 16.7% | 20.5% | .182 | 55.8% | 0.4% | 80% / .690 | +| h3 | 7,606 | 1,131 KB | 27.7% | 35.3% | .307 | 84.6% | 7.3% | 80% / .611 | +| h4 | 7,637 | 1,132 KB | same as h3 | | | | | | +| V2: h3, generic sections folded | 6,713 | 1,094 KB | 28.3% | 35.7% | .312 | 84.6% | 7.9% | 80% / .621 | +| V3: h3, generic sections retitled | 7,606 | 1,112 KB | 27.7% | 35.3% | .307 | 84.6% | 7.3% | 80% / .610 | +| V2 + query split on every `.` | 6,713 | 1,094 KB | 67.9% | 90.5% | .761 | 84.6% | 93.9% | 80% / .621 | +| V2 + symbols field (boost 50) + split on every `.` | 6,713 | 1,143 KB | 77.9% | 97.5% | .862 | 98.1% | 97.2% | 80% | +| V2 + symbols (boost 200) + split | | | | 97.7% | .867 | | | **75%** | +| BM25 `b` 0.5 / 0.25 on V2 + split | | | | 90.4 / 90.3% | .757 / .751 | | | | +| **recommended**: V2 + symbols (50) + smart dot split | 6,713 | 1,143 KB | **88.7%** | **97.9%** | **.929** | 98.1% | 97.9% | 80% | + +Against today, the recommended configuration makes 6,590 queries better and +38 worse, and leaves 1,384 unchanged. Its index build time in the browser is +about 1.1 s against 1.0 s today. Query latency is unchanged at about 0.03 ms. + +### What the numbers showed + +- **h4 adds nothing.** From h3 to h4, no query changes rank. +- **Splitting creates short competitors.** At h3, `vbSolid` falls from rank + 1 to about 12. Its page, `DrawStyleConstants`, is 483 characters. The new + per-property entries such as `DrawStyle` (171 characters) mention the + constant once and now outscore it. Titles play no part; every hit is on + content. Tuning `b` down to 0.25 leaves `vbSolid` at 12, and only `b` = 0 + fixes it, so BM25 tuning is not the answer. +- **Most remaining misses are enum values that live in table rows.** Of the + 444 bare names still outside the top 10 at h3: 410 have a page whose title + is not the symbol's name, mostly enum tables (`KeyCodeConstants` 100, + `MenuAccelConstants` 75, `ShortcutConstants` 49); 24 are operators; 10 are + common words (`And`, `Is`, `Not`, `With`). A symbols field fixes the first + group. +- **The symbols field generalises.** Built from a random 50% of symbols, it + still raises the other 50% from 82.4% to 87.3% in the top 10. At boost 200 + it starts to beat real prose answers (prose drops to 75%); at 50 prose is + unchanged. +- **Splitting on every `.` has two costs.** `Debug.Print` falls from rank 1 + to 61, because `debug` and `print` are both common words and the exact + token no longer counts. And `1.0`, `e.g.` and `i.e.` turn into + single-character tokens that match letter-index pages. The smart split + below fixes both and keeps every qualified-name gain. +- **Generic sections:** folding them into the entry before them is a small, + free gain. Retitling them ("PaintPicture — Example") gains nothing. + +## Design + +### 1. Split entries at h3, and fold generic sections + +- Set `search.heading_level: 3` in `docs/_config.yml`. +- In `extractSections`, after the split, append any section whose title is a + generic section name to the section before it on the same page, rather + than giving it its own entry. The list lives in config as + `search.fold_headings`: `See Also`, `Example`, `Examples`, `Remarks`, + `Parameters`, `Return value`, `Syntax`, `Notes`. It is matched without + regard to case. +- Folding applies at every level, including the h2 `Example` and `See Also` + on normalised `Reference/Core` pages. Measured separately, that is neutral + to slightly positive (V2 on h2 against h2: hit@10 21.0% against 20.5%). +- A page whose first section is generic keeps it as its own entry, since + there is nothing before it to fold into. +- The order of entries stays fixed. `discover.mjs` orders pages + deterministically because reordering used to shuffle search entries + between builds, and splitting headings top to bottom keeps that. + +### 2. Two fields built from the symbol index + +Measurement (below) used a single `symbols` field at boost 50. Building it, +BM25 discounts a match inside a long field, and `Container.Name` forms are +longer than bare names, so the two compete inside one field for no reason. +Splitting them into two fields, each boosted on its own, ranks better than +one field at any single boost. The shipped design is therefore two fields, +not one: + +- `names`: the bare name of every symbol in `tB/symbols.json` whose URL is + the entry's `relUrl` (ignoring a trailing slash or `index`), space- + separated and deduplicated. Client boost 100. +- `qualified`: the `Container.Name` form of the same symbols (only those + that have a container -- a bare statement or operator does not), space- + separated and deduplicated. Client boost 50. + +With the current corpus, 4,084 of 6,713 entries get either field. + +- In the build, `searchData` also waits for `symbolIndex`. Both already run + on the main thread after `renderJoin`. `symbolIndex` returns its symbols, + and `searchData` joins them by URL before it writes (`joinSymbolsToEntries` + in `search.mjs`). Workers keep deriving sections as they do now; the join + is a map lookup per entry. +- `renderEntryString` adds `names` / `qualified` only when non-empty, so a + page with no symbols produces the same bytes as before. +- In the client, `this.field('names', { boost: 100 })` and + `this.field('qualified', { boost: 50 })`. The title keeps 200 and content + keeps 2. + +### 3. Query construction + +In the client's `update()`, for each token the tokenizer produces: + +- **Smart dot split.** Keep the whole token as a term with boost 10, as + today. Where a `.` sits between identifier characters on both sides + (found by `/([A-Za-z_]\w*)\.(?=[A-Za-z_])/g`, not a lookbehind, which + Safari before 16.4 can't parse), also add the parts as extra terms, + with the same boost and trailing wildcard as ordinary tokens. Drop parts of + one character. So `1.0`, `3.9`, `e.g.` and `i.e.` add nothing, and + `Debug.Print` still matches its exact entry first. + Making the parts required (lunr's `presence.REQUIRED`) was measured and + rejected: qualified hit@10 falls from 97.9% to 88.6%. +- **Asterisk guard.** Drop tokens made only of `*`. If nothing is left, show + no results. Today a search for `*` or `**` throws inside lunr's query + engine and search stops working. The guard changes no metric. + +### 4. Both copies of the client + +The online client is the vendored +`builder/vendor/just-the-docs/assets/js/just-the-docs.js`, patched in place. +Add both changes to the patch list in its README. The offline build replaces +`initSearch` with its own copy, `JTD_INITSEARCH_FN_REPLACEMENT` in +[builder/offline.mjs](builder/offline.mjs), which builds its own lunr index, +so it needs the `names` and `qualified` fields too, at the same boosts. It +uses the query code in `update()` unchanged, so it inherits the query +changes. A unit test extracts the field/boost list from both sources and +asserts they agree, so they cannot drift apart silently. + +### 5. Tests and tooling + +- **Unit tests for `search.mjs`.** Today nothing checks what entries + contain; only URL coverage is checked (`checkSearch`, which strips `#`, so + it is unaffected). Add tests for the h3 split, folding (including a generic + section that comes first on a page), the symbols join and URL + normalisation, and byte stability when `symbols` is empty. +- **Fix `eval/site_search.mjs`.** It claims to reproduce the client exactly, + but never sets `lunr.tokenizer.separator`, so it tokenises with lunr's + default `/[\s\-]+/`. That is why it could not show the `*` crash. Mirror + the new field and the new query logic too. +- **Promote the harness into `eval/`**, as `eval/search_quality.mjs` next to + `site_search.mjs`, sharing its index and query code. Commit the ground-truth + derivation and the 20 prose queries, but not the built data files, and + record today's numbers as the baseline. `scripts/` does not fit, because + its tools are checks that pass or fail. +- **Comments.** The byte-parity-with-Jekyll wording in `search.mjs` is + history, since nothing has compared against Jekyll since the cutover. + Reword it where this change diverges. + +## Known regressions and remaining gaps + +- **38 queries get worse.** The largest: `VB` (rank 1→15), `Do` (1→13), + `For` (1→12), `With` (28→54), `Is` (151→176). Short keywords now also + match many `symbols` fields. Step 5 of the rollout diagnoses them and + fixes them. +- **Operators (24 symbols) remain unsearchable.** See [Future work](#future-work). +- **Prose hit@10 stays at 80%.** Four of the twenty prose queries miss + whatever the configuration, so they are content or wording problems. +- **Ground-truth bias.** Most queries are symbol names, so the numbers favour + API lookup. The symbols field comes from the same data as the ground truth. + The held-out check limits that concern, but does not remove it. The prose + set is small and hand-picked. +- **After step 5** (below), 18 queries are still worse than the *original* + (pre-rollout) baseline, not the 38 above: `VB` (rank 1→16, wildcard + crowding -- `VB` is a prefix of many `vbXxx` constant names in `names`), + `Lock`, `Time$` (same crowding, on `names`/`qualified`), `Column`, + `ColumnHeader`, `CheckBox`, `PropertyPage`, `ToolWindow`, `Node`, several + `vbXxx` bare names at rank 2 instead of 1, and the `symbol index` prose + query (rank 1→39, because `Index` is a property on many controls, so the + `names` field now outranks the prose page that used to be the only match). + "Worse than the original baseline" turned out to be the wrong criterion: + that index was nearly useless, so its rank 1 was often not what a reader + wanted. These 18 and the rest of the work list are now judged by reader + intent; see [Reader intent](#reader-intent-after-the-rollout). + +## Design §5: stop words, dot runs, and a lazy build + +Step 5 of the rollout (see "Rollout" below) is three independent, measured +changes to both client copies and the eval replica. + +### A. Keep lunr's stop words in the index + +lunr's **index** pipeline runs `lunr.stopWordFilter` by default, dropping +words like `do`, `for`, `if`, `is`, `on`, `with`, `each` from every field +before it's added. Its **search** pipeline never ran that filter -- upstream +never removes it, so a query for `Do` still carried the token `do`, which +by that point existed nowhere in the index. Since a large fraction of +twinBASIC's own keywords are English stop words, every keyword query was +guaranteed to miss its own entry (`Do` ranked 18th, `For` 17th, `With` 55th, +`Is` 177th, `On` 3rd only because of an unrelated coincidental match, `Each` +6th). The fix is one line inside the `lunr(function(){...})` builder: ` +this.pipeline.remove(lunr.stopWordFilter);`. Keeping stop words does let a +handful of short, extremely common tokens (mostly other `vbXxx` constant +name prefixes and `Index`/`Lock`/`Column`/`Time$`) crowd into more results +than before -- see "Known regressions" above -- but every one of the +keyword queries this fixes goes to rank 1, and the fix generalises (it +isn't specific to twinBASIC's keyword list). + +### B. Split runs of two or more dots + +Titles like `Do...Loop`, `For Each...Next` and `If...Then...Else` tokenised +as a single opaque token (`do...loop`), because lunr's tokenizer decides +where to split by testing *one character at a time* against `separator`; a +`\.{2,}` alternative inside that regex can never match, since no single +character is "two or more dots". The fix wraps `lunr.tokenizer` instead of +trying to extend the separator regex: for a string input, every run of 2+ +dots becomes the same number of spaces (`s.replace(/\.{2,}/g, m => new +Array(m.length + 1).join(' '))`, chosen over deleting the dots so that +character offsets used for match highlighting stay valid), then the +original tokenizer runs as usual. Non-string input (arrays, `null`) passes +through unchanged, since lunr also calls its tokenizer with those. + +The wrapper is installed once, at the same place the separator used to be +set, and it has to carry `separator` itself: the original tokenizer reads +`lunr.tokenizer.separator` at call time (not a value captured when it was +first defined), so once `lunr.tokenizer` points at the wrapper, that lookup +resolves to the *wrapper's* own `.separator` property, not the original +function's. The wrapper sets `wrapper.separator = /[\s\-\/]+/`, the client's +existing separator, and `/` still splits tokens, e.g. `a/b` tokenises to +`a`, `b` exactly as before. Since it runs on the shared global `lunr` +object, it applies to the index build and every query alike -- `update()` +already calls `lunr.tokenizer(input)` unchanged. + +### C. Build the index lazily, on the first keystroke, with visible feedback + +Before step 5, `initSearch()` ran unconditionally on every page load, +fetching `search-data.json` and synchronously building the lunr index -- +about 1.3s and 240MB of heap on a desktop (see the cost table below), paid +by every reader whether or not they ever opened search. `initSearch()` now +only defines *how* to build the index (`loadIndex(onSuccess, onError)`, +which fetches via XHR as before and installs the stop-word/dot-run patches +above) and hands that function to `searchLoaded()`, which wires up the +search box's listeners immediately but doesn't call `loadIndex()` until the +first `keyup` that leaves the box non-empty. + +That first keystroke calls `loadIndexNow()`, which shows a status message +("Loading search index...") in the same slot and style as "No results +found" (`.search-no-result`, reused rather than adding a class, since +visually it's the same single centred message) and the same text in the +`a11y-status` live region, then yields (a `requestAnimationFrame` raced by +a 100 ms timer, since frames never fire in a hidden tab, then +`setTimeout(fn, 0)`) so that message actually paints before the +synchronous, comparatively expensive index build runs on the main thread. +A keystroke during the load leaves the message in place. Checked in a +browser: the message stays up until results replace it, with no empty +panel in between, both online and in the offline tree. +When `loadIndex()` finishes, `finishLoad()` searches whatever is in the box +*at that moment* -- not whatever was typed when the load started, since the +reader may have kept typing while the fetch and build were in flight. Only +one load is ever in flight at a time (`indexLoading`, checked at the top of +`loadIndexNow()`); a keystroke that lands mid-load just returns from +`update()`, having already updated `currentInput`, and `finishLoad()`'s +fresh read of the search box at completion is what ends up searched for. A +failed load (a bad XHR status, a parse error, or an exception building the +index) shows "Search is unavailable" plus the matching `a11y-status` text +and leaves the index `null`, so the next keystroke retries from scratch. + +The restructuring keeps `update()`'s existing prefix (the dedup check, the +`search-active` class, the iOS scroll workaround) and its keyboard/focus/ +blur listeners untouched; only the part that used to run the query +unconditionally was split out into `doSearch(input)`, gated behind `if +(index === null) { loadIndexNow(); return; }`. `searchLoaded()` itself now +takes `loadIndex` instead of `(index, docs)`, since those two are no longer +available at the time it's called. + +`builder/offline.mjs`'s `JTD_INITSEARCH_FN_REPLACEMENT` gets the same +treatment: its `window.SEARCH_DATA` is already sitting in memory (preloaded +via a `