fix(ms-ai-architect): registry-herding Fase 1a — +5 sitemap-prefiks + skjemaløs URL-ekstraksjon

Update-mekanismen lot siterte URL-er stå som not_in_sitemap fordi deres docset
aldri ble pollet, og skjemaløse citater (tabellceller: bare learn.microsoft.com/...
uten https://) ble aldri ekstrahert.

- Taxonomy: +5 docsets (graph, ai-builder, power-apps, power-automate,
  microsoftsearch) — hver ett child-sitemap, verifisert live mot indeksen.
  Bevist: 21/24 not_in_sitemap-URL-er i disse docsetene blir nå tracked.
- url-normalize: extractUrls fanger skjemaløse citater-med-sti (krever /path
  etter domenet → avviser bare-domene-prosa og JSON-eksempler); normalizeUrl
  kanonikaliserer scheme til https. Bevist: +19 nye URL-er ekstraherbare.
- Bugfix: backtick fra inline-kode-citat (`learn.microsoft.com/...`) lekket inn
  i URL-en — ekskludert i regex + trailing-strip.

TDD: tests/kb-update/test-url-normalize.test.mjs (ny, 14) + test-taxonomy
prefiks-count 18→23 + Fase 1a-assertion. Full suite 338/338, 0 regresjon.
Registry-refresh (build-registry --merge + poll) bevisst utsatt — unngår
552KB re-order-churn; effekten er empirisk verifisert read-only.
This commit is contained in:
Kjell Tore Guttormsen 2026-06-26 10:15:46 +02:00
commit e74646d395
4 changed files with 152 additions and 7 deletions

View file

@ -4,6 +4,7 @@
/**
* Normalize a learn.microsoft.com URL to a canonical form.
* Rules applied in order:
* 0. Canonicalise scheme accept schema-less + http(s) citations, store https
* 1. Strip trailing punctuation leaked from markdown
* 2. Strip fragment (#anchor)
* 3. Strip ?view= and other query params
@ -18,8 +19,16 @@ export function normalizeUrl(raw) {
let url = raw;
// 0. Canonicalise scheme — citations appear schema-less (table cells: bare
// learn.microsoft.com/...) or as http; strip any leading scheme and force
// https so a schema-less citation matches the same key as its https form
// (and the https sitemap entries) instead of creating a duplicate entry.
url = url.replace(/^https?:\/\//i, '');
url = 'https://' + url;
// 1. Strip trailing punctuation that leaked from markdown context
url = url.replace(/[.,;:!?'")}\]]+$/, '');
// (incl. backtick from inline-code citations: `learn.microsoft.com/...`)
url = url.replace(/[.,;:!?'")}\]`]+$/, '');
// 2. Strip fragment
const hashIdx = url.indexOf('#');
@ -43,18 +52,23 @@ export function normalizeUrl(raw) {
/**
* Extract all learn.microsoft.com URLs from a text string.
* Handles all 5 citation formats found in reference files:
* Handles all citation formats found in reference files:
* - Markdown links: [text](https://learn.microsoft.com/...)
* - Bare URLs on their own line
* - URL: prefix format
* - Dash-bullet format
* - Table cell format
* - Schema-less citations: bare learn.microsoft.com/<path> (no https://)
* The scheme is optional but a path (`/...`) after the domain is REQUIRED this
* captures schema-less citations-with-path while rejecting bare-domain prose
* mentions ("(learn.microsoft.com, blogs...)") and JSON example strings
* ({"url": "learn.microsoft.com"}), which carry no trackable source.
* @param {string} text
* @returns {string[]} array of normalized unique URLs
*/
export function extractUrls(text) {
if (!text) return [];
const regex = /https:\/\/learn\.microsoft\.com[^\s)"'<>\]|]+/g;
const regex = /(?:https?:\/\/)?learn\.microsoft\.com\/[^\s)"'<>\]|`]+/g;
const seen = new Set();
const results = [];
let match;