Commit graph

6 commits

Author SHA1 Message Date
544934dc57
refactor(examples): replace sector-specific example material with generic, fictitious examples
Reference files, test fixtures, the playground demo project and one design
document now use generic, fictitious examples (buildings, energy, water,
grants, municipal services). The playground demo (17 fixtures plus the
embedded demo state) tells one consistent story: a municipal customer
chatbot that pre-screens housing-benefit applications, classified under
Annex III point 5(a). The embedded demo copies were edited in place rather
than regenerated, because they already carry newer AI Act dates than the
fixture files.

Legal text is unchanged. Test semantics are unchanged. Four dark-theme
onboarding screenshots with outdated placeholder text are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:03:53 +02:00
6224487987 fix(ms-ai-architect): RX-REG-KB fixtures-halen — 2027-08-02→2027-12-02 Annex III + Transparens-label i fixtures/playground [skip-docs] 2026-07-15 09:36:44 +02:00
6af3624b17 feat(ms-ai-architect): Sesjon 4 - verifisering-ut (lag 5) + adversarial regresjonstest
Lag 5 kjører ETTER transformasjon (lag 4) og FØR en kandidat-endring skrives til
en KB-fil. Den fanger regresjons-klassen: en status-påstand «korrigert» mot en
tilfeldig sitert side og stille auto-applyet (den kjente agentic-retrieval-
regresjonen — hovedkontekst måtte rette manuelt).

Regel (spec §21): status-påstander (GA/preview/versjon/pris) flagges ALLTID for
operatør, aldri auto-applyet — uansett hvor sikker evidensen ser ut.

- lib/verify-out.mjs: ren klassifiserer, null deps (speiler decisions-io).
  detectStatusClaim(text) → {isStatus, kinds}; classifyChange(change) →
  {verdict: flagged|auto-applied, status_claim, reasons}. Tre flag-regler:
  status-gate (§21) · adversarial refutering · autoritets-mismatch (regresjonens
  rotårsak). Konservativ med vilje: ved tvil flagges. SKRIVER ALDRI.
- Fixtur tests/fixtures/kb-update/agentic-retrieval-regression.json: den kjente
  regresjonen (flat «GA» mot nyansert «delvis GA … resten preview»).
- TDD: 13 tester før kode, inkl. import-invariant (ingen write-utils).
- Wiret i kb-update.md §4 c2 (lag 5 mellom identifiser-endring og skriv).

Kriterium møtt: fixtur → flagged:true, ikke auto-applied.
Tester: validate 239 · kb-update 95 (+13) · kb-eval 13 · kb-integrity 115/115.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01REiKFhP4w6xGXXqWKpPCJJ
2026-06-19 22:00:36 +02:00
665836caf1 test(ms-ai-architect): add ros-analysis fixture for E2E suite
Synthetic ROS-analyse output for "Acme Kunde-chatbot" (Acme Kommune)
following the same pattern as security-assessment, cost-estimation,
ai-act and summary fixtures. Satisfies all 29 assertions in
tests/test-ros-output.sh:

- 8 phases (Fase 1-8) plus Ledelsessammendrag
- 12 trusler i T-XXX-NN format (MAESTRO + OWASP-mapping)
- 9 risikoer i R-N format
- 10 tiltak i M-N format
- 7 ROS-dimensjoner med X/5-scoring
- 5x5 risikomatrise + restrisiko-tabell
- NS 5814 + ISO 31000 metodikk-referanser
- AI Act, GDPR, OWASP regulatoriske referanser
- MAESTRO + supply-chain referanser (Vedlegg O coverage)

Tar bort den siste pre-eksisterende run-e2e-feilen
(`bash tests/run-e2e.sh` exits 0).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-05 12:38:49 +02:00
Kjell Tore Guttormsen
781d98f62f chore(privacy): scrub real-org references from plugin internals (phase 2)
Same bulk replacement applied to plugin-internal KB, examples, fixtures,
tests, and docs. Real organization names, persona names, internal system
identifiers, and domain-specific terms replaced with fictional generic
public-sector entity (DDT) and generic terminology.

Scope:
- okr/ — examples, governance, framework, integrations, sources
- ms-ai-architect/ — KB references (engineering, governance, security,
  infrastructure, advisor), tests/fixtures, agents, docs
- linkedin-thought-leadership/ — voice samples, network-builder,
  examples (genericized identifying headlines to "[your organization]")
- llm-security/ — research notes, scan report

Manual genericization beyond bulk replace:
- okr SKILL.md "Primary user / Domain" — generic Norwegian public sector
- linkedin-voice SKILL.md headline placeholder
- network-builder.md headline placeholder
- high-engagement-posts.md voice sample employer line + hashtag

Phase 3 (factual-attribution review) remains: a few KB files attribute
publicly known transport-sector docs/datasets (e.g. håndbok V440, NVDB)
to the fictional DDT after bulk replace. Needs manual semantic review
to either remove or restore correct citation without re-introducing
affiliation references.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-03 04:28:15 +02:00
Kjell Tore Guttormsen
baa2d0220b feat(ultraplan-local): v1.6.0 — /ultraresearch-local deep research command
Add /ultraresearch-local for structured research combining local codebase
analysis with external knowledge via parallel agent swarms. Produces research
briefs with triangulation, confidence ratings, and source quality assessment.

New command: /ultraresearch-local with modes --quick, --local, --external, --fg.
New agents: research-orchestrator (opus), docs-researcher, community-researcher,
security-researcher, contrarian-researcher, gemini-bridge (all sonnet).
New template: research-brief-template.md.

Integration: --research flag in /ultraplan-local accepts pre-built research
briefs (up to 3), enriches the interview and exploration phases. Planning
orchestrator cross-references brief findings during synthesis.

Design principle: Context Engineering — right information to right agent at
right time. Research briefs are structured artifacts in the pipeline:
ultraresearch → brief → ultraplan --research → plan → ultraexecute.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-08 08:58:35 +02:00