Operatørdirektiv 14.08 etter ekstern test: mottakeren skal få kjørbar Python, ikke en Docker-innpakking. `Dockerfile` og `azure.yaml` er SLETTET fra treet. Sømmen er valgt av den eksisterende invarianten, ikke av smak: pakka ER `git archive HEAD`, så å ekskludere filene fra arkivet ville krevd en kurerings-mekanisme — den andre kopien av «hva mottakeren får», fri til å drifte fra HEAD (kø-(p)). Fjerning holder arkivet ukurert og gjør fraværet til en egenskap ved HEAD, som er det eneste en gate kan måle. De to gatene som pinnet flaten er håndtert bevisst: - 4e-rå-tekst-gaten (`--platform linux/amd64` + ÉN kopi av startkommandoen) er SLETTET, med et notat der den sto. En gate som pinner en fjernet flate kan bare bli grønn. - handover-gatens `_REQUIRED_MEMBERS` er ikke bare fratatt de to navnene, men erstattet av en POSITIV fraværs-assert pluss en dokument-gate. Å kun slutte å KREVE dem ville gitt en gate som ikke kan skille «fjernet» fra «shippes fortsatt». Startkommandoen har nå ÉN kopi igjen: DEPLOY.md-ens `python main.py`, som navngir inngangen subprosess-testen faktisk kjører.
197 lines
9.4 KiB
Markdown
197 lines
9.4 KiB
Markdown
# Deploying and using this package
|
||
|
||
This archive is a complete, self-contained copy of `portfolio-optimiser`. You do not need to clone
|
||
anything or have an account anywhere to deploy it. Unzip it and follow this document.
|
||
|
||
It is written in English because its readers are implementing developers and platform operators.
|
||
The two documents that describe the *human* side of the loop are written in Norwegian and ship in
|
||
this same archive: [`docs/bestille-en-kjoring.md`](docs/bestille-en-kjoring.md) (ordering a run) and
|
||
[`docs/ekspert-svar.md`](docs/ekspert-svar.md) (giving the verdict).
|
||
|
||
---
|
||
|
||
## 1. What this is — and what it is not
|
||
|
||
It finds cost savings **inside** a single project: agents propose candidate measures, a
|
||
**deterministic validator** decides whether the numbers hold, a second agent (the "checker") gates
|
||
the *reasoning*, and a domain expert gives the final verdict. The system learns from those verdicts.
|
||
|
||
**There is no chat interface, and that is a design decision rather than a missing feature.**
|
||
|
||
The hosted surface is a task API — `POST /invocations` with a JSON body, `GET /readiness` — not a
|
||
conversation. The system is deliberately **not** exposed as a Microsoft Foundry agent via
|
||
`as_agent()`, because the validator, the cost-baseline anchoring, the checker gate and the savings
|
||
ledger all live *outside* the agent graph. Wrapping it as a chat agent would route requests around
|
||
precisely the components that make an answer trustworthy, and you would get a fluent assistant
|
||
instead of a validated result.
|
||
|
||
The dialogue with humans happens in two places instead, both file-shaped and both auditable:
|
||
|
||
- **before** a run, through a mandate file (what should be considered), and
|
||
- **after** a run, through a verdict file (what a qualified person decided).
|
||
|
||
If you want an interactive front end, build it on top of `POST /invocations` — but keep the verdict
|
||
step human. An approval loop that a model can close by itself is not an approval loop.
|
||
|
||
---
|
||
|
||
## 2. Who does what
|
||
|
||
The package has three distinct audiences. Only the first one is technical.
|
||
|
||
| Role | What they do | Where it is documented |
|
||
|---|---|---|
|
||
| **Platform operator** | Installs the package, sets two environment variables, starts the service, verifies with the offline preflight. One-time work. | This document, sections 3–6. |
|
||
| **Ordering domain expert** (*bestiller*) | Writes the mandate file: the objective, which approaches are worth trying, and why. The *why* is the part the system cannot read out of cost figures. | `docs/bestille-en-kjoring.md` |
|
||
| **Reviewing domain expert** (*fagperson*) | Receives a validated proposal and answers `approved` or `rejected` with a written rationale. That verdict is what the system learns from. | `docs/ekspert-svar.md` |
|
||
|
||
**The ordering expert steers what is *considered*, never what is *approved*.** If you ask for
|
||
something the numbers do not support, the deterministic validator rejects it and the rejection comes
|
||
back with its reason. That property is the point of the whole design; do not build around it.
|
||
|
||
---
|
||
|
||
## 3. The process, end to end
|
||
|
||
1. **Prepare a knowledge base** for the project — cost documents, technical sources, prior
|
||
experience — as a document bundle. A working five-document example ships in
|
||
`shared/examples/bygg-energi-mikro`, and `docs/knowledge-base-recipe.md` describes how to build
|
||
your own.
|
||
2. **Order a run** with a mandate file (optional but recommended). See
|
||
`docs/bestille-en-kjoring.md`.
|
||
3. **Run it** — from the CLI, or by `POST /invocations` against the running service (section 6).
|
||
The run navigates the knowledge base, debates candidate measures, validates the numbers against
|
||
the project's actual cost baseline, and returns either a **validated proposal** or a
|
||
**rejection with a reason**. A rejection is a successful run: the negative outcome belongs to the
|
||
answer, not to the transport, and the API returns HTTP 200 for it.
|
||
4. **A domain expert judges the outcome** and writes a verdict — either inline during the run, or as
|
||
a JSON file dropped into an inbox directory afterwards, possibly days later. See
|
||
`docs/ekspert-svar.md`.
|
||
5. **Approved verdicts can be promoted into the knowledge base**, so the next run reads them as
|
||
context and does not re-propose something the organisation has already rejected. Promotion is
|
||
gated: only an approved verdict can be promoted, and raw agent output never enters the knowledge
|
||
base on its own.
|
||
|
||
Steps 4 and 5 are what make this a learning loop rather than a report generator.
|
||
|
||
---
|
||
|
||
## 4. Prerequisites in your own tenant
|
||
|
||
You need, in Microsoft Foundry:
|
||
|
||
1. A **Foundry resource** with project management enabled and a **project** inside it. A Foundry
|
||
project is required — an Azure OpenAI resource (`kind: OpenAI`) is a different client surface and
|
||
will not work.
|
||
2. At least one **model deployment** (any current chat model; a small one is enough).
|
||
3. The **Foundry User** role (role definition id `53ca6127-db72-4b80-b1b0-d745d6d5456d`) for the
|
||
identity that will run it, assigned at project scope. Azure `Owner` and `Contributor` grant
|
||
management permissions only, **not** the data-plane permissions this needs.
|
||
|
||
---
|
||
|
||
## 5. Configuration — the two variables that matter
|
||
|
||
**5.1 The project endpoint.** Resolved as the first non-empty of:
|
||
|
||
```
|
||
PORTFOLIO_FOUNDRY_PROJECT_ENDPOINT # yours; wins if set
|
||
FOUNDRY_PROJECT_ENDPOINT # injected by the Foundry hosting platform
|
||
```
|
||
|
||
Where something injects the second one for you — the Foundry hosting platform does — you need to
|
||
set nothing. Everywhere else (a laptop, a pipeline, your own runtime) set the first. The value looks
|
||
like `https://<resource>.services.ai.azure.com/api/projects/<project>`.
|
||
|
||
**5.2 The model map — this is the one that catches people.** The packaged
|
||
`src/portfolio_optimiser/data/model_map.json` ships **placeholders**:
|
||
|
||
```json
|
||
"azure": { "default": "REPLACE-WITH-FOUNDRY-DEPLOYMENT", ... }
|
||
```
|
||
|
||
Deployment names are tenant-specific, so they are deliberately never committed. The code
|
||
**fail-fasts** on a `REPLACE-WITH-` value rather than guessing. Point `PORTFOLIO_MODEL_MAP` at your
|
||
own file:
|
||
|
||
```json
|
||
{
|
||
"local": { "default": "qwen3:4b", "proposer": "qwen3:4b", "checker": "qwen3:4b" },
|
||
"azure": { "default": "my-deployment", "proposer": "my-deployment", "checker": "my-deployment" }
|
||
}
|
||
```
|
||
|
||
```bash
|
||
export PORTFOLIO_MODEL_MAP=/path/to/model_map.json
|
||
```
|
||
|
||
If you skip this, the service starts and answers `/readiness` — and fails every invocation. That
|
||
is the single most likely first-deployment failure, which is why it has its own section.
|
||
|
||
**5.3 Verify before you pay for a model call:**
|
||
|
||
```bash
|
||
uv run python -m portfolio_optimiser.preflight --profile azure
|
||
```
|
||
|
||
Exit code 0 means every offline-checkable precondition holds: endpoint shape, model-map structure,
|
||
no placeholders left. It is **necessary but not sufficient** — RBAC (403), token or tenant problems
|
||
(401) and a well-formed but non-existent deployment (404) can only surface at the live call.
|
||
|
||
---
|
||
|
||
## 6. Running it
|
||
|
||
**This package is Python and nothing else.** There is no image to build and no platform-specific
|
||
deployment manifest — that was removed deliberately after an external trial, because it forced a
|
||
container toolchain on receivers who only needed to run the code. How the process is supervised,
|
||
packaged or placed in your environment is your decision, and this package does not pre-empt it.
|
||
|
||
Install the locked environment and start the service:
|
||
|
||
```bash
|
||
uv sync --frozen --no-dev
|
||
uv run python main.py
|
||
```
|
||
|
||
`uv sync --frozen` installs the exact resolution every measurement in this repository ran against —
|
||
never a fresh resolve, which picks other versions. Two dependencies are git-tag-pinned direct
|
||
references, so **`git` must be on PATH** for the install to work.
|
||
|
||
`main.py` serves port 8088 (override with `PORT`), answers `GET /readiness`, accepts
|
||
`POST /invocations`, and exits cleanly on SIGTERM. For a single run without the service, the CLI is
|
||
the same code path:
|
||
|
||
```bash
|
||
uv run portfolio-optimiser --help
|
||
```
|
||
|
||
**The invocations payload** is whitelisted against the run signature — unknown fields are refused by
|
||
name with HTTP 400, never silently ignored:
|
||
|
||
| Field | Required | Notes |
|
||
|---|---|---|
|
||
| `project_id` | yes | The project to analyse. |
|
||
| `docs_dir` | yes | Source documents. |
|
||
| `verdict_input` | yes | The expert verdict for this run. |
|
||
| `bundle_dir` | no | Knowledge base to navigate. |
|
||
| `profile` | no | Defaults to `azure` on this HTTP surface. |
|
||
| `max_rounds` | no | Debate round cap. |
|
||
| `max_tokens` | no | **Per-invocation token ceiling.** Set it. The built-in default is 100 000 tokens per run, which is a ceiling nobody chose for your workload. |
|
||
| `top_k` | no | Retrieval breadth. |
|
||
|
||
HTTP 400 means the request was malformed; 500 means the run itself failed, with `error_type` and
|
||
`error` in the body. A rejected proposal is **200** — see section 3.
|
||
|
||
---
|
||
|
||
## 7. Honest limits
|
||
|
||
- **This framework has never been run against a live model in our environment.** Everything above is
|
||
measured against tests and an offline simulation. The first live run in *your*
|
||
tenant is the first live run, full stop.
|
||
- The system is a **technical framework**. Data protection assessments, lawful basis and risk
|
||
ownership belong to whoever deploys it — see `README.md` and `SECURITY.md`.
|
||
- Readiness does not respond during CPU-bound stretches of a run (the optimisation solve).
|
||
- Chunked request bodies are not supported.
|
||
- Verdict promotion writes to the knowledge base non-atomically; it assumes a single writing process.
|