# SAIG Migration Connector — Manifest & Integration Spec

**Version:** 0 (draft) · **Status:** open for app manifests · **Owner:** SAIG

A single, input-agnostic migration service owned by SAIG. Any ecosystem app (and, later, any
external SaaS) plugs in so a customer arriving from an existing tool — a spreadsheet today, a
rival CRM/finance tool tomorrow — onboards in minutes with their data intact and cleanly mapped.

The engine is built once and reused everywhere. Each **target app** does only two small things:

1. Publish a **Migration Manifest** (this document defines the format) describing its entities,
   fields, relationships, validation and dedup keys.
2. Expose a **bulk Import API** that the engine writes to (contract defined below).

Everything else — ingest, source understanding, field mapping, dry-run, legal review, execute,
verify — lives in the engine and the human-in-the-loop review UI. Apps never build mapping logic.

---

## 1. Architecture & boundaries

- **Engine (SAIG-owned):** ingest adapters, source profiler, AI mapper, dry-run planner, legal
  module, execute/rollback engine, verifier, review UI, connector SDK.
- **Target app (per app):** one manifest + one Import API. That's the whole integration surface.
- **Source:** spreadsheets (CSV/XLSX) first; later API/DB-dump adapters and named-SaaS adapters.
- **Identity:** every migration is scoped to a single **`brain_id`** — the SAIG brain minted at
  login (the same login-derived id used everywhere in the ecosystem; **no company-number / CRN
  matching**). Migrated *people* are mapped to SAIG identities so the same person resolves to the
  same brain across apps (see §6).

```
 SOURCE ──ingest──▶ PROFILE (AI) ──┐
                                   ├─▶ MAP (AI, field-by-field, confidence) ──▶ DRY-RUN ──▶ LEGAL ──▶ EXECUTE ──▶ VERIFY
 TARGET MANIFEST ──────────────────┘            (human review at MAP and DRY-RUN)
 TARGET IMPORT API ◀──────────────────────────────────────────────────────────────── writes here
```

Nothing is written to the target until **execute**. MAP and DRY-RUN are reviewable and reversible.

---

## 2. Pipeline (contractual stages)

1. **Ingest** — read the source (CSV/XLSX now). Normalise to a tabular intermediate: one logical
   table per source entity, with raw column headers + sampled rows preserved.
2. **Profile (AI)** — infer entities, field types, semantics (e.g. "this column is an email"),
   and relationships (which columns are foreign keys to which entity).
3. **Map (AI + human)** — for each target field in the manifest, propose a source field +
   transformation + a **confidence score**. Low-confidence and required-but-unmapped fields are
   flagged for human review. Transformations are drawn from a fixed library (§5).
4. **Dry-run** — apply the mapping in memory, in dependency order, producing: per-entity counts,
   a sample of fully-transformed records, a validation report (rule failures, dedup collisions,
   unresolved references). **Nothing is written.**
5. **Legal & compliance** — generate the migration's legal artifacts and notification plan (§7).
   Human sign-off gate.
6. **Execute** — write to the target's Import API in dependency order. Idempotent, resumable,
   audited; supports rollback (§4).
7. **Verify** — reconcile written counts against planned counts and run integrity checks
   (referential integrity, dedup, required-field coverage). Produce a reconciliation report.

---

## 3. The App Migration Manifest

A JSON document the app publishes (recommended path: `https://<app>/ecosystem/migration.json`).
The engine reads it to understand the target. Top-level shape:

```json
{
  "app": "minaos",
  "manifestVersion": 1,
  "specVersion": 0,
  "importApi": {
    "baseUrl": "https://<app-domain>/api/ext/migrate",
    "auth": "saig-service-token",
    "batchMax": 500
  },
  "entities": [ Entity, ... ]
}
```

### 3.1 Entity

```json
{
  "name": "customer",
  "label": "Customer",
  "writeOrder": 10,
  "dedupKeys": [["email"], ["phone"]],
  "fields": [ Field, ... ]
}
```

- **`writeOrder`** (integer) — ascending dependency order. Parents (lower number) are written
  before children that reference them. The engine sorts the execute phase by this.
- **`dedupKeys`** — array of key-sets; a record matching any existing record on a full key-set is
  treated as a duplicate (the engine upserts rather than inserts). Each key-set is an AND of
  fields; multiple key-sets are OR.

### 3.2 Field

```json
{
  "key": "email",
  "label": "Email address",
  "type": "email",
  "required": true,
  "pii": true,
  "unique": true,
  "enum": null,
  "ref": null,
  "description": "Primary contact email; used as a dedup key.",
  "validation": { "maxLength": 254 }
}
```

| Field        | Meaning                                                                                  |
|--------------|------------------------------------------------------------------------------------------|
| `key`        | Stable machine name the Import API expects.                                               |
| `type`       | One of the canonical types in §3.3.                                                       |
| `required`   | If true and unmapped/empty, the migration cannot execute until resolved.                  |
| `pii`        | Marks personal data → feeds the legal PII inventory (§7). **Set this honestly.**          |
| `unique`     | Values must be unique within the entity.                                                  |
| `enum`       | Allowed values (array) for `type: "enum"`; the mapper normalises source values to these. |
| `ref`        | `{ "entity": "customer", "field": "id" }` — declares a foreign key (drives `writeOrder`). |
| `validation` | Optional constraints: `minLength`, `maxLength`, `min`, `max`, `pattern`.                  |

### 3.3 Canonical field types

`string` · `text` · `integer` · `decimal` · `boolean` · `date` · `datetime` · `email` · `phone`
· `url` · `currency` (amount + ISO-4217 code) · `enum` · `json` · `ref` (foreign key) ·
`person_name` (the mapper can split a single source "name" into given/family across two fields).

The engine maps source data **to** these types; the app never sees raw source formats.

---

## 4. Import API write contract

The engine calls the app's `importApi.baseUrl` per entity batch.

**Request** — `POST {baseUrl}`

```
Headers:
  Content-Type: application/json
  x-saig-service-token: <SAIG_SERVICE_TOKEN>
  X-SAIG-App: migrate
Body:
  {
    "migrationId": "mig_...",         // stable per migration run
    "brainId": "<target brain_id>",   // whose data this is
    "entity": "customer",
    "mode": "upsert",                 // "insert" | "upsert" (upsert keyed on dedupKeys)
    "records": [
      { "_srcId": "row-42", "email": "a@b.com", "givenName": "A", ... },
      ...
    ]
  }
```

- **`_srcId`** — the engine's stable id for the source row. The app MUST echo it back per record
  so the engine can map source→target ids and resume safely.
- **Idempotency (required):** the app keys each write on `(migrationId, entity, _srcId)`. A repeat
  of the same `_srcId` under the same `migrationId` MUST NOT create a second row — update or no-op.
  This is what makes execute resumable after an interruption.

**Response** — `200`

```json
{
  "results": [
    { "_srcId": "row-42", "id": "cus_abc", "status": "created" },
    { "_srcId": "row-43", "id": "cus_def", "status": "updated" },
    { "_srcId": "row-44", "status": "skipped", "reason": "duplicate" },
    { "_srcId": "row-45", "status": "error",   "error": "phone invalid" }
  ]
}
```

- The app returns the **target `id`** per `_srcId` (needed so the engine can resolve `ref` fields
  in dependent entities written later).
- Per-record `status` ∈ `created | updated | skipped | error`. Errors are row-local; one bad row
  never fails the batch.

**Rollback:** if the app supports it, `DELETE {baseUrl}?migrationId=mig_...` removes everything
written under that migration. If it can't, it declares `"rollback": false` in the manifest's
`importApi` and the engine falls back to a documented compensating report instead.

---

## 5. Transformation library (fixed, auditable)

The mapper may only apply transformations from this list (each is logged in the audit trail):

- **Date/datetime:** parse heterogeneous formats → ISO-8601; timezone normalisation.
- **Currency:** split "£1,200.00" → `{ amount: 1200.00, code: "GBP" }`; code inference.
- **Name:** split one `person_name` into given/family; title-case.
- **Enum normalisation:** map source values to the target `enum` (e.g. `"WON"`/`"closed-won"` →
  `"won"`), with a per-value mapping table surfaced for human review.
- **Phone:** normalise to E.164 where a country is known.
- **Boolean:** `"yes"/"1"/"true"` → `true`.
- **Dedup/merge:** collapse rows matching a `dedupKey`.
- **Default/constant:** fill an unmapped non-required field with a constant.
- **Trim/case/split/concat:** basic string ops.

No free-form code execution. New transforms are added to the engine, versioned, never ad-hoc.

---

## 6. Identity & brain scoping

- A migration targets exactly one **`brain_id`** (passed on every Import API call). It's the
  login-minted SAIG brain for that org — **never** resolved by company number.
- When the source contains **people who are platform users** (staff, agency owners), the engine
  maps them to **SAIG identities** so the same person is the same brain across apps. Customer/
  contact records that are not platform logins are migrated as plain data under the org's brain.
- Finance law still holds: migrating financial data into an app does not let that app recompute
  finances — Hisaab computes, others present. A finance migration lands records, not recalculations.

---

## 7. Legal & compliance module (first-class)

For every migration the engine produces, for human/legal review (never as authoritative advice):

- **PII inventory** — every field marked `pii: true`, by entity, with row counts.
- **Lawful-basis check (UK-GDPR/GDPR)** — prompts for the basis under which the data is processed
  and migrated; flags special-category data; flags whether consent needs re-confirming.
- **Retention & cross-border flags** — retention expectations per entity; a flag if the source or
  destination implies a cross-border transfer.
- **Notification plan + draft notices** — who must be told (customers, staff, the source vendor as
  a processor, regulators *only where genuinely triggered*), when, and a **draft** message for each.
  A migration is not a breach, so regulator notification is rarely required — the module says so
  rather than over-warning.

All outputs are clearly labelled **drafts for review**. This is a record-keeping aid, not legal
sign-off.

---

## 8. Versioning & publishing

- This spec is versioned by `specVersion` (currently `0`). Apps record the `specVersion` they
  implement in their manifest.
- Apps publish their manifest at `https://<app>/ecosystem/migration.json` and (optionally)
  register its URL in the ecosystem manifest so the engine can discover it.
- Breaking changes bump `specVersion`; the engine supports the current and previous version.

---

## 9. Build sequence (for the engine, after this spec)

1. **This manifest spec** — done (v0). Apps can author manifests now; MinaOS publishes first.
2. **Vertical slice:** spreadsheet → MinaOS, end-to-end (ingest → map → dry-run → execute →
   verify) through the review UI, on one source and one target, before generalising.
3. **Execute engine** — the idempotent/resumable/rollback core (the real work) + the connector SDK.
4. **Legal module v1** — PII inventory + lawful-basis check + draft notices.
5. **Generalise** — more source adapters (API/DB dumps, named SaaS) and more target manifests.

The engine is a separate service. The integration surface for every app, forever, is just §3 +
§4: one manifest, one Import API.
