Application Overview

What this app is for

The DSLCore Orchestrator is the integration, scheduling and governance control plane for a network of DSLCore applications and the external / legacy systems wired alongside them. It exists so that cross-application behaviour — "when a work order needs a part that is out of stock, start a purchase", "when EIS approves an opportunity, create the project in EPA and ELWPM", "every night evaluate ELWPM obligations", "prove the tenement state in EPA still matches ELWPM" — lives in one place with a full, auditable history, instead of being scattered as point-to-point integrations inside each app.

The same code runs as one instance per network (its own port and database): the live orchestrator for the exploration suite, orchestrator_r9 for the work_supply network, orchestrator_work_supply for work_supply_00 — see the table in the section's introduction.

Its defining principle is a boundary:

Domain apps own domain data and business logic. The orchestrator owns execution and integration logic. It stores mappings, events, routes, schedules, controls and execution/audit state — never the master copy of a business record.

The operating chain

REGISTER            CONNECT IDENTITY         MOVE                    SCHEDULE
apps + connectors → canonical entities   →  events → routes →     jobs → schedules →
                    + ownership + maps       field maps → delivery  executions + checkpoints
                                             (→ dead-letter)
                                                         │
                                                         ▼
                                          ASSURE                    RECORD
                                       controls → executions →   health checks +
                                       governance exceptions →   audit trail
                                       actions (→ Verified)

Read left to right, that is exactly the sidebar: Registry → Integration (and Message Protocol) → Scheduling → Governance → Administration.

Two kinds of traffic

Entity sync Message protocol (message_v1)
Used by the exploration suite (eis / epa / elwpm), work_supply_00 work_supply (ten apps with an integration contract)
What moves an event; the orchestrator transforms the record and upserts it into each target a complete event, command or query envelope, delivered unchanged
Who receives it route definitions (source app + event type, optional condition) an explicit destination, the entity's owner, or every matching subscription
Defined in routes, field mappings and controls (the entity_sync section of the network's system.manifest.yaml, or the seed) the network manifest (system.manifest.yaml)
Read more below, and the Exploration Demo Runbook 04 — Message Protocol

The two never touch each other's messages: each Event Message is marked entity_sync or message_v1.

The modules

1. Registry — who is out there and how do we reach them

  • ApplicationInstance — every participating app: type (DSLCore / External / Legacy / Infrastructure), environment, base URL, health endpoint, connector, schema version and operational status.
  • Connector — how to communicate with a system: type (DSLCoreHTTP / REST / Webhook / JSON / …), direction, auth type and a credential reference (never a raw secret), timeout and retry policy.
  • CanonicalEntity — the shared business entities and their canonical key field (projectCode, tenementCode, …). This is the vocabulary that lets different apps talk about the same thing.
  • EntityOwnership — which application is authoritative for an entity, or for a single field of one (ELWPM owns a tenement's licence_status; the external ERP owns a budget period's actual_cost).
  • EntityMapping — the crosswalk: canonical key ↔ each app's local key, with a sync status.
  • HealthCheck — technical liveness history per application.

2. Integration — what moves, where, and did it arrive

  • EventDefinition — the catalogue of event types, their producer and materiality.
  • EventMessage — a received event envelope: id, type, source, correlation, canonical key, payload, status and a duplicate flag (dedup by event_sid).
  • RouteDefinition — what a trigger causes: source event/app → target app/entity/action, an operation (Create / Update / Upsert / Invoke / …), an optional condition and a retry policy.
  • FieldMapping — per-route field transforms: Direct, ValueMap, Expression, TypeCast, UnitConversion, Lookup or Constant.
  • DeliveryAttempt — one row per attempt to deliver an event down a route: timing, request/response summary, status.
  • DeadLetterItem — where deliveries go when their retries are exhausted; worked to Resolved / Ignored.

2b. Message Protocol — events, commands and queries between contract apps

  • Subscription — which application receives which message types (a pattern such as inventory.*), loaded from the network manifest.
  • RouteDecision — one per destination of a message: why it was chosen (explicit / owner / subscription), its status, attempts and next attempt.
  • Acknowledgement — the receiving application's answer (completed / rejected / duplicate) and what it created.
  • MessageReconciliation — open delivery problems (unacknowledged, stale, dead-lettered), recorded once each and closed when the route is delivered.
  • Shared with entity sync: EventMessage, DeliveryAttempt, DeadLetterItem, RetryPolicy (MSG-V1-DEFAULT).

3. Scheduling — what runs, when, and what happened

  • ActionDefinition — an invokable operation on an app (endpoint, method, idempotency).
  • JobDefinition — a schedulable unit that runs an action, with concurrency and retry policy.
  • Schedule — Cron / Interval / OneTime, timezone, next/last run and a misfire policy.
  • JobExecution — the run history: trigger, timing, attempt, status, result.
  • SyncCheckpoint — how far an incremental sync has progressed (timestamp / cursor / sequence) and whether it is Current, Stale or Failed.
  • RetryPolicy — max attempts, initial delay, back-off multiplier, max delay, dead-letter-after-failure.

4. Governance — catch absence, inconsistency and staleness — not just events

  • ControlDefinition — an assurance control of one of ten types: Existence, Absence, Reconciliation, Freshness, Timeliness, Threshold, Completeness, Cardinality, State, Sequence — with scope, authoritative vs compare app, match key, the fields/tolerance it checks, severity, owner role and schedule.
  • ControlExecution — each run of a control: Pass / Fail / Error, records evaluated and exceptions created.
  • GovernanceException — a persistent finding worked through the lifecycle Open → Acknowledged → Investigating → Remediation → Resolved → Verified → Closed (or Accepted), carrying expected vs actual state, owner, due date and escalation level.
  • ExceptionAction — the audit-grade action log on an exception (who did what, and the resulting status).

5. Administration

  • AuditRecord — the integration/governance audit trail: actor, action, object, before/after and correlation id.

Two ideas that make it trustworthy

Verification is rerun-and-pass. A governance exception cannot jump to Verified on someone's say-so. The originating control is rerun; only if it now passes is the exception verified and eligible to close. In the demo, CTL-TENEMENT-RECON fails and raises EXC-0001; after EPA is corrected the same control is rerun (a second ControlExecution that passes) and only then does EXC-0001 move Verified → Closed.

Delivery is at-least-once; targets must be idempotent. Routes can retry, so a target may see the same upsert twice. Every action is modelled as idempotent (create-or-update by canonical key). Events are deduplicated by event_sid so a replay (EVT-0001-DUP in the demo) is ignored rather than re-delivered.

Build status (updated 2026-10-06)

The framework renders the whole control plane as AI-Safe CRUD, and the runtime engine (server/lib/orchestrator_engine/, peer-facing internal API in server/lib/internal_api/) runs in any instance started with ORCHESTRATOR_ENGINE=1 — one engine per instance, in the process that serves requests.

  • Entity sync — event dispatcher (event → matching routes → condition → field transform → delivery, dead-lettering exhausted failures), transform evaluator (Direct / Constant / ValueMap / TypeCast; Expression deliberately deferred as a code-execution risk), control evaluators (Reconciliation, Freshness; Existence / Absence record a Skipped execution until their target config is restored), rerun-and-pass verification, scheduler + retry / dead-letter. Covered by server/tests/test_orchestrator_flows.py (the four acceptance flows, against a fake transport).
  • Message protocol — ingest, routing (explicit / owner / subscription), delivery with retries, dead letters, acknowledgements and reconciliation; push delivery (a stored message wakes the engine), per-tick handling of unreachable apps, wake-on-recovery, a dead-letter rule that waits for the whole retry window, engine-created indexes. Covered by server/tests/test_orchestrator_messages.py and the ten-app acceptance suite; load-tested in R12 (docs/implementation/r12_load_testing.md).

Running where: locally, orchestrator_r9 (work_supply, engine every 2 s plus push) and orchestrator_work_supply (work_supply_00, every 15 s) run with their networks. The live orchestrator.dslcore.net has the internal API mounted with the engine dormant; turning it on there, and deploying a message-protocol network, is part of deployment (R14).