04 — Message Protocol

The orchestrator carries two kinds of traffic, and they never touch each other's messages (every Event Message row is marked entity_sync or message_v1):

  • entity sync — events matched to route definitions and field mappings, with the orchestrator transforming records between apps (the exploration suite, eis / epa / elwpm, and the older work_supply_00 network);
  • the message protocol (message_v1) — applications built with an integration contract exchange complete events, commands and queries; the orchestrator routes and delivers each envelope unchanged. This page is about this protocol.

The message protocol is what the work_supply network uses: ten applications — Work Management, Inventory, Purchasing, Catalogue, Supplier, SAP AP, Asset Register, Asset Lifecycle, Warranty and Assurance — on their own orchestrator instance, orchestrator_r9. Each network has its own orchestrator instance and database; the code is the same.

How a message moves

sequenceDiagram
  participant A as Sending app
  participant O as Orchestrator
  participant B as Receiving app(s)
  Note over A: business change + outbox row, one transaction
  A->>O: PUT /_orchestrator/messages/{id} (envelope)
  O->>O: store Event Message (dedupe by id) and choose destinations (Route Decisions)
  Note over O: the stored message wakes the engine
  O->>B: PUT /integration/messages/{id} (same envelope)
  B-->>O: 2xx with completed, rejected or duplicate
  O->>O: record the Acknowledgement and mark the route Delivered
  1. The sending app writes the event in the same transaction as the change (its outbox), so a change and its event commit — or roll back — together. Nothing is lost if the orchestrator is down: the event waits in the outbox.
  2. The app's publisher sends it at once. A commit that wrote outbox rows wakes the publisher; it sends in up to four lanes in parallel, keeping the order of each business flow (its correlation). A publisher pass every 2 s remains as the fallback.
  3. The orchestrator stores and routes it in one short transaction and answers accepted (or duplicate for an id it already has).
  4. The stored message wakes the engine, which delivers it to each destination. An engine tick every few seconds remains as the fallback (retries, schedules).
  5. The receiving app processes it once: its inbox recognises a message (or a command's idempotency key) it has already handled and answers duplicate without doing the work again.

A hop normally takes well under a second (measured on the work_supply network: median 0.06 s from the sending app's commit to delivery, 0.4 s at the 95th percentile).

Who receives it

  • an explicit destination in the envelope, if there is one;
  • a command or query without one goes to the owner of its entity type (Entity Ownership);
  • an event goes to every active Subscription whose pattern matches its type — never back to the sender.

Subscriptions, owners, member applications (with their addresses) and the retry policy are loaded from the network's manifest (networks/<name>/system.manifest.yaml, by load_network.py) when the network starts — never typed in by hand.

Queries for reading only. An app that only needs to read another app's data — a picker searching the catalogue, SAP AP checking an order's receipts — asks the owner directly (POST <owner>/integration/query, same service token) instead of sending a message through the orchestrator. Queries never change data and are not routed.

When delivery fails

What happened What the orchestrator does
no answer (network error), 408, 425, 429 or 5xx retry on the schedule of retry policy MSG-V1-DEFAULT (2, 10, 30, 120, 600 s)
any other 4xx (e.g. handler not found), or the app has no address dead-letter at once
retries used up dead-letter — but only once the route has also been failing for the schedule's whole window, so a briefly flickering app does not use up its attempts in seconds

Two rules keep an outage from spreading:

  • An app that does not answer is tried once per engine tick. Its other waiting messages wait for the next tick without spending an attempt, so one stopped app does not slow the delivery of everyone else's messages.
  • When an app answers again, its waiting retries go at once instead of sitting out the rest of their back-off. Measured: with Inventory stopped for a minute, the backlog cleared in about 20 s after it came back.

A dead letter can be put back in the queue with the same envelope — the engine's retry_dead_letter (there is no button for it on the screens yet) — and when it is then delivered it is marked Resolved. Re-sending is always safe because the receiving app answers duplicate for a message it has already processed.

Reconciliation (at most every 30 s) records open problems once each — deliveries without an acknowledgement, routes still undelivered after an hour, open dead letters — and closes the record when the route is delivered and acknowledged.

Where to look

In this app (menus Message Protocol and Integration):

Screen Shows Views
Event Messages every message; protocol tells the two kinds apart default · Troubleshooting (ids, correlation, errors) · All fields
Route Decisions one row per destination: status, attempts, next attempt default · Troubleshooting
Delivery Attempts every delivery: endpoint, HTTP status, duration, error, next retry default · Troubleshooting
Acknowledgements the receiving app's answer per route
Dead Letter Items routes that gave up, with the reason
Subscriptions who receives which message types (from the manifest)
Message Reconciliation open delivery problems

Status columns are colour badges: green delivered / completed, amber routing / retrying, red dead letter / failed, blue selected. To follow one business flow, filter Event Messages by its correlation.

In the control panel (python server/control_panel.py, http://localhost:5000/networks): the network console shows every member's status, live message counts, and a message-flow diagram (one column per app, one arrow per delivered message) — with scenario buttons and simulated operating days to watch the protocol work.

Endpoints and security

Endpoint Used by Protection
PUT /_orchestrator/messages/<id> the apps' publishers service token (ORCHESTRATOR_SERVICE_TOKEN); not behind the login screen
PUT <app>/integration/messages/<id> the orchestrator's deliveries the same token
POST <app>/integration/query apps reading another app's data the same token

The engine runs only in a process started with ORCHESTRATOR_ENGINE=1, and its message step only in the process that serves requests (one engine per orchestrator instance).

Running a network locally

python networks/work_supply/network.py start      # seeds missing databases, loads the manifest, starts all 11
python networks/work_supply/network.py status
python networks/work_supply/network.py stop
python networks/work_supply/network.py reset [--world S|M|L] [--start]

or Load network in the control panel's Networks view — both use the same launcher and the same process manager. Operating guide: networks/work_supply/RUNBOOK.md; the network's members and flows: networks/work_supply/NETWORK_MAP.md.

The live orchestrator (orchestrator.dslcore.net) runs the same code with the exploration suite's entity-sync data; running a message-protocol network there is part of deployment (R14).