R9 Work & Supply network — local runbook (message protocol v1)

Runs the R1/R2 reference flows across real processes:

Work (5087) --publish--> Orchestrator R9 (5089) --deliver--> Inventory (5088)
     ^                         |    ^                             |
     +------- item_reserved / item_unavailable <-------publish----+
     |                         |    |
     |                         +--> Purchasing (5090)  item_unavailable -> purchase requirement
     +--- requirement_created / delivery_commitment_changed / delivery_delayed ---+

Catalogue (5091) and Supplier (5092) answer item / supplier lookups directly (no orchestrator); SAP AP (5093) reads orders and receipts directly when it matches an invoice. Supplier also receives order_placed, delivery_delayed and receipt_rejected through the orchestrator (performance and supplier issues).

The same flows run automatically (no servers) in server/tests/r9_acceptance/ (pytest server/tests/test_integration_apps.py).

Separate from networks/work_supply_00/ (the older entity-sync demo on the 5086 sandbox).

1. One shared service token

All processes must use the same token. Simplest: put it in the project's .env (gitignored; run.py loads it):

ORCHESTRATOR_SERVICE_TOKEN=r9-local-dev-token

(or set $env:ORCHESTRATOR_SERVICE_TOKEN = "..." / export ORCHESTRATOR_SERVICE_TOKEN=... in the terminal you start things from — started processes inherit it.)

2. Start / stop everything — one command

python networks/work_supply/network.py start      # orchestrator_r9 and every member app (see README)
python networks/work_supply/network.py status
python networks/work_supply/network.py stop       # add --force if a member is still on its port
python networks/work_supply/network.py restart
python networks/work_supply/network.py start --only inventory

start seeds member databases that do not exist yet, loads the manifest into the R9 orchestrator DB (when the orchestrator is down), starts only what is down, in order, and waits for each port. Members are started and stopped through the control panel's ProcessManager, so the panel's Networks view (python server/control_panel.py, port 5000) shows the same processes and can Load / Stop the network too. Output and PIDs: logs/<app>.log / logs/<app>.pid (the panel's files); the manifest load log: networks/work_supply/logs/. The same launcher drives the older network: python networks/run_network.py [--network <name>] [--status|--stop].

Manual alternative (one terminal each)

python networks/work_supply/load_network.py               # once, orchestrator stopped
python server/run.py orchestrator_r9                         # 5089 — engine on, 2 s ticks
python server/run.py work_management                         # 5087 — publisher on
python server/run.py inventory                               # 5088 — publisher on
python server/run.py purchasing                              # 5090 — publisher on
python server/run.py catalogue                               # 5091 — lookups only
python server/run.py supplier                                # 5092 — publisher on
python server/run.py sap_ap                                  # 5093 — publisher on
python server/run.py asset_register                          # 5094 — publisher on
python server/run.py asset_lifecycle                         # 5095 — publisher on
python server/run.py warranty                                # 5096 — publisher on
python server/run.py assurance                               # 5097 — publisher on

load_network.py creates apps/orchestrator/data/orchestrator_r9.db if needed and loads the members (base URLs from applications.yaml), owners, the subscriptions derived from system.manifest.yaml consumes, and the retry policy (idempotent).

The engine / publisher settings come from each entry's env: in applications.yaml (ORCHESTRATOR_ENGINE, INTEGRATION_PUBLISHER, ORCHESTRATOR_URL). Background work starts on each process's first request (so the debug reloader's watcher process never runs a second copy): the trigger below provides it; opening any page does too.

3. Run the flows

python networks/work_supply/trigger.py r2    # oil filter x2  -> reserved
python networks/work_supply/trigger.py r1    # transmission x1 -> unavailable, WO awaiting_parts, PR raised

Each prints the requirement, the reservation and the work order status once the answers have come back through the orchestrator (normally within a few seconds). For R1 it waits until Purchasing's requirement_created has reached Work (requirement status purchasing) and prints the purchase requirement too.

R1c — the goods arrive: trigger.py r1c ships the seeded transmission order PO-004411 (Purchasing), records its arrival and accepts it on inspection (Inventory). The accepted transmission fills the waiting reservation RES-1001, so Work's MR-1001-01 becomes reserved and WO-1001 leaves awaiting_parts (back to planned, see its Status History); Purchasing marks the shipment complete and the order and PR-1001 received. It runs once per demo data set — reset (§5) to run it again.

Purchase to pay (after R1c): trigger.py p2p has SUP-0042 invoice PO-004411 at the ordered price, then SAP AP matches it against the order and the accepted transmission, posts the AP liability (document 51…) and records payment. Look in SAP AP 5093: Invoices, Matching (match lines and any exceptions), AP & Payment (accounting entries). Every order placed also gets a SAP PO number, shown on the purchase order as External ERP Reference.

Warranty: trigger.py warranty claims the failed transmission's eligible candidate (WC-00001: 11 850 h, 22½ months — inside 24 months / 12 000 h), records a 90 % approval, registers SUP-0042's credit memo in SAP AP naming the recovery, and waits until Warranty hears it was posted (claim recovered). Look in Warranty 5096 (Candidates & Claims, Recovery) and SAP AP (Invoices → Supplier Credit Memo). It runs once per demo data set.

R1b — supplier dates (by hand, in Purchasing 5090): open the new purchase requirement, create a Purchase Order for it (or send purchasing.create_purchase_order), add a Delivery Commitment to its line — then a second revision with status at_risk and a reason. Work's material requirement shows the new Expected Available date and the delay in its notes.

4. Where to look

  • Orchestrator 5089 (PIN login — this instance is login-only; machine calls use the token): Message Protocol menu — Subscriptions, Route Decisions, Acknowledgements, Message Reconciliation; Integration — Event Messages (protocol message_v1), Delivery Attempts, Dead Letters.
  • Work 5087: the work order's Material Requirements tab (status, reservation cid) and Status History (awaiting_parts, caused by the inventory event).
  • Inventory 5088: Reservations, Stock Balances, Inventory Transactions.
  • Purchasing 5090: Purchase Requirements (raised from item_unavailable), Purchase Orders, Delivery Commitments (every revision kept), Procurement Exceptions.

Try a fault: stop Inventory (network.py stop --only inventory), run trigger.py r2 --timeout 5 (times out), watch Delivery Attempts record RetryableFailure with a growing schedule (2, 10, 30, 120, 600 s), start Inventory again (network.py start --only inventory) — the next due attempt delivers and the flow completes. After 5 failed attempts the route is dead-lettered.

5. Reset

python networks/work_supply/network.py reset            # stop, rebuild every member from its seed file
python networks/work_supply/network.py reset --start    # ...and start again
python networks/work_supply/network.py reset --only purchasing,sap_ap

The member databases are not in git: each is generated from apps/<app>/schema/<app>_seed.yaml (python -m codegen.cli seed <app>), and start seeds any that is missing. reset also clears the sandbox orchestrator's database (orchestrator_r9.db), which start reloads from the manifest. Use it after a schema change too.

6. Deployments — running another copy

The manifest's deployments: says where this network runs. local (the default) is the registry's ports and databases, its processes managed by the control panel. load is an isolated copy for load testing: every member on its port + 1000 on 127.0.0.1, with its own databases, logs and state under loadgen/instance/, and every peer address pointing inside the copy — it never touches local. Every command takes --deployment:

python networks/work_supply/network.py reset --deployment load --world M   # fresh data for the copy
python networks/work_supply/network.py start --deployment load
python networks/work_supply/network.py status --deployment load
python networks/work_supply/loadgen/drive.py --deployment load --days 3     # simulated days on the copy
python networks/work_supply/loadgen/report.py <run-id>                      # knows which deployment a run drove
python networks/work_supply/loadgen/audit.py --deployment load
python networks/work_supply/network.py stop --deployment load

loadgen/instance.py start|stop|status|reset is the same as --deployment load. In the control panel, choose the deployment next to the network: Load / Stop / Reset, the live activity, the message flow and simulations then work on it (scenario buttons and per-member controls are for local). A new deployment is a few lines in the manifest — it must have its own port_offset and data_dir, so it can never share the local ports or databases.