Technical whitepaper v1.0 · what supply-chain AI has to stand on
Supply-chain AI is not failing at the model. It is failing at everything the model has to stand on: consistent meaning, real relationships, situational context, and a record of what was already decided. This paper sets out the layers we put underneath the answer, where the boundary between computation and language sits, and why we draw it where we do.
Supply chain & data leaders
1.0
August 2026
25 min read
Over the past several weeks, supply-chain practitioners: heads of global supply chain, procurement managers, demand-planning leaders, technology founders, have been describing, independently and in public, why AI is not landing in their operations. They are not describing the same project. They are describing the same shape.
None of them said the model was the problem. That is the finding worth sitting with, because almost all vendor attention, and almost all buyer scrutiny, is pointed at the model. This paper takes those reports at face value, works out what they have in common, and then shows what we built underneath our own AI layer in response, including the parts that are not finished.
The practitioner material is qualitative. It establishes that these themes recur across roles: it does not establish how common they are, and no claim of prevalence is drawn from it.
Every figure quoted in Section 08 is raw output from a running instance, captured in one pass and committed with its provenance at docs/whitepaper/data/engine-capture-2026-08-04.json. Running node scripts/verify-ai-whitepaper-figures.mjs re-derives all forty-one of them from that capture and fails on any mismatch: a figure it cannot derive does not belong in this paper. The underlying records are seeded demonstration data, so the numbers are genuine outputs of the pipeline and not facts about a real supplier.
Two of the layers described here are specified and not built, and one dependency is disqualifying for classified use as things stand. Section 11 lists all six gaps rather than leaving them to be found.
A companion paper, “From supply-chain data to defensible decisions,” argues the same case for an operational audience over a fully synthetic dataset. This one is the architecture and the running system: that one is the argument and the adoption case.
The observations below are drawn from public posts by named supply-chain professionals during July and August 2026. We have deliberately not attributed or quoted them. These are people describing their own working conditions, not endorsing a product, and it would be a misuse of their words to place them beside ours. What follows is our synthesis, and any error in it is ours.
Asked what was actually stopping AI adoption, 140 supply-chain chiefs did not name model accuracy. Fifty-six percent named integration with the legacy systems they already run. That figure reaches us through a practitioner's summary of research we have not read in the original: we cite it as reported rather than as established. The advice that followed was unusually candid for the industry: evolve the stack rather than replace it, and build the intelligence as a layer on top of what is already there. The WMS stays. The ERP stays. The spreadsheet everyone quietly depends on stays.
Practitioners in procurement and planning describe the same division of their day: the decision itself is a small fraction of it. The bulk is collecting data, validating it, reconciling one system's version against another's, and shaping the result into something a person can act on. The opportunity, in their reading, is not a system that decides. It is a system that removes the preparation so that expert judgement has room to operate.
A procurement manager put it more concretely: the job is quotes buried in email threads, a lead-time slip mentioned in the fourth paragraph of a reply-all, three open RFQs and no clean answer about who owes what. Pricing sits in forty PDFs. Deviation from standard terms sits on page nineteen of a master agreement. This is not strategy work, and it is where the hours go.
Demand planning, one practitioner noted, is measured by the absence of problems. When it is done well, nobody notices. When it is not, everybody asks why inventory is high, why customers are waiting, why shipments are being expedited. A function whose success is invisible is a function whose value is chronically underestimated, and, we would add, a function whose reasoning is never written down anywhere.
A technology founder working in aviation MRO drew a distinction that most stacks blur. A semantic layer defines your metrics once, so a dashboard and an agent return the same number for the same question. An ontology defines your entities, so that the same physical engine referenced in three different systems is understood to be one engine. A context layer wraps both and answers three questions: who is allowed to see this, where did this figure come from, and what did we already decide.
His test is the sharpest thing we read all month. A semantic layer with no ontology underneath it gives you a confident number for the wrong entity. And an ontology relabelled “context layer,” with no lineage or access policy attached, is a better diagram, not a better system.
A related argument concerns the layer that gets skipped entirely. An audit log records that a value changed and who changed it. It does not record why, or what else the decision was assumed to hold true for. That second record, decision memory, is absent from most systems, and the reasons given for its absence are not technical. Nobody asks for it in an RFP. It has no dashboard, because what it produces is the absence of a future problem, which demonstrates badly. And its cost is deferred: a system without it works fine for a year.
Finally, and least tractable: the people who can interrogate an AI recommendation, override it when context demands, and design the workflows that make it useful are not the people the industry hired for. That skill set cannot be bought with a software licence, and it did not exist as domain knowledge two years ago. The sharper form of the argument asks whether an organisation built around managing uncertainty that AI now processes in seconds can be incrementally converted at all, or whether AI-native and AI-enabled are different species.
Line those seven up and the temptation is to file them as a list of separate complaints. They are not separate. They form one system, and the order matters: each band makes the one below it worse.
Fragmented sources and constraining legacy systems produce a day dominated by preparation. Terms that mean different things and entities that do not match produce analysis that cannot be trusted: the confident number computed over the wrong part. Absent context and an unrecorded rationale produce decisions that cannot be defended or reused.
The return path is the part worth dwelling on. A decision that cannot be defended gets made again, and making it again is preparation work. So the third band feeds the first, and an organisation in this state does not drift towards equilibrium, it accumulates. That is a reasonable explanation for why the reported experience is of things getting harder rather than plateauing.
Six of the seven sit below the intelligence. One, the capability gap, is not a band at all: it runs across all three, because the skill to interrogate an answer is what converts any of this into a better decision. None of the seven is a question about the model.
This matters commercially, because it inverts the usual buying conversation. If the binding constraint were model quality, the right move would be to wait for a better model, and the right question to a vendor would be about benchmarks. If the binding constraint is the substrate, a better model changes nothing: it produces more fluent answers over the same unreliable ground.
Our position is that the substrate is the product, and the model is a rendering surface attached to the end of it. The rest of this paper describes that substrate concretely, using our own system and its actual data.
Before meaning, relationships or context, there has to be a curated dataset that is worth reasoning over. Ours is deliberately small and deliberately closed: six datasets covering the value chain end to end, and no seventh.
| DATASET | CARRIES | FEEDS |
|---|---|---|
| Product design | Part identity, specification, origin declarations | Origin-threshold evaluation |
| Planning | Committed dates, schedule position | Schedule-slip detection |
| Procurement | Supplier, contract, single-source flags | Single-source exposure |
| Manufacturing | Lots, yields, defect rates | Component-quality detection |
| Logistics | Shipments, lanes, days late, holds | Logistics-delay detection |
| Finance | Contract values, remediation cost | Financial-exposure evaluation |
Six is a constraint, not a milestone. A new reporting requirement is answered with a column on an existing dataset or a composition across several, never a seventh table. The moment the curated layer starts growing per-question, it stops being a shared vocabulary and becomes the same sprawl it replaced.
Every row carries lineage back to the raw asset it came from, and every derived figure carries an evidence state: verified, attested, estimated or missing. That last distinction does more work than any model we could attach. A number that is estimated and a number that is verified are different objects, and a system that renders them identically is lying to its operator regardless of how accurate the underlying estimate is.
The distinction between semantic layer, ontology and context layer is the most useful thing we have adopted from outside our own work this year. Our AI layer is being restructured around it explicitly, because each omission fails in its own recognisable way.
One authoritative definition of each term the supply chain is described in. What counts as on-time delivery. What tier means. What makes an item controlled. Varied source fields map onto one vocabulary, and the rule is that a term is defined in exactly one place: the ontology and the context layer reference it and never redefine it. Without this, the same question asked twice returns two numbers, both defensible, and trust collapses on the first discrepancy a user finds.
Typed relationships and the rules for traversing them. Supplier supplies component; component belongs to lot; lot feeds sublot; sublot is assembled into a programme deliverable; certificate covers lot; supplier is located in jurisdiction. Without it, a system can report that a licence lapsed but cannot tell you what that lapse touches. With it, the traversal itself is reportable: not just the conclusion, but the path taken to reach it.
Who is asking and what their role permits, which programme is in view, what deadline is live, what the operating rules are. This is the layer that turns a finding into the right recommendation for this person now. It answers the three questions above: who may see this, where did the figure come from, and what did we already decide.
The third question is the one that has no home in most architectures, which is the subject of Section 07.
Everything so far describes what the model stands on. This section describes what it is not permitted to do, which we consider the more important half.
Detection and categorisation are deterministic. A rule catalogue, eleven rules in the running system, evaluates facts from the curated datasets against stored thresholds and produces a categorised risk case. No model participates. The rules are versioned and editable by an administrator without a deploy, so tuning is an operational act with a history, not a code change.
| RULE | METRIC | AMBER AT | ESCALATES AFTER |
|---|---|---|---|
| Origin below threshold | eu_origin_pct | < 65% | 10 days |
| Logistics delay | days_late | > 3 days | 7 days |
| Component quality issue | defect_rate_pct | > 2% | 10 days |
| Certificate expiry | days_to_expiry | < 30 days | 14 days |
| Schedule slip | days_late | > 15 days | 14 days |
| Sanctions screening hit | n/a | on match | 1 day |
The model writes grammar, not values. When a risk case is rendered into a sentence for an operator, the model never sees the figures. It receives a slot manifest, keys and human labels only, and returns prose containing {{placeholders}}. The server substitutes the real, deterministically formatted values afterwards. Because the model never handles a number, a name or a date, it is structurally incapable of altering one. This is not a policy enforced by review: it is enforced by what the model is given.
A validator rejects any output containing a bare value, a real entity name or an unknown placeholder, and falls back to deterministic prose. In the system state used for this paper, with no model configured at all, that fallback fired seven times and the response payload was unchanged. The AI layer is genuinely optional: turn it off and you lose fluency, not function.
One surface admits no AI at all. The quantitative view, tables, percentages, counts, is rendered directly from facts with no narration, summarisation or model-derived figure, permanently. One fact bundle, two renderers: one produces sentences, one produces numbers, and only the first involves a model.
Sections 03 to 05 describe the layers as a stack. This section describes them as a sequence, because that is how a transmission actually experiences the system. There are eleven stages. The model is the eighth, and it is the only one that cannot change a number.
The first four stages ingest raw records from the six source systems, resolve them against the ontology so that the same supplier, component or lot referenced in different systems collapses to one entity, attach lineage back to the raw asset, and mark each derived figure with its evidence state. Nothing downstream is allowed to treat an unresolved or unlabelled record as fact.
The rule catalogue evaluates the curated facts against stored thresholds, produces a categorised risk case, and attaches the context needed to make that case actionable: who is allowed to see it, what programme it belongs to, and what was already decided about similar cases. None of this involves a model.
The eighth stage is the only one where a model is involved, and it is given a slot manifest rather than figures. It writes the sentence around the case. It cannot alter a number, a name or a date, because it never receives one.
A validator checks the model's output for bare values, real entity names or unknown placeholders and falls back to deterministic prose on any failure. The response is delivered to the operator through one of two renderers. The interaction, and the decision made from it, is written back into the record that the next pass of the pipeline will read.
It would be easy to describe stage 11 as the system learning from operator feedback. It does not. What is written back is a record of what was decided and why, available for the next case to reference, not a weight update to the model. The distinction matters: the pipeline gets more useful over time because its memory grows, not because its model changes underneath it.
Prefer to see the layers running instead of reading about them? Thirty minutes, on screen, nothing to prepare.
An audit log answers: what changed, when, and by whom. That is necessary, and we have it on every action. It is not sufficient, because it cannot answer the question that actually arrives eighteen months later, usually from an auditor or a new owner: why did we accept that?
A decision memory record captures the judgement applied around a deterministic outcome: what was decided, by whom, on what evidence and with what evidence state, which alternatives were weighed, and what happened as a result. Decisions supersede rather than overwrite, so the earlier judgement and its replacement are both retained and linked.
Where AI is involved, the record states that AI advised and a named human decided. There is no path by which a decision is attributed to the system. This is partly a regulatory position and partly a practical one: an organisation that cannot name the person behind a judgement cannot defend it, and will not learn from it either.
We should be plain that this layer is specified and not yet built. It is the clearest illustration of the deferred-cost problem described in Section 01: a system without it works perfectly well for a year, and then does not.
What follows is a single thread taken end to end from a running instance. The figures are real outputs of the system described above. The underlying records are seeded demonstration data, not a customer's operation: the numbers are genuine products of the pipeline, not genuine facts about a real supplier.
The instance holds twelve programmes. Three are at risk, down from five over four weeks. Two sit below the EU-origin threshold of 65%, drawing on five connected sources. The nearest audit is an EDIP eligibility review of a programme on 16 August 2026 against control C-EUO-001.
Banding. For a metric where lower is worse, the band is green at or above the green threshold, amber between the amber and green thresholds, and red below the amber threshold. For this rule the stored values are 65 and 60, so a measurement of 58 bands red. The verification script re-derives the band from the thresholds rather than reading the stored band, so a drift between configuration and outcome would fail the check.
What 58% is, and what it is not. It is EU-origin content for one component, expressed as a percentage and carrying an evidence state of estimated. The derivation of that percentage happens in the origin service and is not re-derived by the verification script: the capture records the measurement, not the arithmetic behind it. We would rather state that limit than imply a depth of verification we have not done.
Two exposure figures, two denominators. €1.85M is the exposure the engine attaches to this case across the two programmes on it. €4.6M is the Command Centre roll-up across all three programmes the shared component touches. They are not inconsistent, they are different scopes, and quoting either without the other would misstate the position.
EU-origin content for one silicon die, measured across the curated product-design dataset, came to 58%. The stored threshold is 65%, with a red band below 60%. Nothing about that evaluation is a judgement call: it is a comparison between a computed figure and a number an administrator set, and it would produce the same result on every run.
The comparison produced a risk case categorised ORIGIN_BELOW_THRESHOLD, banded red, carrying exposure of €1.85M and touching two programmes. It was raised on 29 July and de-duplicated against a key combining the rule, the supplier and the component, so a second detection of the same condition updates the case rather than creating a parallel one. Ten days later, unclosed, it escalated: the case is at level 2 with the next escalation due 8 August.
In the Command Centre the same facts surface as a decision with a recommended next action: request origin confirmation from the supplier, and a stated reason for that action: one confirmation re-derives EU-origin across every shared lot. The decision carries the wider blast radius, €4.6M across three programmes, and an evidence state of estimated, because origin has been declared rather than confirmed against a source record.
That last field is the one we would point to. The system is not claiming to know the origin percentage. It is claiming to have computed 58% from evidence it has marked as estimated, and it is recommending the action that would convert that estimate into something verified. A layer that could not represent the difference would have presented 58% as fact.
Nothing that matters. Auto-executable remediation steps run inside a window the operator can redirect, every one of them audited. The case does not close until a person confirms it. If nobody acts, the clock escalates the case upward through the buyers above rather than quietly letting it lapse: which is the behaviour the “measured by the absence of problems” observation in Section 01 demands, since the failure mode of an invisible function is silence.
Everything above describes a destination. This section describes a route, because the most common failure we see is not choosing the wrong architecture, it is starting at the layer that demonstrates well rather than the layer everything else depends on.
The exit conditions are the substance here, not the increment names. “Ground the data” is a slogan; “the same records produce the same metric and the same denominator” is a test that either passes or does not. An increment without a falsifiable exit condition cannot be finished, only abandoned, and a programme made of such increments reports progress indefinitely.
Note where the model appears: increment five of six. Increments one to four, measuring the preparation cost of a real decision, resolving identity, typing the relationships, moving thresholds into configuration, deliver value on their own and are not blocked on choosing a model, a vendor or a hosting posture. If budget runs out after increment four, the organisation is better off than it started. That is deliberate sequencing, not modesty.
The capability gap in Section 01 is an ownership problem before it is a training problem. Four roles have to exist, and in most organisations three of them are unassigned:
Software can make the interrogation skill cheaper to acquire by showing the traversal, the evidence state and the prior decision rather than only the conclusion. It cannot assign these four ownerships, and a deployment into an organisation where they are unassigned will produce fluent answers that nobody is accountable for.
Most AI measurement we encounter reports on the model: acceptance rate, satisfaction, questions asked. Those describe how the interface is being received. They say nothing about whether the decisions got better, and they will look healthy in a system that is confidently wrong.
Measure the operating system instead. These are the indicators we would hold ourselves to, grouped by the layer they test:
| LAYER | INDICATOR | WHAT A BAD READING MEANS |
|---|---|---|
| Data | Share of displayed figures with inspectable lineage | Numbers are being trusted rather than checked |
| Data | Unresolved-entity rate; identity-match precision | Arithmetic is landing on the wrong part or supplier |
| Semantic | Definition consistency across surfaces | Two surfaces will disagree, and trust will go on the first discrepancy |
| Context | Permission-leakage incidents | Someone saw something their role does not permit |
| Model | Unsupported-claim rate under adversarial evaluation | The boundary is policy rather than mechanism |
| Decision | Rationale completeness; supersession frequency | The why is not being captured, or is never revisited |
| Decision | Decision-to-outcome linkage; repeat-exception recurrence | The same judgement is being made again from scratch |
| Operation | Median preparation minutes per decision | The preparation tax has not moved |
Two of these deserve emphasis because they are the ones that fail quietly. Repeat-exception recurrence is the direct measure of the loop in Figure 2: if the same exception keeps arriving, decision memory is not doing its job regardless of how complete the records look. And unsupported-claim rate under adversarial evaluation is the only honest test of the boundary in Section 05: we assert that placeholder narration makes figure tampering structurally impossible, and an assertion of that kind should be attacked deliberately rather than accepted because the design implies it. We have not yet published that evaluation, which belongs in the list in Section 11 as much as anywhere.
A whitepaper that ends at the previous section would be dishonest about the state of the work. Six things are unresolved, and a reader evaluating us should weigh them.
The architecture in Figure 3 places the model inside the trust boundary. Today the call goes to a hosted third-country API. For a customer operating at classification, that is not a tuning issue: it is disqualifying, and it must be replaced by an accredited model hosted inside the enclave before any such deployment. It is tracked as a blocker, not as a backlog item.
The decision memory layer in Section 07 and the three-layer restructuring in Section 04 both exist as requirements documents. The AI layer as it runs today is closer to intent-keyed assistance than to the fundamental analysis those documents describe. We have published the specifications and the reasoning; we have not yet published the implementation.
Text carried inside a partner document should never reach the model as an instruction, and a release crossing an organisational boundary should be checked against the receiving party's clearance and cut to the fields they are entitled to see. Both are designed and neither is implemented. The placeholder mechanism in Section 05 constrains what the model can emit; it does not by itself constrain what reaches it.
Section 08 is a real execution of a real pipeline over demonstration data. It demonstrates that the mechanism works and that the numbers are computed where we say they are. It does not demonstrate accuracy against any real supply chain, and we would resist anyone citing 58% or €1.85M as evidence about a supplier.
Building intelligence as a layer over systems that stay in place is the right shape, and it is the shape we chose. It is still integration work. A first connection to an ERP or a PLM is measured in weeks, not hours, and anyone quoting hours is describing a demo. What the framework buys is that the second and fifteenth connections cost a fraction of the first, because identity, staging, validation, retry and audit are already solved.
The seventh problem in Section 01 is the one no vendor can close. People who can interrogate a recommendation, override it on context, and design the workflows around it are built inside an organisation over years. Software can make that cheaper by being interrogable: by showing the traversal, the evidence state and the prior decision rather than only the conclusion, and that is a real design objective for us. But a licence does not produce the skill, and a paper that implied otherwise would be selling the thing the practitioners specifically warned about.
The practitioners quoted in aggregate at the start of this paper are describing a market that has spent several years optimising the visible layer and skipping the ones underneath. The costs of that are deferred, which is exactly why it keeps happening: a system with no consistent semantics, no traversable relationships, no access-aware context and no memory of its own reasoning works acceptably for about a year.
Our bet is that the durable product is the substrate: the curated datasets, the deterministic rules over them, the lineage and evidence state attached to every figure, and the record of what was decided and why. The model sits at the end of that chain and writes the sentence. It is genuinely useful there, and it is replaceable there, which is the point. If the model can be swapped, or switched off entirely, without changing a single number the operator sees, then the value was never in the model.
That is a testable claim rather than a marketing one, and Section 05 records the test: with no model configured, the system fell back seven times and returned the same payload.
Practitioners describing stalled AI adoption name integration with legacy systems, the preparation tax on every decision, inconsistent definitions, unmatched entities, missing access-aware context and unrecorded rationale. Six of the seven reported problems sit below the intelligence layer and none of them is a question about model quality, so a better model produces more fluent answers over the same unreliable ground.
A semantic layer defines each term once, so a dashboard and an agent return the same number for the same question. An ontology defines the entities and typed relationships, so the same physical part referenced in three systems is understood to be one part. A context layer wraps both and answers who is allowed to see this, where the figure came from, and what was already decided. A semantic layer with no ontology under it returns a confident number for the wrong entity.
Detection, categorisation, banding, exposure and the remediation plan are deterministic: a versioned rule catalogue evaluates curated facts against stored thresholds and no model participates. The model receives a slot manifest of keys and human labels with no values, and returns prose containing placeholders that the server substitutes with deterministically formatted figures. Because the model never handles a number, a name or a date, it cannot alter one.
You lose fluency, not function. A validator rejects any model output containing a bare value, a real entity name or an unknown placeholder and falls back to deterministic prose. In the state used for this paper, with no model configured, the fallback fired seven times and the response payload was unchanged. The quantitative view is rendered directly from facts and admits no model at all.
An audit log records what changed, when and by whom. A decision memory record captures the judgement around a deterministic outcome: what was decided, by whom, on what evidence and with what evidence state, which alternatives were weighed and what happened as a result. Decisions supersede rather than overwrite. Nobody asks for it in an RFP and it has no dashboard, so it gets skipped, and a system without it works acceptably until the first challenge to the reasoning arrives.
TECHNICAL WHITEPAPER V1.0 · PUBLISHED AUGUST 2026
The practitioner material behind Section 01 is qualitative and drawn from public posts during July and August 2026. It is deliberately unattributed and unquoted, and no claim of prevalence is drawn from it. Every figure in Section 08 is raw output from a running instance over seeded demonstration records, so the numbers are genuine products of the pipeline and not facts about a real supplier. Two of the layers described here are specified and not built, and Section 11 lists all six open gaps rather than leaving them to be found.
A companion paper, From supply-chain data to defensible decisions, argues the same case for an operational audience over a fully synthetic dataset. This one is the architecture and the running system; that one is the argument and the adoption case. For the regulatory side of the same problem, see the EDIP and the 35% rule guide. For the integration layer that feeds this pipeline, see the Common Interface Framework paper.
Thirty minutes: how evidence, lineage and thresholds work in your company today, Skansar on screen against data shaped like yours, then closing remarks and next steps. Nothing to prepare and no data needed for the call.