AI doesn’t need more documents, it needs knowledge it can trust.

 

Ask a language model about a Phase III study and it will answer. Fluently, instantly, convincingly, and with no inherent way of knowing whether the answer is valid.

That gap between fluent and valid is the central problem with AI in regulated life science.

It is tempting to treat it as a model problem. It mostly isn’t. You cannot fine-tune your way out of a knowledge base that has no opinion about what is true, and you cannot prompt-engineer your way around five systems that disagree about the current state of the same study. Retrieving the most semantically similar paragraph tells you nothing about whether that paragraph is current, superseded, applicable, authoritative, permitted, or contradicted by something that happened yesterday.

 What AI needs is knowledge deliberately made fit for machine consumption: knowledge carrying its own identity, context, boundaries, history and provenance. Knowledge that knows what is true now, what must be true, what happened to make it true, and which part of that truth belongs to the person asking.

 That is the problem Miosis is built to address.

 

From documents to governed knowledge

In The Biology of Knowledge I argued that four billion years of evolution already solved knowledge management, and that the answer is a living cell. A cell does not expose every molecule to every process. It maintains identity, responds to context, operates inside boundaries, records change through state transitions, and expresses the right information in the right place for the function at hand. The Cellular Architecture set out the technical layers that follow from that.

 Both were arguments; this article comes with an artifact: 

www.miosis.io is live, and it carries the first clinical study holon.


AXION-3

AXION-3 is a synthetic randomized, double-blind, double-dummy Phase III trial of MRX-302 (voxacerib, an oral RAS multi-selective inhibitor) against chemotherapy in second-line metastatic pancreatic ductal adenocarcinoma. It carries dual primary endpoints, overall survival and progression-free survival, assessed in the RAS-mutant primary population by blinded independent central review.

500 patients randomized from 560 screened. 58 sites, 6 countries, two protocol amendments, three DSMB reviews, and roughly 16,000 events in its history.

Every bit of it is synthetic. The sponsor is fictional, the compound is fictional, the patients and sites are generated. No real patient or clinical site exists anywhere inside it.

That is the point rather than a caveat. A synthetic study can be opened, inspected, queried, challenged and argued with in public. Try that with a real Phase III.

The word for the resulting object is holon: one governed knowledge object holding a single identity across its whole lifecycle, composed of four coordinated graphs.

Scene graph: what is true now

The Scene is the study’s current governed state. Not the version somebody exported yesterday, not the spreadsheet sitting in an inbox, not whichever document happened to rank highest in retrieval. The current state.

 For AI that distinction does real work, because a model cannot reason reliably if it must first guess which representation of reality is authoritative.

 

Boundary graph: what must be true

The Boundary holds the rules governing the object: ICH GCP requirements, inclusion and exclusion criteria, visit-window rules, endpoint-derivation rules, controlled terminology, authority-specific submission constraints, and the other validity conditions that decide which states are acceptable.

 Storing those rules is the easy half. The Boundary makes them **machine-checkable**. Validity stops living only as prose inside a protocol, an SOP or a guidance document, and the knowledge object itself starts carrying executable constraints about what is permitted, required or inconsistent.
 

Event graph: what happened, when, by whom and why

The Event graph records the lifecycle. Every meaningful change becomes a governed event rather than a historical footnote: what changed, when, who or what caused it, under what authority, what the previous state was, and what evidence supports the new one.

AXION-3 holds roughly 16,000 such ledger entries. That history is exactly what you need the moment an AI is expected to defend an answer rather than merely produce one.
 

Projection graph: what is relevant here

The Projection graph computes context-specific views of the same object. A medical monitor does not need the same representation of a study as an investigator; a statistician does not need the submission lead’s view. An AI agent assessing protocol deviations should not receive the same unrestricted context as one preparing an operational site summary.

AXION-3 carries seven role views plus explainable predictions, all computed from the same underlying knowledge rather than re-entered as separate versions of reality.

 

Why retrieval over documents is not enough

Retrieval over documents gives a model text about a study. It does not give it the study.

Ask a document-grounded assistant whether a patient visit fell inside the permitted protocol window and it will return several highly relevant passages. Meanwhile the actual visit sits in the EDC, the permitted window sits in a protocol PDF, the applicable protocol version depends on the site’s amendment implementation date, a special exception may live in a third system, and the definition of the visit itself may depend on terminology stored somewhere else again.

Every one of those facts can be individually correct. Nothing in a conventional document repository knows they belong together.

Similarity is not semantics. Retrieval is not validity. And a fact does not become knowledge just because it fits inside a context window.

 A holon changes four things.

1. It grounds

One identity. One current state. One semantic point of reference across the study lifecycle.


That matters because a great many apparent “AI accuracy” failures begin before the model generates a single token. They begin when two systems disagree about which entity is under discussion, which version applies, what a field means, or what state the object is in. Model intelligence cannot repair that kind of ambiguity. It can only render it more fluently.

So the first requirement for accurate AI is a better definition of what the answer is allowed to stand on. 

 

2. It refuses

This is the capability I think matters most, and the one enterprise AI discussion reaches for least. A trustworthy system has to be able to reject an answer.

Because the Boundary carries validity as executable constraints rather than prose, an assertion that violates a known rule can be challenged or failed before a human relies on it. That shifts the nature of the problem. Hallucination is normally treated as a statistical property of a language model, something to be reduced through better prompting, retrieval, tuning or model choice. Where the domain contains explicit rules, some of those failures stop being probabilistic and become gates:

Does this answer violate an inclusion criterion?
Does this date fall outside the protocol-defined visit window?
Does this endpoint derivation contradict the approved rule?
Does this assertion rely on a superseded state?

Does this conclusion cross an authorization boundary?

 

A substrate that can evaluate those questions gives an AI system something worth more than fluency: the ability to say no.

A system that cannot say no cannot be trusted when it says yes.

3. It proves

Enterprise AI needs more than citations stapled to generated prose. It needs lineage: which fact supports this assertion, which event changed that fact, which rule applies and in which version, which source supplied the evidence, and under whose authority it entered the canonical state.

The Event graph gives every assertion a history. The output stops being *”the model said so”* and becomes a chain you can inspect.

In regulated environments that is close to everything. Provenance reconstructed three weeks before an inspection is not provenance. It should be a property of the architecture.
 

4. It answers the right person

There is rarely one universally correct representation of a complex regulated object. There are correct views for particular purposes. Investigator and medical monitor need different views of the same study. Regulator and sponsor may legitimately see different projections. A pharmacovigilance agent and a site-management agent should operate over different allowed contexts.

The Projection graph makes that explicit: the truth is governed once, and the view is computed from role, task, authorization and decision. Which is a rather different proposition from maintaining several independently curated spreadsheets and hoping they stay in agreement.

 

Fit-for-purpose knowledge for AI

An enterprise knowledge base does not become “AI-ready” the moment its documents have been embedded. Knowledge that is genuinely fit for AI consumption needs: 

  • stable identity
  • explicit semantics
  • controlled terminology
  • relationships between entities
  • current state
  • applicability
  • rules and constraints
  • lifecycle history
  • provenance
  • versioning
  • authorization
  • uncertainty and exceptions
  • validation
  • and purpose-specific views
It also needs something more basic than any of those: a definition of what questions the knowledge is expected to answer.


Miosis calls these competency questions. AXION-3 carries
500 competency questions across 17 domains, spanning three verbs: describe (what is true), prescribe or govern (what should or must be true), and predict (what is likely to happen). All 500 are answered from the same four graphs.

Competency questions make an unusually honest benchmark for knowledge engineering. You write the question first, then build the knowledge needed to answer it. When the graph stops answering a question correctly, the build has failed. Continuous integration for meaning.

Anyone who has tried to reconcile a study’s state across EDC, CTMS, the TMF and a submission dataset knows the failure mode: several systems, several answers, each locally defensible, none reconcilable without a person in the room. Writing the questions first is what exposes where that happens. It is the least glamorous part of the work and the part that changes the most.

 

The layer underneath the model

We discuss AI accuracy as though it were a property of the model. Choose a better model, add retrieval, widen the context window, tune the prompt, drop the temperature, fine-tune, add an agent. All of that can matter.

Underneath it sits another layer: what does the system know, how does it know it, which version does it know, what changed, what is authoritative, what may it infer, what must it reject, and can the whole path from evidence to answer be reconstructed? Those are knowledge-architecture questions, and a more capable model running over badly governed knowledge just produces a more convincing wrong answer.

Which reframes the objective. A conventional AI system asks whether it can generate a plausible answer from what it retrieved. A system with a governed substrate can ask whether the answer is supported by the current state, consistent with the applicable rules, traceable to its evidence, and valid for the context in which it was asked. That is a considerably higher bar, and it is the whole distance between a plausible answer and a defensible one.

The aim is an appropriately informed AI rather than a maximally informed one: the right knowledge object, the right projection, the relevant constraints, the provenance preserved, the whole thing tested against explicit questions. Then let the model do what language models are extraordinarily good at, which is to reason, synthesize and explain over that substrate.

 

Where this gets hard

Three objections I have not answered well.

The first is cost: Someone has to author the Boundary rules, keep them current through two protocol amendments, and decide what happens when two rules conflict. That is expert human work and it does not disappear because the output is machine-checkable. A holon is cheaper to query than a document estate and considerably more expensive to build.

The second is scope: This architecture suits objects with a defined lifecycle, a regulatory envelope and a stable identity: a study, a product, a submission. It suits exploratory discovery science much less well, where the point is precisely that the entities are not yet stable and the rules are what you are trying to find. I would not model a target-discovery programme this way.

The third is the artifact itself: AXION-3 is synthetic, which is what makes it publishable and also what makes it easy. Synthetic data is clean by construction. It shows that the architecture holds. It does not show that the architecture survives contact with twenty years of real legacy identifiers, half-finished migrations and undocumented local practice. Nobody has run that test in public, including me.

 

The artifact is live

AXION-3 is walkable. Four graphs, the synthetic clinical dataset, roughly 16,000 events of history, 500 competency questions, the architecture, and a value model that prints the unflattering numbers next to the attractive ones.

The next test is obvious enough that somebody should run it. Take one set of clinical questions. Give one AI system conventional document retrieval and the other a governed holon. Then measure answer correctness, unsupported assertions, use of superseded information, entity and version confusion, constraint violations, provenance completeness, and whether either system can name the fact, the event and the rule behind each answer.

I have not run that comparison. What I am claiming is that this architecture has earned the right to fail it in public.

www.miosis.io

If you build, govern or buy knowledge infrastructure in life science, I would rather have your objections than your likes. Tell me where the pattern breaks.