A message arrives under a colleague's name. Most of it reads like them. But some of the sentences were typed by a person and some were generated by a model, and nothing on the page says which is which. Whatever you do next — rely on it, push back on it, quote it in a meeting — you do it blind.
AI provenance is the repair: a record of origin, attached to the words themselves, that says which parts came from which principal — human or model — and with what authority. Not a vibe, not a disclaimer. A record.
Provenance is a record of origin
The word comes from the art world, where provenance is the documented chain of custody that separates a painting from a forgery. Nobody authenticates a canvas by asking it. They check the record of where it has been.
Applied to AI-assisted writing, provenance answers three questions about any given passage:
- Which principal produced it — the human, or the model working for them?
- Through what process — typed directly, drafted and approved, inferred from context?
- With what authority — is this something the person committed to, or the software's best guess?
That last question is the one attribution schemes forget. "Human or AI" is not a binary; it is a ladder. Words a person typed in an authenticated session carry more authority than words they approved, which carry more than words their assistant merely inferred. Provenance that collapses the ladder into a yes/no flag throws away most of the signal.
Misattribution fails in two directions
Without that record, attribution fails two ways, and both are costly.
A person gets held to the model's words. An assistant drafts something plausible, it goes out under a human's name, and the recipient treats it as a commitment. "Your assistant said you agreed" is a genuinely hard argument to untangle, because from the outside there was never a way to know the human didn't say it.
The person's actual words get discounted. The reverse failure is quieter but just as corrosive. Once a reader suspects everything is machine-written, they discount all of it — including the one sentence the sender wrote deliberately, weighed, and meant. Careful writing costs effort; when attribution is ambiguous, that effort buys nothing.
Both failures share a root: attribution applied to the whole message at once. Which is why the popular fix doesn't work either.
Per-field attribution beats blanket disclosure
The common proposal is disclosure — append "written with AI assistance" to anything a model touched. As a norm it is well-intentioned; as information it is nearly empty. Assistance is involved in a vast share of what people write, so a disclaimer that is true of everything distinguishes nothing. Worse, it flattens exactly the distinction a reader needs: which parts are the person, and which parts are the software.
(The disclosure debate deserves its own treatment — our piece on AI email etiquette makes the longer argument that ownership beats disclosure.)
Per-field attribution goes the other way. Instead of one disclaimer per message, each part carries its own label. The reader stops asking "was AI involved?" — assume it was — and starts asking the useful question: which of these words carry the person's authority?
What makes a label trustworthy
A label is only as good as the process that mints it. Four properties separate a working provenance system from decoration:
- Earned by process, not asserted. The label follows mechanically from how the words entered the system. Nobody — human or model — picks it from a dropdown.
- Checked server-side. Enforcement lives where the sender's own software cannot reach it. A label the sending model could alter is theater.
- A single source of truth. Every surface that shows the words — email, web view, an assistant reading the message — renders the same label from the same record.
- Never the model's to grant. A model that can stamp its own output "human-authored" makes the stamp worthless everywhere it appears.
A worked example: four labels
RelayLink, which carries briefings between people's AI assistants, attaches provenance per field using exactly four labels — a compact ladder of authority:
verbatim, human-authored— typed by the sender in an authenticated session. The only label under which a reader's assistant may quote the words as the person's own. It is earned, not asserted: if the sender echoes the AI's draft back unchanged, that is approval, not authorship, and it is recorded as —AI-drafted, approved unchanged— the assistant wrote it; the human signed off.typed via magic link— a reply typed in the email or web view, authenticated by token. Real words, lower-tier authentication, labelled as such.inferred by sender's AI— the assistant's reading of the private session, not something the human stated.
The line between the first two labels is where most systems would cave. Approving a sentence is meaningful — it is sign-off — but it is not the same act as writing one, and readers deserve to know which happened. The verbatim label exists precisely because that line is worth defending.
One honest limit: provenance records origin, not accuracy. A verbatim, human-authored label tells you the sender typed those words; it does not tell you they are right. And every tier is a claim about authentication — a magic-link reply ranks lower precisely because a token proves less than a signed-in session. Provenance narrows what you must take on faith. It does not eliminate it.
Attribution is the unit of trust
You cannot calibrate trust in a message; messages are mixtures. You can calibrate trust in a labeled part. That is the whole case for provenance: it turns attribution from something readers guess at into something the system records.
Provenance is one of two answers to a much older problem — an agent acting on inferences its principal never sanctioned — laid out in the principal-agent problem applied to AI. To see what a fully labeled message looks like in practice, read what a briefing is — or connect your assistant and send one.