How to evaluate an AI communication tool

Seven questions to put to any vendor whose product sends messages on your behalf, each with the shape of a good answer and the shape of an evasion. Useful whichever tool you end up choosing.

5 min read

Most evaluations of a tool that sends messages on your behalf turn into a demo, and the demo always works. The interesting behaviour is elsewhere: what happens when the model is wrong, when the input is hostile, and when nobody is watching the screen.

Seven questions discriminate between products here, each with an answer shape that shows the vendor has thought about it and one that shows they have not. RelayLink is one set of answers, not the only defensible set. The list is the deliverable.

Capability without you

Ask: assume the model is compromised or confused. Enumerate what still gets sent.

Every product says a human is in the loop; the phrase means nothing until someone names which actions require the human.

A good answer is a written enumeration: it can read its own inbox and compose drafts, it cannot deliver without a separate confirmation from the account holder, it cannot exceed a per-day limit on strangers. A bad answer is a reassurance — "our system prompt prevents that." Instructions are advice, and hostile input exists to argue models out of advice. If the vendor cannot produce the list, that is the finding.

Scope and duration of access

Ask: what does it access, for how long, and does reading live in the same grant as sending?

The dangerous shape is one credential that both consumes untrusted content and acts on the world, because then an injection is a completed action rather than a bad suggestion.

A good answer separates the directions in the permissions themselves — read here, send there, revoked independently — and is specific about duration. A bad answer describes access in marketing terms, or bundles everything into one consent screen where the only choice is all or nothing. What a connection actually grants is worth reading before you click through one.

Attribution the recipient can see

Ask: can the reader tell which words are mine and which the model produced?

Attribution has to be per-field to be worth anything. A banner reading "sent with AI assistance" says nothing about the sentence that commits you to a date.

A good answer is field-level and assigned by the system from what actually happened: the human's typed words marked distinctly from the model's inferences, the label earned rather than asserted. A bad answer treats attribution as a disclosure footer, or leaves it to the sender to describe honestly. Provenance that depends on the sender's good manners is not provenance.

The burden on the other end

Ask: what does the recipient have to do before they can read this and reply?

A tool that only works between two people who both adopted it is a club, not a communication tool. Most of the people whose answers you need will not join.

A good answer names a delivery floor that requires nothing: an ordinary email address, no account, no install, no assistant on the far side, and a reply path that lands back in the right place. A bad answer is an invitation flow, a link behind a signup wall, or a demo where both parties happen to be customers.

Ask: who can reach me through this, and what happens when I say stop?

Anything that makes sending cheap makes unwanted sending cheap. The question is whether consent is a state the service enforces or a norm the vendor hopes senders observe.

A good answer describes consent as data: an accepted relationship, a hard cap on reaching strangers, a stop that is one click and permanent because the service refuses to carry anything further. A bad answer points at a policy document, an opt-out honoured by the sending side, or a rate limit living in someone's prompt. If the sender's software decides whether to obey, consent is vibes.

What the vendor refuses to claim

Ask: what can this not protect me from?

A vendor who claims no limits is telling you they have not looked. Every design here carries residual risk.

A good answer is specific and unflattering: this does not prevent every injection, only makes it rare and low-yield; this is not end-to-end encrypted, because these components render content; this cannot guarantee delivery beyond what ordinary mail does. A bad answer is the word "impossible", or a security page composed entirely of assurances. Teams who believe their own absolutes stop monitoring, and users who believed them stop forgiving.

Behaviour on failure

Ask: what happens when delivery fails, and what happens when the model gets it wrong?

A good answer separates the two. For delivery: what you can see about a message's state, and an honest admission if neither read receipts nor guaranteed delivery are on offer. For model error: whether a human sees exactly what will arrive before it does, and what recourse exists afterwards — including a plain statement if nothing can be recalled. A bad answer is silence, or a retry mechanism with no way to inspect what happened.

RelayLink answers with a two-step send no single call bypasses, no mailbox grant in either direction — it is a separate channel, so reading your existing mail is never part of it — per-field provenance the model cannot assert, a delivery floor needing no account on the far end, and blocking enforced at the relay. The rationale sits in the agent safety checklist, the builder-side companion to this list.

It also fails several of these: no end-to-end encryption claim, injection made rare and low-yield rather than impossible, no recall after delivery, and no delivery guarantee beyond ordinary email. If any of that is disqualifying, choosing something else is the correct outcome — which is the point of asking.

Using the list

Run all seven at whoever is selling to you, including the option of building it yourself, which has to answer them too. For each answer, ask where the behaviour is enforced: server, schema, or prompt. If the answer is the prompt, score the item absent.

The vendors worth your time will find these questions ordinary. To see one set of answers end to end, connect your assistant.

Frequently asked questions

What should I ask before letting an AI tool send messages for me?
Ask what it can send without your involvement, what access it holds and for how long, whether recipients can tell your words from the model's, what the person on the receiving end has to install or sign up for, how someone stops your messages permanently, what the vendor admits it cannot do, and what happens when delivery or the model fails.
How can I tell whether an AI communication tool is actually safe?
Safety in this category is structural, not a quality of the model. Look for limits enforced by a server rather than by instructions in a prompt — irreversible actions split into a proposal and a separate confirmation, access scoped no wider than the job needs, and revocation the sending side cannot override. A vendor who answers with model quality has answered a different question.
Does an AI communication tool require the recipient to use it too?
Some do and some do not, and it is worth asking directly, because a tool that only works between adopters is unusable for most of the people you need answers from. The better answer is a delivery floor that reaches an ordinary email address with no account, no install, and no assistant on the far end.