THE STATUS QUO

Token Observe vs doing nothing

With three agents, no regulated data and no incident, doing nothing is often the correct decision. This page is about what changes it.

If you have three agents, one owner, no regulated data and no consequential actions, doing nothing is the right answer and installing a fail-closed control in front of them would be a poor trade. Governance has a cost that is usually left out of the argument: a component in the request path, an on-call decision about what happens when it is unavailable, and an approval queue that nobody reads — and an approval queue nobody reads is worse than no gate at all, which is the product’s own wording for a documented agentic-AI failure mode. What changes the answer is not the number of agents but what one of them can now do: a refund, a deployment, an email to a customer, a ticket transition, a row written to a production database. The second thing that changes it is arithmetic you cannot do — when spend is real and you cannot attribute it to an agent and an owner, an unmetered estate and an idle one look identical on every spend surface you have.
When it is right
Few agents, one owner, reversible actions, no auditor
What it costs
Nothing, until the first question you cannot answer
The tell
An unmetered estate and an idle one look identical
The trigger
A refund, a deployment, an email, a ticket, a database write
What buying does not fixAgents that never route through it stay invisible
On this page
where they win

Premature governance has a real cost, and the vendor is not GA either

The cost of doing nothing is zero and the cost of governing too early is not. Token Observe sits in the request path and fails closed, so adding it to three experiments buys an outage mode in exchange for evidence nobody has asked for. It needs a deployment, a database, a backup owner, a restore authority and an emergency decision about what happens when the gateway is unavailable — the design-partner gate requires those to be named in advance, which tells you the shape of the commitment. If the honest answer to who is going to read this record is nobody, the record is a cost with no reader.

There is a specific failure mode worth naming because it is the usual result of governing early. Gates only work while they are rare enough to be read: an approval queue nobody reads is worse than no gate at all, and a policy set that fires on ordinary work trains people to click through it. That is why every policy here can run in shadow mode first, so you learn your false-positive rate before you start blocking real work, and it is also why starting with one policy on one agent is better advice than starting with a governance programme.

The vendor’s own status is part of the honest comparison. Broad general availability is currently a no-go on the product’s own decision record. There is no SOC 2, no ISO 27001, no ISO 42001 and no independent penetration test; no multi-node high availability, no replica and no vendor-operated uptime SLA; and the first deployment it is sold for is a bounded one — one self-hosted install in your own network, roughly five to fifty agents owned by one platform team, non-production-critical, shadow policies before enforcement. Waiting is a defensible position on the vendor as well as on your estate, and a page that pretended otherwise would be contradicting the product’s own record.

the difference

Token Observe and doing nothing, row by row

One card per dimension rather than a three-column table, because the two sides are rarely the same length and a table of them is a horizontal scroller on a phone.

Cost today
doing nothing

Zero. No deployment, no dependency, no on-call decision.

Token Observe

A licence, a self-hosted deployment, a backup and restore owner, and a fail-closed component in front of every governed agent.

What you can answer
doing nothing

Whatever the provider console shows: spend by key, roughly, with no owner attached.

Token Observe

Which agent, whose owner, under which policy, approved by whom, at what cost, with the trace id on every response.

The first incident
doing nothing

The investigation and the record-keeping start on the same day, and the record you want is the one nobody kept.

Token Observe

The record exists from the first governed request, including the blocked ones — a trace id is minted before the verdict, so a refusal is recorded rather than absent.

Regulatory position
doing nothing

Nothing recorded. The EU AI Act’s high-risk obligations, Article 12 record-keeping among them, phase in through 2026 and 2027.

Token Observe

A flight recorder for the governed request lifecycle and a hash-chained log for governance-plane changes — a control that helps evidence a clause, not a certification. Retention is an operator setting, and left unset it keeps traces indefinitely, so counsel has to choose a period that satisfies storage limitation as well as record-keeping.

Discovery
doing nothing

You find out what is running when the invoice arrives, or when someone leaves.

Token Observe

Five radar evidence sources, and only for what you feed it. What it does not see, it reports as unseen rather than as clean.

When doing nothing is the right answer, said plainly

Three tests, and if all three hold, wait. First, can one person name every agent your organisation runs, without asking anyone. Second, is every action an agent takes reversible by the person who notices it — a draft a human sends, a suggestion a human accepts, a summary a human reads. Third, has nobody outside engineering asked you a question about them with a date attached. An estate that passes all three is being governed adequately by the fact that it is small, and the correct next step is to keep it small deliberately rather than to buy a control.

The advice that goes with waiting is cheap and worth taking: write down the list. Not in a governance tool — in whatever your team already reads. An agent, a named human owner, one sentence of purpose, and which provider it calls. That list is the thing that makes the eventual decision easy, and it is also the artefact whose absence makes the first incident expensive. Every regulatory framework that asks anything about AI systems starts by asking for an inventory, and the reason a spreadsheet eventually fails is not that it is a spreadsheet, it is that nothing enforces against it, so it is updated by whoever remembers while the runtime is updated by whoever ships.

What waiting does not buy is a free option on the past. Records that were never kept cannot be reconstructed later, and the product’s own compliance note puts the point without softening it: retrofitting logging onto agents that have been running ungoverned for a year is the expensive path. That is an argument for starting the record early, not for buying a platform early — those are different decisions, and only the second one costs money.

The four things that change the answer

None of them is a headcount threshold or an agent count. They are all changes in kind rather than in scale, and any one of them on its own is usually enough.

An agent can now do something irreversible
A refund, a deployment, an email that leaves the building, a ticket transition, a row written to production. This is the documented buying trigger, and the sentence that goes with it is that the action cannot be reported complete solely because an API returned success. Once an action has a business effect, the question stops being what did the model say and becomes what happened, and who authorised exactly that.
A second provider appears
The moment the same behaviour has to be constrained on two providers, a rule maintained in two places starts to diverge, and the divergence is invisible until it matters. A policy that fires on one provider but not the other is worse than no policy, because it produces a coverage claim you cannot support.
Someone outside engineering asks
Counsel, a data protection officer, an auditor or a customer’s security team. Their questions are field-shaped — who owns this agent, what is it for, what may it touch, who approved this action, keep the logs for how long — and they are answerable from a record or not at all. Reading the record also becomes something that needs to be attributable, which is a property a log file does not have.
Spend is real and unattributed
The point at which the invoice arrives and nobody can decompose it by agent and owner. The awkward part is that the failure is silent in both directions: an explicitly unbudgeted agent can meter at zero, so an unmetered estate and an idle one look identical on every spend surface you have, and nothing prompts you to check.

What buying does not fix, so the comparison stays honest

Token Observe governs what routes through it. Agents that never present a credential to the gateway are outside it, and the documentation puts them under what the product does not evidence rather than in a footnote. The shadow-AI radar exists for exactly that gap and is deliberately reconciliation-based — bills, network egress, service-account keys, IDE and CLI telemetry, and the deployment’s own caller and price consistency — so it tells you those agents exist if you have fed it, and chasing them down inside your organisation remains your work. Where a source has gone silent, coverage travels with every clean result, because a dead feed must never be indistinguishable from a clean estate.

Several other things stay yours after purchase, and they are listed in the support boundary rather than discovered: your infrastructure, your model providers’ incidents, the models’ behaviour, your own agents’ code, and authoring the policies themselves, which is a professional services engagement rather than support. Compliance outcomes are yours too — the mapping documents say the entries mean this feature helps evidence that clause, not installing this makes you compliant, and deciding risk tiers, running an impact assessment and notifying regulators remain the deploying organisation’s duties.

And the record has limits that are stated where the claim is made. Evidence exports are SHA-256 digest-sealed and carry the audit-chain verdict, but they are not themselves signed; durable origin evidence comes from the keyed audit chain plus an Ed25519 anchor retained off-box. The chain is tamper-evident rather than tamper-proof: what an anchor buys is that any copy you kept off-box beats any rewrite made after you took it, and key theft, pre-anchor history, collusion among all sinks, and proof of when all remain open.

the decision

Which of the two you should actually put in.

Both columns are real answers and both are the same length on the page. Read the left one first: if it describes your estate, it is the cheaper decision and this page has done its job.

When to choose doing nothing

  • One person can name every agent you run, and every action they take is reversible by whoever notices.
  • Nobody outside engineering has asked you a question about them, and no date is attached to anything.
  • The cost you would be accepting — a deployment, a backup owner, an on-call decision about a fail-closed dependency — is larger than the risk you are carrying.
  • You would be buying a product whose own decision record says broad general availability is a no-go, with no independent certification and no availability commitment, to solve a problem you do not yet have.

When to choose Token Observe

  • An agent can now do something that moves money or changes a customer’s record, and somebody would have to explain it.
  • You cannot say how many agents you have without asking, or who owns two of them.
  • Spend is material and cannot be attributed to an agent and an owner, which is the point where an unmetered estate and an idle one become indistinguishable.
  • Someone has asked for evidence with a date on it, and the honest answer today is that the record was not kept.

If the left-hand column is the one that describes you, that is still worth an email: a straight answer about which of these to buy costs both of us less than an evaluation that ends in the same place.

Ask which one fits

We have three agents and no incident. Should we do anything at all?

Probably not, beyond writing down the list. An agent, a named human owner, one sentence of purpose and which provider it calls, kept wherever your team already looks. That costs an afternoon, it is what every regulatory framework asks for first, and it makes the eventual decision straightforward instead of archaeological. Installing a fail-closed control in front of three experiments is a poor trade: you would be accepting an outage mode and an operational commitment in exchange for evidence with no reader. Revisit when one of those agents can do something irreversible.

What is the smallest useful step if we do want to start?

One agent, one policy, in shadow mode. Every policy can run in shadow mode first, recording what it would have done without stopping anything, so you learn your false-positive rate before you start blocking real work — and where the deployment turns the gate on, no rule may begin enforcing until a backtest of that exact rule has been replayed against recorded traffic and acknowledged by a named person. Starting with one agent also matches the pilot boundary the product itself recommends: one self-hosted deployment in your own network, roughly five to fifty agents owned by one platform team, non-production-critical, shadow policies first.

What does waiting a year actually cost us?

Mostly the record. Everything else is recoverable — you can add permissions, budgets and approvals to an existing estate in an afternoon each — but a year of agent activity that was never recorded cannot be reconstructed, and the compliance note states the consequence directly: retrofitting logging onto agents that have been running ungoverned for a year is the expensive path. The second cost is discovery. Agents accumulate quietly, and the usual moment of finding out is an invoice or a departure, at which point you are doing an inventory and an investigation at the same time.

Is Token Observe ready to buy?

Not as a general-availability product, and the product’s own decision record says so before any salesperson would. Broad, production-critical general availability is a no-go; there is no SOC 2, no ISO 27001 or 42001 certification, and no independent penetration-test result; there is no multi-node high availability, no replica and no vendor-operated uptime SLA; and the storage at this scale is a single writer on a single host. What is being offered is a design-partner arrangement against a bounded pilot, with a named engineer, direct access, roadmap influence and honest limits rather than a support desk and a service-credit schedule neither party believes in. If that shape does not suit you, waiting is the correct decision and this page would rather say so.

Does this page compare Token Observe against any vendor’s claims?

No — the alternative here is your own status quo, so there is nothing vendor-authored to repeat. On the pages that do name competitors, every claim about another product comes from that vendor’s own public material and has not been independently tested, which is the caveat the product’s own competitive benchmark states about itself. There is no reference production deployment and no independently witnessed bake-off; both are listed among the evidence gates the product has not yet cleared.

get in touch

Tell us which way you are leaning, and why.

Write to hello@tenhaw.com with what your agents do, which providers they call and what would have to be true for you to put something in front of them. James Rooney replies. You will get a straight answer about whether Token Observe fits, including when it does not.

no form · no qualification step · no sales desk · the other three ways in