Token Observe vs doing nothing
With three agents, no regulated data and no incident, doing nothing is often the correct decision. This page is about what changes it.
The alternative on this page is your own status quo, so there is no vendor claim to test — only what your agents can already do.
the other comparisonsOn this page
Premature governance has a real cost, and the vendor is not GA either
The cost of doing nothing is zero and the cost of governing too early is not. Token Observe sits in the request path and fails closed, so adding it to three experiments buys an outage mode in exchange for evidence nobody has asked for. It needs a deployment, a database, a backup owner, a restore authority and an emergency decision about what happens when the gateway is unavailable — the design-partner gate requires those to be named in advance, which tells you the shape of the commitment. If the honest answer to who is going to read this record is nobody, the record is a cost with no reader.
There is a specific failure mode worth naming because it is the usual result of governing early. Gates only work while they are rare enough to be read: an approval queue nobody reads is worse than no gate at all, and a policy set that fires on ordinary work trains people to click through it. That is why every policy here can run in shadow mode first, so you learn your false-positive rate before you start blocking real work, and it is also why starting with one policy on one agent is better advice than starting with a governance programme.
The vendor’s own status is part of the honest comparison. Broad general availability is currently a no-go on the product’s own decision record. There is no SOC 2, no ISO 27001, no ISO 42001 and no independent penetration test; no multi-node high availability, no replica and no vendor-operated uptime SLA; and the first deployment it is sold for is a bounded one — one self-hosted install in your own network, roughly five to fifty agents owned by one platform team, non-production-critical, shadow policies before enforcement. Waiting is a defensible position on the vendor as well as on your estate, and a page that pretended otherwise would be contradicting the product’s own record.
Token Observe and doing nothing, row by row
One card per dimension rather than a three-column table, because the two sides are rarely the same length and a table of them is a horizontal scroller on a phone.
Zero. No deployment, no dependency, no on-call decision.
A licence, a self-hosted deployment, a backup and restore owner, and a fail-closed component in front of every governed agent.
Whatever the provider console shows: spend by key, roughly, with no owner attached.
Which agent, whose owner, under which policy, approved by whom, at what cost, with the trace id on every response.
The investigation and the record-keeping start on the same day, and the record you want is the one nobody kept.
The record exists from the first governed request, including the blocked ones — a trace id is minted before the verdict, so a refusal is recorded rather than absent.
Nothing recorded. The EU AI Act’s high-risk obligations, Article 12 record-keeping among them, phase in through 2026 and 2027.
A flight recorder for the governed request lifecycle and a hash-chained log for governance-plane changes — a control that helps evidence a clause, not a certification. Retention is an operator setting, and left unset it keeps traces indefinitely, so counsel has to choose a period that satisfies storage limitation as well as record-keeping.
You find out what is running when the invoice arrives, or when someone leaves.
Five radar evidence sources, and only for what you feed it. What it does not see, it reports as unseen rather than as clean.
When doing nothing is the right answer, said plainly
Three tests, and if all three hold, wait. First, can one person name every agent your organisation runs, without asking anyone. Second, is every action an agent takes reversible by the person who notices it — a draft a human sends, a suggestion a human accepts, a summary a human reads. Third, has nobody outside engineering asked you a question about them with a date attached. An estate that passes all three is being governed adequately by the fact that it is small, and the correct next step is to keep it small deliberately rather than to buy a control.
The advice that goes with waiting is cheap and worth taking: write down the list. Not in a governance tool — in whatever your team already reads. An agent, a named human owner, one sentence of purpose, and which provider it calls. That list is the thing that makes the eventual decision easy, and it is also the artefact whose absence makes the first incident expensive. Every regulatory framework that asks anything about AI systems starts by asking for an inventory, and the reason a spreadsheet eventually fails is not that it is a spreadsheet, it is that nothing enforces against it, so it is updated by whoever remembers while the runtime is updated by whoever ships.
What waiting does not buy is a free option on the past. Records that were never kept cannot be reconstructed later, and the product’s own compliance note puts the point without softening it: retrofitting logging onto agents that have been running ungoverned for a year is the expensive path. That is an argument for starting the record early, not for buying a platform early — those are different decisions, and only the second one costs money.
The four things that change the answer
None of them is a headcount threshold or an agent count. They are all changes in kind rather than in scale, and any one of them on its own is usually enough.
- An agent can now do something irreversible
- A refund, a deployment, an email that leaves the building, a ticket transition, a row written to production. This is the documented buying trigger, and the sentence that goes with it is that the action cannot be reported complete solely because an API returned success. Once an action has a business effect, the question stops being what did the model say and becomes what happened, and who authorised exactly that.
- A second provider appears
- The moment the same behaviour has to be constrained on two providers, a rule maintained in two places starts to diverge, and the divergence is invisible until it matters. A policy that fires on one provider but not the other is worse than no policy, because it produces a coverage claim you cannot support.
- Someone outside engineering asks
- Counsel, a data protection officer, an auditor or a customer’s security team. Their questions are field-shaped — who owns this agent, what is it for, what may it touch, who approved this action, keep the logs for how long — and they are answerable from a record or not at all. Reading the record also becomes something that needs to be attributable, which is a property a log file does not have.
- Spend is real and unattributed
- The point at which the invoice arrives and nobody can decompose it by agent and owner. The awkward part is that the failure is silent in both directions: an explicitly unbudgeted agent can meter at zero, so an unmetered estate and an idle one look identical on every spend surface you have, and nothing prompts you to check.
What buying does not fix, so the comparison stays honest
Token Observe governs what routes through it. Agents that never present a credential to the gateway are outside it, and the documentation puts them under what the product does not evidence rather than in a footnote. The shadow-AI radar exists for exactly that gap and is deliberately reconciliation-based — bills, network egress, service-account keys, IDE and CLI telemetry, and the deployment’s own caller and price consistency — so it tells you those agents exist if you have fed it, and chasing them down inside your organisation remains your work. Where a source has gone silent, coverage travels with every clean result, because a dead feed must never be indistinguishable from a clean estate.
Several other things stay yours after purchase, and they are listed in the support boundary rather than discovered: your infrastructure, your model providers’ incidents, the models’ behaviour, your own agents’ code, and authoring the policies themselves, which is a professional services engagement rather than support. Compliance outcomes are yours too — the mapping documents say the entries mean this feature helps evidence that clause, not installing this makes you compliant, and deciding risk tiers, running an impact assessment and notifying regulators remain the deploying organisation’s duties.
And the record has limits that are stated where the claim is made. Evidence exports are SHA-256 digest-sealed and carry the audit-chain verdict, but they are not themselves signed; durable origin evidence comes from the keyed audit chain plus an Ed25519 anchor retained off-box. The chain is tamper-evident rather than tamper-proof: what an anchor buys is that any copy you kept off-box beats any rewrite made after you took it, and key theft, pre-anchor history, collusion among all sinks, and proof of when all remain open.
Which of the two you should actually put in.
Both columns are real answers and both are the same length on the page. Read the left one first: if it describes your estate, it is the cheaper decision and this page has done its job.
When to choose doing nothing
- One person can name every agent you run, and every action they take is reversible by whoever notices.
- Nobody outside engineering has asked you a question about them, and no date is attached to anything.
- The cost you would be accepting — a deployment, a backup owner, an on-call decision about a fail-closed dependency — is larger than the risk you are carrying.
- You would be buying a product whose own decision record says broad general availability is a no-go, with no independent certification and no availability commitment, to solve a problem you do not yet have.
When to choose Token Observe
- An agent can now do something that moves money or changes a customer’s record, and somebody would have to explain it.
- You cannot say how many agents you have without asking, or who owns two of them.
- Spend is material and cannot be attributed to an agent and an owner, which is the point where an unmetered estate and an idle one become indistinguishable.
- Someone has asked for evidence with a date on it, and the honest answer today is that the record was not kept.
If the left-hand column is the one that describes you, that is still worth an email: a straight answer about which of these to buy costs both of us less than an evaluation that ends in the same place.
Ask which one fitsThe other comparisons
Same template, same order, same concession first. Claims about every named product on all of them are that vendor's own and have not been independently tested.
Token Observe vs LLM gateways
Token Observe is a gateway in delivery. The gateway is how it arrives, not what it is for.
Token Observe vs LLM observability
One refuses the call inline. The other scores it afterwards. Most estates need both, and they are not substitutes.
Token Observe vs AI security platforms
Token Observe is not a complete AI security suite, and the product’s own strategy document forbids selling it as one.
Token Observe vs cloud-native controls
If every agent, model and tool lives in one cloud, use that cloud’s controls. The argument here is for the estate that does not.
Token Observe vs building it yourself
For a small estate, a few hundred lines of proxy is usually the right answer. The cost arrives later, and it arrives in specific places.
We have three agents and no incident. Should we do anything at all?
Probably not, beyond writing down the list. An agent, a named human owner, one sentence of purpose and which provider it calls, kept wherever your team already looks. That costs an afternoon, it is what every regulatory framework asks for first, and it makes the eventual decision straightforward instead of archaeological. Installing a fail-closed control in front of three experiments is a poor trade: you would be accepting an outage mode and an operational commitment in exchange for evidence with no reader. Revisit when one of those agents can do something irreversible.
What is the smallest useful step if we do want to start?
One agent, one policy, in shadow mode. Every policy can run in shadow mode first, recording what it would have done without stopping anything, so you learn your false-positive rate before you start blocking real work — and where the deployment turns the gate on, no rule may begin enforcing until a backtest of that exact rule has been replayed against recorded traffic and acknowledged by a named person. Starting with one agent also matches the pilot boundary the product itself recommends: one self-hosted deployment in your own network, roughly five to fifty agents owned by one platform team, non-production-critical, shadow policies first.
What does waiting a year actually cost us?
Mostly the record. Everything else is recoverable — you can add permissions, budgets and approvals to an existing estate in an afternoon each — but a year of agent activity that was never recorded cannot be reconstructed, and the compliance note states the consequence directly: retrofitting logging onto agents that have been running ungoverned for a year is the expensive path. The second cost is discovery. Agents accumulate quietly, and the usual moment of finding out is an invoice or a departure, at which point you are doing an inventory and an investigation at the same time.
Is Token Observe ready to buy?
Not as a general-availability product, and the product’s own decision record says so before any salesperson would. Broad, production-critical general availability is a no-go; there is no SOC 2, no ISO 27001 or 42001 certification, and no independent penetration-test result; there is no multi-node high availability, no replica and no vendor-operated uptime SLA; and the storage at this scale is a single writer on a single host. What is being offered is a design-partner arrangement against a bounded pilot, with a named engineer, direct access, roadmap influence and honest limits rather than a support desk and a service-credit schedule neither party believes in. If that shape does not suit you, waiting is the correct decision and this page would rather say so.
Does this page compare Token Observe against any vendor’s claims?
No — the alternative here is your own status quo, so there is nothing vendor-authored to repeat. On the pages that do name competitors, every claim about another product comes from that vendor’s own public material and has not been independently tested, which is the caveat the product’s own competitive benchmark states about itself. There is no reference production deployment and no independently witnessed bake-off; both are listed among the evidence gates the product has not yet cleared.
Tell us which way you are leaning, and why.
Write to hello@tenhaw.com with what your agents do, which providers they call and what would have to be true for you to put something in front of them. James Rooney replies. You will get a straight answer about whether Token Observe fits, including when it does not.
no form · no qualification step · no sales desk · the other three ways in