Skip to content

Agent Observability

See exactly what your AI did

NOTHING UNSEEN

An autonomous agent you cannot inspect is a liability. Every action a Xiilio agent takes is recorded in full — what it did, why it did it, what it touched, what it cost and who signed it off.

RECORDED ON EVERY ACTION

Nine signals, captured whether the action succeeded, waited, was blocked or failed. Each one exists because a specific failure is invisible without it.

Timestamp, agent and task

Every action an agent takes, timestamped to the second, with the agent, the customer workspace, the trigger that started it and the run it belongs to.

Without a complete log there is no way to answer 'what happened on the 14th' six months later.

Input and data sources

What the agent was given to work with, and every system and dataset it read during the run.

A decision can only be audited if you know what the agent could see when it made it.

Decision summary and evidence

One plain sentence stating the decision, and the checkable facts it rested on. Never raw model chain-of-thought.

A score with no explanation cannot be challenged, corrected or defended to a regulator.

Tools used and action taken

Each system the agent touched, what it asked for, what came back, how long it took, and what the agent then did or was stopped from doing.

Most agent failures are integration failures. The tool call is where you see them.

Policy checks

Every guardrail evaluated on the run — consent, suppression, spend caps, scope, data residency — with the result of each.

Showing the check ran and passed is the difference between compliance and a claim of compliance.

Confidence and risk class

A confidence figure on every decision, the risk class of the action, and the threshold that decided whether it ran, waited or escalated.

Confidence is what lets you automate the routine 80% and route the rest to a human.

Human approval

The rule that required approval, who decided, what they decided, when, and any note they left.

This is the accountability record. Named person, timestamp, decision.

Outcome and error status

What the action produced — a reply, a booked meeting, a CRM update, a budget shift, nothing at all — and the error if it failed.

Activity is not achievement. Outcome is the only measure worth reviewing.

Cost, model and version

Model spend for each individual action, the token counts in and out, and the exact model version running when the decision was made.

Autonomous systems get expensive quietly, and when behaviour changes the model version is the first thing to check.

WHAT DID MY AI DO TODAY?

Every action, filtered how you need it, with four dashboards reading the same records. Open any action to see the decision, the evidence behind it, every system it touched, the policy checks it passed and what it cost — then replay the whole run step by step. The activity below is an illustrative demonstration, not a client's record.

Actions taken

14

Needing a person

7

Model cost

£7.33

Tokens used

803,500

Showing 14 of 14 recorded actions.

  • Why did the agent do this?

    Ranked 31 of 136 new accounts above the fit threshold and queued them for review; no account was contacted.

    Evidence

    • 12 of the 148 accounts already existed in the CRM and were removed before scoring.
    • 129 of the remaining 136 returned firmographic and technology data; 7 had none and were parked, not scored.
    • Fit weights were derived from closed-won patterns over the last 18 months, not set by hand.
    • 31 accounts scored above the fit threshold agreed with the customer.

    Decision summary and evidence only. Raw model chain-of-thought is not captured or shown.

    Input

    148 accounts added to the source list overnight, plus the ICP definition and 18 months of CRM history.

    Trigger

    Scheduled — daily account refresh, 06:00

    Data sources accessed

    • CRM · accounts
    • CRM · closed-won history
    • Enrichment provider
    • Company register

    Action taken

    Created 31 ranked accounts in the CRM with the fit reasoning attached to each.

    Policy checks

    • PassData source licensingAll enrichment drawn from licensed and public sources.
    • PassNo contact without approvalScoring only — outreach requires a separate approval.
    • PassData residencyProcessing stayed within the customer's configured region.

    Tools used

    • CRM · searchMatch 148 domains against existing accounts → 12 matches found2.1s
    • Enrichment · lookupFirmographics and technology for 136 domains → 129 enriched, 7 no data41.6s
    • CRM · writeCreate 31 ranked accounts with fit reasoning → 31 created8.9s

    Outcome

    31 accounts queued for review. No contact attempted — outreach requires a separate approval.

    Model

    Scoring model

    Model version

    v4.2.1

    Tokens in

    214,800

    Tokens out

    18,400

    Cost

    £1.94

WHERE THIS STOPS

  • The activity on this page is an illustrative demonstration, not a client's real record. Customer names are fictional.
  • A log proves what an agent did. It does not prove the agent was right — that is what review and approval are for.
  • Decision summaries are written for people to read. Raw model chain-of-thought is not captured, stored or shown: you get the decision and the evidence behind it.
  • Cost figures cover model usage for the action. They do not include your subscription or the connected platforms' own charges.
  • Confidence is the model's own estimate. It is useful for routing decisions to people, not a guarantee of accuracy.

The controls that decide what an agent may do unattended are on the Responsible AI Centre. Live approvals and system health sit in the AI Operations Centre.

ASK TO SEE THE LOG

Bring your risk, audit or procurement team. We will walk through the activity record action by action and answer what happens when an agent gets it wrong.

  • ICO Registered

We use a single first-party cookie to remember your currency and consent choice. Page-view analytics are sent to our own servers — no third-party trackers, no profiling. See our Privacy Policy.