Blog · · 11 min read

AI agent result verification before claiming success

AI agent result verification before claiming success

AI agent result verification stops a confident model response from becoming false product state. A tool may return success before a write is visible. Every attempted call may fail even though the workflow ends normally. Empty output can look like proof that nothing exists. If your application accepts the model's summary as the result, users see completed work that never happened. Completion needs to be an application decision, based on a runtime-issued receipt, a typed outcome, and an authoritative postcondition check. This guide defines that contract and shows how to retry without duplicating side effects.

Why a successful tool call is not proof

An agent runtime usually sees several different facts during one action:

  1. The model asked to call a tool.
  2. The adapter accepted the arguments.
  3. A downstream service returned a response.
  4. The intended business state changed.
  5. The model described the result correctly.

Those facts are related, but none proves the next one. An HTTP response can acknowledge receipt while processing continues. An adapter can catch an exception and return friendly text. A write can reach one service but remain invisible in the system that defines the real state. The model can also infer success from silence because it has no explicit negative result.

Issue reports show both failure modes. A Harness issue describes an agent step recorded as successful after every tool call returned 401. A Simple Wiki issue describes five edit attempts returning success: true while authoritative rereads and the version hash showed no persisted change. These are author reports about specific systems, not claims about every framework. In both cases, workflow completion depended on an unverified result surface.

The model should propose actions and explain verified results. It should not decide whether infrastructure evidence is sufficient. Keep that decision in deterministic product code.

Define a three-layer completion contract

Treat completion as the conjunction of three independently checked layers:

  • An execution receipt records which tool ran, the action ID, the tenant and target, and what the runtime observed.
  • A typed outcome classifies the call as accepted, rejected, retryable failure, permanent failure, or unknown. Empty text is not a valid outcome.
  • A postcondition verifier reads the system of record and decides whether the intended business state now holds.

A result may be operationally successful without satisfying the business postcondition. For example, a ticket API may accept a close request but reject the transition later because a required resolution note is missing. Conversely, a request may time out after the downstream service commits it. Retrying immediately could then repeat the effect.

The current Model Context Protocol Tools specification supports structured tool content, output schemas, and an isError flag. Those are useful transport primitives. They still do not prove that a business effect occurred. Your product needs a stronger internal contract even if its external tool protocol is MCP.

Represent the three layers in one machine-readable object:

{
  "action_id": "act_01K2A7M9Q4",
  "tool_id": "support.close_ticket",
  "tenant_id": "tenant_42",
  "target": {"ticket_id": "T-1842"},
  "execution": {
    "receipt_id": "rcpt_01K2A7P1BZ",
    "runtime_status": "completed",
    "started_at": "2026-08-10T09:14:22Z",
    "finished_at": "2026-08-10T09:14:23Z",
    "attempt": 1
  },
  "outcome": {
    "kind": "accepted",
    "provider_code": "202",
    "retryable": false
  },
  "verification": {
    "status": "pending",
    "postcondition": "ticket.status == closed",
    "checked_at": null,
    "evidence_version": null
  }
}

Do not let the model create or edit receipt_id, runtime timestamps, provider codes, or verification evidence. The runtime issues those values after it observes the call. A current research proposal, Tool Receipts, Not Zero-Knowledge Proofs, describes the same separation: runtime-issued signed receipts record execution facts, then a deterministic verifier checks claims against those facts.

Implement AI agent result verification in six steps

1. Assign a stable action identity

Create action_id before the first write attempt. Derive it from the product operation, tenant, target, and logical user request, or generate it once and persist it with the workflow. Reuse that ID across retries and resumes.

Pass the action ID as the downstream idempotency key when the provider supports one. If it does not, maintain a product-owned action ledger that maps the logical action to provider identifiers and observed state. This prevents a timeout from turning into a second independent write.

The action ID identifies intent. The receipt ID identifies one execution attempt. Do not collapse them. One action may have several attempts, but each attempt must have exactly one immutable receipt.

2. Issue receipts outside the model context

Wrap every effectful tool in runtime code that records the attempt before returning control to the model. The receipt should contain:

  • action, attempt, tool, tenant, actor, and target identifiers;
  • a hash or version of normalized arguments, with secrets excluded;
  • start and finish timestamps;
  • transport status and normalized provider outcome;
  • downstream operation or resource identifiers;
  • the policy, tool-contract, and agent-release versions used;
  • a pointer to protected raw evidence when incident review needs it.

Store the receipt in an append-only or audit-protected record. Give the model a safe projection, not the signing key, raw credentials, or unrestricted logs. The runtime must be able to reject a claimed receipt that was never issued.

A receipt proves what the runtime observed. It does not prove the final product state. Mark that limitation explicitly so downstream code cannot equate runtime_status: completed with business success.

3. Return typed positive and negative outcomes

Every adapter path must return a typed result. Ban ambiguous conventions such as empty strings, missing fields, or generic prose like done.

Use a small outcome vocabulary:

  • accepted: the provider accepted work, but the postcondition may still be pending;
  • rejected: the provider made an explicit business decision and retrying unchanged input will not help;
  • retryable_failure: the call did not complete and policy permits another attempt;
  • permanent_failure: the call cannot succeed without a code, policy, authority, or data change;
  • unknown: the runtime cannot determine whether the effect happened.

Make explicit negatives first-class. A search tool should return found: false with the searched scope and a completed receipt. It should never use silence to mean no result. A Claude Code issue describes silent verifier output being interpreted as a confirmed negative. The proposed repair uses explicit positive and negative sentinels, which is the right boundary for deterministic consumers.

4. Verify the postcondition against the authority

Define the business postcondition when the tool is registered, not after an incident. For support.close_ticket, it might be:

status == "closed"
AND resolution_note_id == requested_resolution_note_id
AND version >= version_observed_before_action + 1

Read from the system that owns each fact. Do not verify a database write through the agent's cached context, the adapter's response echo, or the same replica that may be stale. If the product owns the database, use a consistent read where practical. If an external provider owns the state, query its authoritative resource endpoint using the provider operation ID captured in the receipt.

Keep verification deterministic. Anthropic's agent-evaluation guidance recommends outcome-oriented criteria and deterministic graders where possible. At runtime, the same principle means comparing exact state, versions, identifiers, and invariants instead of asking another model whether the action probably worked.

5. Reconcile delayed and unknown outcomes

Some systems are eventually consistent. A postcondition that fails one second after an accepted write may pass ten seconds later. Do not treat that delay as immediate proof of failure, and do not retry the write before checking whether the first attempt committed.

Use a bounded reconciliation policy:

  1. Read the authoritative state immediately when the provider claims completion.
  2. If the state is not visible but the provider says processing continues, mark the action unverified and schedule another read.
  3. Retry verification with bounded backoff until the visibility deadline.
  4. Before any write retry, query by action ID, idempotency key, or provider operation ID.
  5. Retry the write only when evidence establishes that the original effect did not occur and policy still permits it.
  6. Escalate an unresolved unknown outcome instead of guessing.

A 2026 preprint on verified tool calls under non-atomic failures describes this verify-before-retry pattern with postcondition checks and idempotency keys. Its controlled simulation supports the mechanism, but your production thresholds still need to come from your own provider behavior and risk limits.

6. Gate user-visible success on verification

The final response builder should consume the deterministic completion status, not free-form model confidence. Give it only three user-visible states:

  • verified: the receipt is valid, the outcome permits success, and the postcondition holds;
  • unverified: execution evidence exists, but the product cannot yet prove the postcondition;
  • failed: the receipt or outcome proves failure, or the postcondition is definitively false.

The model may explain those states in natural language, but templates or policy checks should prevent it from converting unverified into completed. Persist the state in the workflow record so the UI, alerts, audit log, and follow-up jobs agree.

For a delayed ticket closure, say: "The close request was accepted, but the ticket is not yet confirmed closed. Verification is still running." That is less fluent than a false success claim and far more useful.

Work through a realistic failure

Assume an agent must add a resolution note and close ticket T-1842. The close tool returns success: true, but the note write silently fails. The product requires a resolution note before a ticket can remain closed, so an asynchronous rule reopens it.

Without result verification, the model sees success: true and tells the user the ticket is closed. Minutes later, the ticket is open again and no system can explain which part failed.

With the completion contract:

  1. The runtime creates one action ID for the combined business intent.
  2. The note and close attempts receive separate immutable receipts under that action.
  3. The close adapter returns accepted, not verified.
  4. The verifier reads the ticket, required note ID, and record version from the support system.
  5. The missing note makes the compound postcondition false.
  6. The workflow records failed, links both receipts, and asks for recovery instead of claiming completion.

If the close call times out, the verifier reads the ticket before retrying. If the ticket is already closed with the expected note, the workflow marks it verified and does not send another close request. Verification determines whether a retry is safe. The stable action ID then prevents that permitted retry from becoming a new logical action.

Handle failures without inventing certainty

Use explicit decision rules for each evidence state:

  • With no receipt, treat the attempt as not executed. Model text cannot substitute for runtime evidence.
  • For a rejected or permanently failed receipt, mark the action failed and surface the safe reason. Do not retry unchanged input.
  • For a retryable failure where no effect is possible, retry within the attempt and time budget.
  • For an unknown outcome, read the authoritative state before retrying the write.
  • For an accepted outcome with a pending postcondition, reconcile until the visibility deadline.
  • When the postcondition is definitively false, mark the action failed or invoke an explicit compensating workflow.
  • When the verifier is unavailable, mark the action unverified. Missing verification cannot count as success.
  • When the receipt and state disagree, preserve both. Stop autonomous escalation for high-risk actions and alert an owner.

Keep tool errors separate from protocol errors. A malformed request, expired credential, rejected business operation, provider timeout, and verifier outage require different recovery. Flattening them into one string deprives the runtime of the information needed to decide safely.

Test the completion contract with failure injection

Unit tests should cover each adapter branch without calling a model. Integration tests should inject failures between the downstream write and the receipt update, between receipt persistence and verification, and during the authoritative read.

Require these cases before an effectful tool ships:

  • every execution attempt creates one receipt, including exceptions and timeouts;
  • no success response can be emitted without a valid receipt ID;
  • explicit negative results remain negative after model summarization;
  • a successful provider response with unchanged state remains unverified or failed;
  • a timeout after a committed write is discovered before another write attempt;
  • repeated reconciliation reads never repeat the side effect;
  • stale replicas cannot satisfy the postcondition when a stronger read is required;
  • verifier failure produces unverified, never verified;
  • forged or mismatched receipt IDs fail closed;
  • UI, workflow status, alerts, and audit records expose the same completion state.

Run model-facing cases too, but assert the boundary rather than exact prose. The model may choose different wording. It must not claim success for unverified or failed, omit the unresolved state, or recommend an unsafe manual retry.

Shortcuts that weaken the evidence boundary

Do not ask a second model to judge whether the first model's claim is true. A model grader can help evaluate communication or fuzzy task quality, but it cannot replace execution evidence and authoritative state reads.

Do not treat observability as verification. Traces can show that a tool was called and returned text. They become proof only when the trace references immutable receipts and state evidence with defined semantics.

Do not reserve verification for high-risk tools. The depth can vary by risk, but every tool needs an explicit result contract. Read-only tools need a clear distinction between found, not found, denied, failed, and unknown. Write tools also need receipts, idempotency, postconditions, and reconciliation.

Start with one false-success incident

Choose one effectful tool that has produced a disputed or hard-to-explain result. Define its stable action ID, runtime receipt, typed outcomes, authoritative postcondition, visibility deadline, and verify-before-retry rule. Then inject a timeout after the downstream commit and prove that the product detects the completed effect without repeating it. Do not ship the gate until the UI and workflow record show unverified when the verifier is unavailable. That single AI agent result verification test establishes the contract your other tools can adopt.

References

Be first in line.

Join the waitlist and we'll email you the moment it's ready. No sales call.

Try the demo →
or talk to us