Blog · · 11 min read

AI agent tool contract testing before API changes ship

AI agent tool contract testing before API changes ship

AI agent tool contract testing catches a class of failure that ordinary API tests miss. A backend can still return HTTP 200, satisfy its OpenAPI schema, and pass service tests while the agent starts choosing the wrong tool, inventing required values, mishandling errors, or claiming that an action succeeded. The break sits between the product API and the model-facing tool definition. Fix it by versioning that complete boundary, testing the provider and consumer independently, running deterministic schema and error cases, then checking model behavior before promotion. This guide gives API platform and agent engineers a concrete contract artifact, compatibility rules, CI sequence, and rollout check for existing products.

Why ordinary API compatibility is not enough

An agent does not consume an endpoint the way application code does. It first reads a tool name and description, chooses among competing tools, generates arguments from a schema, receives a success or error result, and decides what to do next. Every stage adds a contract that may change even when the underlying endpoint remains technically compatible.

OpenAI defines function tools through JSON Schema and recommends strict mode so generated calls adhere to the declared schema. Strict mode also imposes concrete rules: object schemas must set additionalProperties to false, and fields in properties must be listed as required, with nullable types used for optional values. A converter that handles those rules incorrectly can alter model behavior without changing the product API.

OpenAI Agents SDK issue #3733 reports a generated schema that marked optional fields as required. The author observed models filling the forced fields with empty strings, zeroes, or literal null strings. This is an author report, not a universal SDK failure. It shows how a schema can validate while expressing the wrong requirement.

Record these six parts of the tool boundary:

  1. Discovery semantics: tool name, description, examples, and the other tools presented beside it.
  2. Input semantics: types, required fields, nullability, formats, defaults, and business constraints.
  3. Output semantics: structured success data, identifiers, versions, and evidence that an effect occurred.
  4. Error semantics: stable error types, retryability, user-action requirements, and safe text for the model.
  5. Authority and effect semantics: required scopes, tenant boundary, approval class, idempotency, and side-effect level.
  6. Runtime behavior: deadlines, rate limits, concurrency expectations, and version compatibility.

If your tests snapshot only the HTTP request and response, they cover part of items two and three. The rest can drift unnoticed.

Define one versioned tool-contract artifact

Put the model-facing and operational contract in one reviewed artifact. Do not reconstruct it independently in the API gateway, agent service, prompt builder, and documentation site. Those copies will diverge.

A minimal artifact can look like this:

tool_id: billing.issue_refund
tool_version: 4
provider:
  service: billing-api
  operation: POST /v2/refunds
  minimum_version: 2026-08-01
model_surface:
  name: issue_refund
  description: Issue one refund for an eligible order after approval.
  input_schema_ref: schemas/issue-refund-input-v4.json
  output_schema_ref: schemas/issue-refund-output-v3.json
errors:
  - code: ORDER_NOT_FOUND
    retryable: false
    user_action: verify_order
  - code: PROVIDER_TIMEOUT
    retryable: true
    retry_after_source: response
policy:
  required_scope: refunds.write
  effect: irreversible_external_write
  approval: always
  idempotency_key: action_id
compatibility:
  consumes: [billing-api.refund.v2]
  replaces: billing.issue_refund@3

The artifact is not a second implementation of your API specification. It records what the agent adapter promises to the model and what it expects from the provider. Keep the referenced JSON Schemas machine-readable, but keep effect, authority, and error policy explicit too.

The current Model Context Protocol Tools specification defines inputSchema and optional outputSchema. Servers must return conforming structured results when they declare an output schema, and clients should validate those results. The specification also separates protocol errors from tool-execution errors. Use the same separation internally even when you do not expose the tool through MCP.

Classify changes by what they can break

A useful compatibility checker compares JSON trees and classifies what each change can break: the model, the adapter, the provider, or an in-flight run.

ChangeDefault classificationWhy it can break the agentSafe release path
Remove or rename an input fieldBreakingExisting traces and adapters still send itAdd a new tool version and keep the old adapter
Make an optional field requiredBreakingThe model may invent a filler valueAdd a nullable field first, collect data, then version
Narrow an enum or numeric rangeBreakingPreviously valid calls now failVersion and add explicit migration cases
Add an optional output fieldUsually compatibleConsumers can ignore it if unknown fields are allowedProvider verification plus consumer test
Remove an output fieldBreakingThe agent may lose effect evidence or an identifierRun both output versions during migration
Change a tool descriptionBehavior-sensitiveTool selection can change without a schema diffRun selection and no-call golden cases
Change an error shapeBreaking unless normalizedRecovery, retries, and user messages can changeMap both provider errors to one typed tool error
Expand required authorityPolicy-breakingExisting runs may lack the new grantStop or reauthorize in-flight work
Change read behavior into a writeSafety-breakingApproval and idempotency assumptions become falseCreate a new tool ID and require policy review

A wording change in a tool description can look harmless in source control but shift which tool the model chooses. Compare normalized text for review, then run behavior cases against the exact candidate tool set.

Run a separate gate for error changes. MCP issue #3003 identifies an open gap around schema-governed tool-execution errors: structured success results can be validated against outputSchema, while the error example uses unstructured content. The report warns that unstructured error text can carry internal identifiers or personal data into model context. Your product can define a typed internal error envelope before any text reaches the model, whether or not your runtime adopts the proposed MCP change.

Build the contract test harness in six layers

Run fast deterministic layers on every pull request. Run model-dependent checks on changes that touch the tool surface, and use a canary before general release.

1. Lint and fingerprint the candidate contract

Validate every JSON Schema against its declared dialect. Reject duplicate tool names, missing descriptions, unbounded free-form objects, undocumented error codes, and write tools without authority or idempotency metadata.

Create a canonical representation and hash it. Store that fingerprint in the agent release manifest and every execution trace. A later incident should answer exactly which contract the model saw.

2. Run a semantic compatibility diff

Compare the candidate against every supported contract version. JSON Schema diff tools can detect structural changes, but add product rules for descriptions, scopes, effects, approvals, and errors.

Return a typed result instead of a pass/fail string:

{
  "tool_id": "billing.issue_refund",
  "baseline": 3,
  "candidate": 4,
  "classification": "breaking",
  "changes": [
    {
      "path": "input.reason",
      "kind": "optional_to_required",
      "affects": ["model_arguments", "saved_runs"]
    },
    {
      "path": "errors.PROVIDER_TIMEOUT",
      "kind": "retryability_changed",
      "affects": ["recovery_policy"]
    }
  ]
}

A release rule can allow compatible additions, require review for behavior-sensitive edits, and block breaking or safety-breaking changes until a new version and migration plan exist.

3. Verify the provider against consumer expectations

Treat the agent adapter as a consumer and the product service as a provider. Pact's contract-testing model has the consumer declare the minimal request and response it needs, then verifies that the real provider supplies them. The two sides do not need to run together for every consumer test.

For each tool, verify normal success, every documented business error, malformed input, expired authority, duplicate action ID, timeout, and stale record version. Assert only what the agent consumes. Overly broad response snapshots make harmless provider additions look breaking and encourage teams to bypass the gate.

4. Test the adapter as a deterministic consumer

Stub provider responses and prove that the adapter emits one stable tool result. The adapter should normalize backend details into model-safe outcomes.

For a refund tool, verify that:

  • a duplicate action returns the original receipt instead of issuing a second refund;
  • a provider timeout becomes TEMPORARY_FAILURE, not an empty success;
  • an ineligible order becomes ACTION_REJECTED with a user-safe reason;
  • an unknown response shape fails closed and never claims completion;
  • internal stack traces, tokens, and personal data are removed before model delivery.

These checks should not call a model. They are ordinary deterministic tests and belong in the fastest CI tier.

5. Run model-facing golden cases

Model-facing tests cover the part ordinary contract frameworks cannot: how the model uses the tool surface. Use the candidate definition and the same neighboring tools used in production.

Include positive cases, negative no-call cases, ambiguous requests, missing required information, similar tools, denied authority, provider failures, and old conversation state. Assert properties rather than exact prose:

  • the correct tool is selected when an action is eligible;
  • no tool is called when the user has not authorized the action;
  • argument values come from the conversation or current product context;
  • the model asks for missing information instead of inventing it;
  • a failed provider result never becomes a success claim;
  • the trajectory does not retry beyond the declared policy.

Current AI agent testing guidance from Confident AI separates tool selection, argument validation, trajectory checks, regression cases, and simulated failures. Use those checks after deterministic contract verification, not instead of it. A model test can expose behavioral drift, but it cannot prove that every provider response satisfies the declared schema.

6. Canary the exact contract and watch deltas

A green test suite does not justify replacing every tool definition at once. Route a small cohort to the new contract, preferably in shadow mode for write tools. Compare tool-selection rate, missing-argument rate, validation failures, error-code distribution, retries, user corrections, approved actions, and verified outcomes against the prior version.

Keep both contract versions available while paused or long-running work drains. Rollback must restore the tool definition, adapter behavior, policy, and provider compatibility together. Reverting only the prompt or description can leave an incompatible adapter active.

Work through a realistic schema-drift failure

Assume an existing billing endpoint accepts an optional reason field. The tool generator is upgraded and now lists every property as required under strict mode. The product API still accepts all old requests. Unit tests pass because fixtures already include a reason.

The model now receives a required reason for requests where the user supplied none. It invents "" or "customer_request". The refund succeeds, but the audit reason is false. This is a contract failure even though no HTTP or JSON validation error occurred.

The layered harness catches it in four places:

  1. The semantic diff flags reason as optional-to-required.
  2. A deterministic missing-reason case proves the adapter should request user input instead of executing.
  3. Golden cases check that the model does not fabricate the reason.
  4. The canary watches for a jump in empty or default reason values.

The fix is not to loosen every schema. Decide whether the product truly needs the reason. If it does, create a new tool version that asks for it explicitly and migrate workflows. If it does not, represent absence correctly for the model and keep the provider field optional.

Handle tradeoffs without weakening the gate

Contract tests add artifacts and release coordination. Keep the burden proportional to risk.

For read-only tools, compatible output additions may pass with provider and consumer tests. For writes involving money, messages, access, deletion, or external commitments, require policy review, deterministic error cases, model regressions, and a canary. A single rule for every tool will either slow harmless changes or under-test dangerous ones.

Do not use model nondeterminism as a reason to tolerate structural failures. Separate deterministic guarantees from statistical behavior. Schemas, error mappings, authorization, idempotency, and redaction should pass every time. Tool selection and trajectory can use repeated cases and thresholds, but safety-critical cases should remain non-negotiable.

Avoid turning the tool registry into a second API gateway. The contract describes expectations and policy; the provider still owns business logic. Generate what you can from authoritative sources, then require explicit review for model descriptions, effects, authority, and error semantics that code generation cannot infer safely.

Verify the release gate closes the gap

Before approving a tool or product API change, require evidence for each item:

  • The candidate contract has a canonical fingerprint and explicit predecessor.
  • Input and output schemas validate against the declared JSON Schema dialect.
  • The compatibility diff covers schema, description, errors, authority, effect, approval, and idempotency.
  • Every breaking change has a new version, migration owner, and retirement condition.
  • Provider verification covers success and every declared business error.
  • Adapter tests fail closed on unknown, malformed, or sensitive provider results.
  • Golden cases cover correct selection, no-call behavior, missing facts, denied authority, and recovery.
  • The test trace records the model, neighboring tools, tool-contract fingerprint, and policy version.
  • The canary has comparison metrics, a rollback threshold, and a named owner.
  • Paused and long-running work remains pinned to a compatible contract.

Contract testing for AI pipelines applies consumer-driven schema checks to probabilistic component handoffs. An existing product also needs to cover model-visible tool semantics and operational policy, not only the data shape. When every checklist item above has machine-readable evidence, reviewers can approve an API change against the contract the agent consumes.

Start with one write tool

Choose one high-impact write tool and export its current name, description, input schema, output shape, errors, scope, approval rule, and idempotency behavior into a versioned ToolContract. Add one semantic diff and three deterministic cases: missing required information, a provider timeout, and an unknown response field. Run them in the provider's pull request workflow. The results will show where the product API and agent platform disagree before a routine backend change turns that disagreement into a customer-facing failure.

References

Be first in line.

Join the waitlist and we'll email you the moment it's ready. No sales call.

Try the demo →
or talk to us