All posts
agent-docs

Hand Off an AI Pass Feature with Reproducible Evidence

A precise handoff separates implementation, mocked checks, OAuth observations, deployment, and genuinely observed wallet-funded inference.

EiliyaOctober 9, 20263 min read

Completion is a claim that needs a bounded evidence bundle

An integration can compile, render a connection button, and pass mocked tests without completing a wallet-funded action. A useful agent handoff makes those distinctions visible. The reviewer should be able to reproduce the checks and see exactly where evidence stops, rather than interpreting “done” as a promise about production funding.

Consider an illustrative document organizer called Archivebox. The agent has added an optional AI Pass action that proposes labels for an uploaded document. Existing sign-in, hosting, storage permissions, and direct BYOK are preserved. Local tests pass, but no user has approved a paid model request. This is a legitimate implementation milestone, not evidence of live wallet-funded success.

The canonical integration skill provides a verification reference and explicitly requires honest pending status when paid verification has not occurred. Use that current reference for exact runtime expectations rather than expanding a successful build into a broader claim.

Organize evidence by what it proves

A practical bundle can use the following local structure. These filenames are a proposed handoff convention, not AI Pass API resources.

evidence/
  scope.md
  environment.json
  commands.json
  changed-files.txt
  test-results.txt
  route-preservation.md
  catalog-observation.json
  runtime-status.json
  reproduction.md
  limitations.md

scope.md identifies the exact action, selected integration path, and exclusions. environment.json records non-secret runtime and package-manager versions. commands.json records executed commands, working directories, timestamps, and exit codes. Avoid recording full environment dumps, which frequently contain credentials.

route-preservation.md maps host login, existing BYOK, subscriptions or credits, and deployment behavior to specific checks. catalog-observation.json records discovery evidence separately from inference. A successful read of the public catalog demonstrates public metadata availability, not an authenticated model call.

Give each claim its own status

For Archivebox, an example status record might read as follows. These are illustrative values for a hypothetical handoff, not reported execution results:

{
  "feature": "document-label-suggestions",
  "implementation": "built",
  "localMockTests": "passed",
  "publicCatalog": "observed",
  "liveOAuth": "not_observed",
  "walletFundedInference": "pending_approval",
  "productionDeployment": "not_verified"
}

Do not collapse those fields into one green badge. Likewise, deployment success does not establish that the intended user can connect, that an operation uses the right payer, or that a model result reaches the actual product screen.

A recorded exit code should belong to a command actually executed against the identified revision. Copying a previous run's output into a new bundle produces misleading evidence. Preserve raw useful output, then add a short explanation of its scope. Redact secrets before retaining logs or screenshots, without changing the substantive test outcome.

Make checks reproducible without spending

reproduction.md should start with the repository revision and documented dependency installation command. Then list the precise build and test commands used by that repository. Include the fixture selection or environment switch that disables network inference. If a fake adapter is necessary, describe how the reviewer can verify it was selected.

For Archivebox, the reproducible tests should cover document authorization, route selection, one dispatch per operation, result rendering, and existing BYOK preservation. A mock label result must be visibly identified as fixture output. Do not include a live inference command in a supposedly safe smoke-test script.

Use the SDK documentation or REST documentation according to the implemented path. Link the relevant current contract in the bundle, but do not store token-bearing callback URLs, runtime OAuth responses, setup recovery records, or provider credentials as debugging evidence.

Separate the later live check from this handoff

When the user separately approves a specific paid action and its cost basis when knowable, the live check can observe one user action, one model request, and the real result rendered in the product. Authenticated reuse can be inspected without issuing another paid call. Record what was actually observed and leave anything unavailable explicitly unverified.

Do not fund a wallet automatically to remove a blocker. Setup approval is not spending approval, and a missing balance is not permission to change payment settings. If the live action cannot proceed, preserve the implementation evidence and identify the blocked criterion.

The final handoff should say what changed, which commands passed, what existing behavior was preserved, and what remains pending. For Archivebox without paid verification, the appropriate summary is “implemented and built; live wallet-funded verification pending.” That statement is more useful than an unsupported production-ready claim because the next operator knows exactly which check remains.

For AI agents

Skill file

Download
---
name: aipass-feature-evidence-handoff
description: Use when handing off an AI Pass feature. Build a reproducible evidence bundle and separate local checks from live funding claims.
---

# Feature verification and handoff evidence

## Scope
Audit a specified implementation revision and produce a redacted evidence bundle. No publishing, deployment, wallet funding, or paid inference is authorized by this skill. Do not describe a build or mock as live production validation.

## Prerequisites
Read the [canonical integration skill](https://aipass.one/skills/aipass-integration/SKILL.md) and its verification reference. Exact connection, setup, and token handling come from its selected implementation references. Consult [SDK docs](https://aipass.one/docs/sdk) or [REST docs](https://aipass.one/docs/rest) for the actual path.

## Procedure
1. Identify the revision, working tree changes, visible action, selected path, deployment target, and preserved host login and billing behavior.
2. Discover repository commands from manifests and documentation. Run build and relevant tests with an explicitly selected fake adapter or otherwise network-disabled inference boundary.
3. Record command, working directory, timestamp, exit code, and useful redacted output. Never recycle results from another revision.
4. Map preservation checks to host authorization, cross-user resource denial, direct BYOK, subscription or credit rules, duplicate dispatch prevention, and result rendering.
5. Record catalog discovery separately. A public metadata response is not OAuth completion or funded inference.
6. Assign independent statuses for implementation, local tests, public discovery, live OAuth, wallet-funded inference, and production deployment. Use pending or not observed where evidence is absent.
7. Write safe reproduction instructions including dependency setup, fixture configuration, and exact checks. Do not hide paid calls inside a smoke-test command.
8. Inspect bundle contents for tokens, provider keys, recovery records, private prompts, and token-bearing callback URLs. Exclude or redact before handoff.

## Deliverables
Create `evidence/scope.md`, `environment.json`, `commands.json`, `changed-files.txt`, `test-results.txt`, `route-preservation.md`, `catalog-observation.json`, `runtime-status.json`, `reproduction.md`, and `limitations.md`. Mark missing observations explicitly instead of inventing file contents.

## Completion wording
Report actual changed files, actual checks, and blocked criteria. Without separately approved and observed paid inference, say: “implemented and built; live wallet-funded verification pending,” but only if implementation and build really passed. A later approved check must observe the real action and result; never automatically fund a wallet or infer production funding from setup success.