All posts
agent-docs

Test Duplicate AI Actions and Unknown Paid Outcomes

Separate duplicate intent from lost responses, preserve unknown outcomes, and test dispatch counts without claiming remote exactly-once behavior.

EiliyaOctober 7, 20263 min read

A timeout is not proof that inference failed

An AI action can finish remotely while the application loses its response. Retrying immediately may create a second paid action. Separately, a double-click can dispatch two requests before either completes. These are different failure modes: duplicate intent and an unknown outcome after dispatch.

Consider an illustrative storyboard editor called Frameboard. A user requests one concept image, clicks again while the spinner appears, and then loses connectivity. The app must preserve the original operation rather than interpreting every click or reconnect as authorization for another image. AI Pass does not remove this application-level responsibility.

Consult the canonical integration skill for the real integration and spending boundary. Its setup provisioning idempotency is not evidence that every inference endpoint deduplicates retries. Check the selected method in the REST documentation before relying on any server-side idempotency or status lookup.

Write a decision table before retry logic

Assign an operation identifier to a deliberate user intent before dispatch. Bind it to the host user, selected payer route, model choice, and an immutable request fingerprint. Reusing an identifier with different content should be a conflict, not an update.

Observed state Repeated same operation User-visible behavior
Prepared, not dispatched Claim once and dispatch Starting
Dispatch claimed Do not start a second request In progress
Result stored Return stored result Completed
Failure proven before dispatch Allow deliberate retry policy Not sent
Connection lost after dispatch Preserve unknown outcome Completion and charge unknown

“Failure proven before dispatch” needs actual transport or application evidence. A generic timeout is not enough. Where dispatch cannot be classified safely, use unknown. If the endpoint offers no documented reconciliation mechanism, explain the uncertainty and require a deliberate new action before another potentially paid attempt.

Exercise the policy with a deterministic fake

This executable Python example uses no network, real credential, or image generation. The fake deterministically loses its response after recording a call. The operation cache then blocks a repeated dispatch. It demonstrates serial policy behavior, not production concurrency safety.

class LostResponse(Exception):
    pass

class FakeAdapter:
    def __init__(self):
        self.calls = 0

    def generate(self):
        self.calls += 1
        raise LostResponse()

class Actions:
    def __init__(self, adapter):
        self.adapter = adapter
        self.states = {}

    def submit(self, operation_id):
        if operation_id in self.states:
            return self.states[operation_id]
        self.states[operation_id] = "dispatch_claimed"
        try:
            self.adapter.generate()
        except LostResponse:
            self.states[operation_id] = "unknown"
        else:
            self.states[operation_id] = "completed"
        return self.states[operation_id]

fake = FakeAdapter()
actions = Actions(fake)
assert actions.submit("storyboard-action") == "unknown"
assert actions.submit("storyboard-action") == "unknown"
assert fake.calls == 1

Extend the test suite with a successful fake, a pre-dispatch validation rejection, and a second deliberate operation. Assert adapter call counts as well as UI text. A screen that says “once” can still hide two network calls.

Move the guarantee into durable application state

An in-memory map is lost on restart and does not coordinate two workers. Production needs an atomic claim backed by durable storage, a uniqueness constraint scoped to user and operation, and result persistence. Concurrent requests must compete for one claim rather than perform separate read-then-write checks.

Test the crash window between claiming and dispatch. Conservatively retaining uncertainty may block an operation that never reached inference; that is preferable to automatically issuing a duplicate paid request without evidence. Also test the window after response receipt but before result persistence. A success that was not durably recorded may still need reconciliation.

Do not describe these controls as universal exactly-once inference. Local deduplication cannot prove what a remote service did during a network partition. Store a documented remote request identifier when available, but do not fabricate a status endpoint or assume that an arbitrary header guarantees deduplication.

Keep the user in control of another attempt

Frameboard can display the operation time, selected route, and uncertainty without exposing credentials or raw private prompts. A retry button should explain that another request may incur additional usage when the original outcome is unknown. Browser implementations should also consult the SDK documentation, keeping connection behavior separate from app operation state.

The useful deliverables are a decision table, deterministic fake tests, and durable-concurrency tests for the actual application. None requires paid inference. A real wallet-funded check remains a separate, specifically approved action, and its result should never be inferred from this fake's passing assertions.

For AI agents

Skill file

Download
---
name: aipass-unknown-outcome-tests
description: Use when AI actions duplicate or time out. Test operation claims and unknown paid outcomes with deterministic fake adapters.
---

# Duplicate action and unknown-outcome policy

## Scope
Design and test application dispatch policy without network inference. This skill cannot prove remote exactly-once processing, reverse charges, or authorize retries that may spend.

## Required discovery
Read the [canonical integration skill](https://aipass.one/skills/aipass-integration/SKILL.md) and verification reference. Exact OAuth and setup details remain delegated there. Read the selected method's [REST documentation](https://aipass.one/docs/rest); consult [SDK docs](https://aipass.one/docs/sdk) for browser behavior. Never transfer setup-control-plane idempotency guarantees to inference.

## Procedure
1. Identify every caller of the paid action: click handlers, form submission, reconnect logic, effects, queue retries, and agent invocation surfaces.
2. Define an operation identity scoped to host user and deliberate intent. Bind request fingerprint, route, and chosen model; reject changed content under the same identity.
3. Write a decision table for prepared, dispatch claimed, completed, proven pre-dispatch failure, and unknown post-dispatch outcome. Unknown must not automatically redispatch.
4. Implement deterministic fakes for success, rejection before dispatch, and lost response after recording dispatch. Use counters, barriers, or explicit scheduler controls rather than timing sleeps.
5. Test duplicate serial calls, concurrent claims, changed-payload conflicts, process restart, response loss, and crash before result persistence. Count adapter invocations, not only UI notifications.
6. Use the app's durable store for concurrency tests with a real uniqueness constraint or atomic claim. Label any in-memory example as serial policy only.
7. If documented remote reconciliation exists, plan its use explicitly. Otherwise preserve unknown status and explain that a deliberate new request may incur additional usage.

## Deliverables
- `operation-decision-table.md` with retry permissions and user-facing uncertainty.
- Runnable deterministic fake-adapter tests, including invocation counts.
- `dispatch-evidence.md` listing commands, durable-store coverage, and untested crash windows.

## Acceptance
Repeated identical intent never silently starts another operation in tested states. Unknown outcomes remain unknown; neither timeout nor disconnect is reported as proof of no charge. No real inference, wallet funding, or invented status endpoint is used.