# A handoff is not done when the message is sent.

Record what the next worker received, which state it used, and what actually committed. A small receipt contract for durable agent workflows.

Published: 2026-10-09 | Updated: 2026-10-09
Author: /init editorial
Method: AI-assisted, source-checked
Canonical: https://init.news/notes/agent-handoff-receipts/

## The answer

Give every handoff an identifier, an explicit input snapshot, a receiving acknowledgement, and a terminal outcome. Keep execution attempts separate from the logical task. A trace explains where work happened; a durable receipt records what the workflow may safely rely on after a timeout, retry, or worker restart.

## Start at the uncertain boundary

Imagine a research worker handing a draft to a publishing worker. The sender times out after dispatch. Did the receiver never start, reject an obsolete draft, or publish successfully before its acknowledgement was lost? A transcript that ends with 'please publish' cannot choose the next safe action. This is a hypothetical example, but the ambiguity is concrete: the orchestrator needs state that survives the conversation.

Our proposed contract has four observable stages: dispatched, received, validated, and completed. Received means the worker acknowledged a particular input version. Validated means application checks passed. Completed means the promised effect has a verifiable outcome. Do not collapse these into one success flag, and do not call an acknowledgement proof that a model reasoned correctly.

Receipt field | Purpose | Example policy
--- | --- | ---
task_id + handoff_id | Identify the logical work and this transfer. | Keep task identity stable across retries; create a distinct handoff for a changed input contract.
attempt_id + input_version | Identify the execution and immutable input snapshot. | Record a new attempt on retry; reject an unexpected snapshot.
required_evidence + received_evidence | Make missing prerequisites visible. | Record source identifiers or content hashes, not sensitive source text in every log.
status + error_class + timestamps | Explain progress and the point of failure. | Separate rejected input, timeout, cancellation, and tool failure.
effect_key + outcome_reference | Reconcile an external write after uncertainty. | Use a stable idempotency key where supported and inspect the external outcome.
trace_id + parent_or_link | Connect the receipt to execution detail. | Keep the durable receipt independently of sampled telemetry.

## Use traces for navigation and receipts for recovery

OpenTelemetry describes spans as units of work and uses propagated context to correlate operations across services. Span links can connect asynchronous work across traces. Those primitives provide a useful navigation layer for the receipt identifiers above; they do not define your business completion rule. [1]

The OpenTelemetry GenAI conventions include agent invocation and tool-execution operations. Map your runtime's observable events onto the conventions version you actually implement, and retain a versioned adapter as the vocabulary changes. Keep payload capture deliberate: identifiers and validation outcomes often answer the recovery question without retaining full prompts or outputs. [2]

## A checkpoint does not make a side effect exactly once

LangGraph's Functional API documentation recommends placing side effects in tasks and designing them to be idempotent. It explains that completed task results can be replayed from checkpoints when a workflow resumes. [3] Our implementation recommendation is to pair that mechanism with the destination's write semantics. If an external service has no idempotency support, define a lookup or reconciliation step before issuing another write after an ambiguous timeout.

- Stop the sender immediately after dispatch. On recovery, find the existing handoff before creating more work.
- Stop the receiver after the external effect but before recording completion. Confirm that reconciliation finds the effect or safely identifies uncertainty.
- Deliver the same handoff twice. Confirm that the result is one logical outcome, or that duplicate attempts are surfaced for repair.
- Change the input version while a worker is queued. Confirm that the stale attempt cannot silently publish a newer task's result.
- Remove a required evidence record. Confirm that the worker declines to advance and records the missing prerequisite.

## Decide what deserves durable storage

Retain receipts for operations whose repetition or omission matters. A cheap, reversible read may need only ordinary tracing; a publication, payment, or state-changing tool call needs a clear recovery contract. Choose retention and access policies by data sensitivity and operational need. This article proposes a design and test plan; it does not claim to have run these failures or prove that any framework eliminates duplicate effects.

## Sources

- [OpenTelemetry: Traces](https://opentelemetry.io/docs/concepts/signals/traces/) — [1] Span units of work, context propagation, events, and links for asynchronous causality.; checked 2026-10-09
- [OpenTelemetry GenAI semantic conventions: Agent spans](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-agent-spans.md) — [2] Agent invocation and tool-execution operation vocabulary; a moving main-branch document inspected on 9 October 2026.; checked 2026-10-09
- [LangGraph: Functional API overview](https://docs.langchain.com/oss/python/langgraph/functional-api) — [3] Task boundaries, checkpoint replay, and idempotent side-effect recommendations.; checked 2026-10-09
