Transcript Hygiene
Reference: provider-specific transcript sanitization and repair rules
This document describes provider-specific fixes applied to transcripts before a run
(building model context). These are in-memory adjustments used to satisfy strict
provider requirements. They do not rewrite stored JSONL transcript on disk.
Scope includes:
- Tool call id sanitization
- Tool result pairing repair
- Turn validation / ordering
- Thought signature cleanup
- Image payload sanitization
If you need transcript storage details, see:
Global rule: image sanitization
Image payloads are always sanitized to prevent provider-side rejection due to size
limits (downscale/recompress oversized base64 images).
Implementation:
- ''sanitizeSessionMessagesImages'' in ''src/agents/pi-embedded-helpers/images.ts''
- ''sanitizeContentBlocksImages'' in ''src/agents/tool-images.ts''
Historical behavior (pre-2026.1.22)
Before 2026.1.22 release, OpenClaw applied multiple layers of transcript hygiene:
- A transcript-sanitize extension ran on every context build and could:
- Repair tool use/result pairing.
- Sanitize tool call ids (including a non-strict mode that preserved ''_''/''-'').
- The runner also performed provider-specific sanitization, which duplicated work.
- Additional mutations occurred outside provider policy, including:
- Stripping ''<final>'' tags from assistant text before persistence.
- Dropping empty assistant error turns.
This complexity caused cross-provider regressions (notably ''openai-responses'' ''call_id|fc_id'' pairing). The 2026.1.22 cleanup removed extension, centralized logic in runner, and made OpenAI ''no-touch'' beyond image sanitization.