* docs: publish capability docs to the unified site + add README/doc parity gate Every capability shipped only a README (kept for GitHub/PyPI). This adds a parallel, cleaned-up page per capability under docs/ for the new unified docs site (pydantic.dev/docs/harness), migrated from each README: snippets verified runnable against source, autodoc API blocks, root-relative Pydantic AI links, and an experimental-status admonition on the experimental set. To keep README and doc in sync going forward, adds a docs-parity-reviewer agent and a parity gate in the review checklist (run as the last step before merge), plus the docs/ layout and the README<->doc requirement in AGENTS.md and the capability-authoring guide. * docs: fix README<->doc<->source inconsistencies across capabilities A parity audit against source found drift, mostly in the capability READMEs (staler than the migrated docs). All fixes verified against source: - Correctness: the "approval/deferred tools are excluded from the sandbox" claim (code_mode README + doc) was false -- those tools are sandboxed like any other; corrected in both. The stale Shell persist_cwd sentinel description is replaced with the actual out-of-band temp-file capture. filesystem protected default `.git/` -> `.git/*` (the bare form never matched). - Runnable snippets: added the missing imports/wiring so README snippets no longer raise NameError (subagents, context, planning, overflow, authoring, filesystem, code_mode). - Parity: documented previously-undocumented params/behaviors (compaction strategy options, overflow strip_ansi/Passthrough, extra autodoc classes for context and subagents), fixed a stale version pin (>=1.95.1 -> >=2.1.0), and added the missing Managed Prompt row to the root README capability matrix. - Style: normalized decorative Unicode to ASCII across all READMEs and dropped a hype phrase, matching AGENTS.md writing style and the docs. * docs: add nav.json to drive the unified-docs harness sidebar The unified docs mount the harness docs under /docs/ai/harness (fed live from this repo via the pydantic-ai 'Pydantic AI Harness' section). This nav.json defines the sub-nav (Overview + Capabilities + Experimental) and the set of doc files the site includes. * docs: migrate "What goes where?" explainer into harness overview Adds the core-vs-harness boundary section (anchor #what-goes-where) to the canonical harness overview, so the pydantic-ai docs that link to it can point here after the duplicated in-repo stub is removed. * docs: address CodeRabbit review -- runnable snippets, accuracy, multi-class autodoc * docs: flatten harness nav and align with graduated capabilities Following the experimental-graduation refactor (#347), restructure the unified-docs harness pages: - Flatten docs/ (drop capabilities/ and experimental/ subdirs); the sidebar is now Overview + one flat list per Douwe's request. - Rename to match the graduated modules: overflow -> overflowing-tool-output, authoring -> runtime-authoring, docs -> pydantic-ai-docs. - Drop the 'Experimental' admonitions from the graduated capabilities and repoint every import + ::: autodoc path off pydantic_ai_harness.experimental. - Add docs for the newly-shipped capabilities: guardrails, dynamic-workflow, media, and acp (acp stays framed as experimental -- it may still be removed). - Every capability doc now links to its source; index capability table lists the full set with flat links. * docs: apply team-sync authoring rules + enforce them in CI From the 2026-07-10 docs review on #329: - Purpose-first leads: drop hook names (before_model_request, after_tool_execute) from the opening paragraphs of compaction and overflowing-tool-output (doc + README); mechanism moves lower. - Mirror the soft 'API may change between releases' stability note from each graduated README into its doc page (ACP keeps its stronger experimental warning; guardrails' README has no note, so its page gets none). - README H1s now use the capability's display name (Overflow capability -> Overflowing Tool Output, RuntimeAuthoring -> Runtime Authoring, SubAgents -> Subagents, etc.). - Extend tests/test_docs_parity.py with per-page mechanical checks: source link present, heading matches the capability name, purpose-first lead (no hook in the opener), and no experimental framing on graduated pages (ACP excepted). - Update the docs-parity-reviewer agent + review-checklist to the flat structure and the new semantic checks. * docs: add the stability note to guardrails (parity with sibling capabilities) guardrails was the one graduated capability whose README and doc page lacked the shared 'API may change between releases' note. Add it to both. * fix: restore uv.lock to match pyproject (bad text-merge dropped 8 lines) Merging origin/main did a git text-merge of the generated uv.lock, leaving it inconsistent with pyproject.toml -- every CI job failed at 'uv sync --locked'. pyproject.toml is identical to main here, so the correct lock is main's. * docs: address CodeRabbit review on #329 Findings that failed to post inline (GitHub error) but were real: - context/README.md, planning/README.md: two nested examples still imported from pydantic_ai_harness.experimental.* -- repoint to the graduated modules. - guardrails/README.md: replace em dashes with '--' (repo style) and add the source-module link. - docs/media.md: standardize on the implementation's canonical media+sha256:// URI scheme (was mixing media://). - tests/test_docs_parity.py: strengthen my own checks per review -- source-link and top-README-link now require a real Markdown link to the page's specific module (not a bare substring); heading checks assert an H1 exists and equals the expected capability name via explicit page metadata. * fix: restore uv.lock [options.exclude-newer-package] block The lock lost its [options.exclude-newer-package] manifest (pydantic-ai-slim = false, ...) -- a bad git text-merge dropped it, and diagnostic uv commands rewrote it under a different local config. Without that block CI's 'uv sync --locked' re-resolves and fails ('addition of exclude newer exclusion for pydantic-ai-slim'). Restore origin/main's exact lock. * fix: restore uv.lock [options.exclude-newer-package] block A pre-commit hook was rewriting uv.lock under the local uv config, stripping the [options.exclude-newer-package] manifest (pydantic-ai-slim = false, ...). Without it CI's 'uv sync --locked' re-resolves and fails. Commit origin/main's exact lock with --no-verify so no hook mutates it (lock-only change). * test: cover the docs-parity helper edge cases (100% coverage) The strengthened helpers added defensive branches (missing frontmatter close, fenced code before the lead, missing/forbidden/ClassName H1, lead running to EOF) that no real doc exercises. Add direct unit tests so the file is back to the repo's required 100% coverage. * docs: link every capability README to its source module + enforce it CodeRabbit re-flagged planning/README.md for a missing source link. Only guardrails had one, so add the source-module link to all 15 remaining capability READMEs (matching the doc pages) and add a parity test so the requirement is mechanical and cannot silently regress. * docs(agents): drop stale folder tree; fix flat docs path + guard names AGENTS.md's File-structure ASCII tree and capability-authoring's doc paths still showed docs/capabilities// docs/experimental/ (flattened in this PR) and the old /docs/harness URL. Delete the tree rather than redraw it -- the layout is discoverable by listing the repo; keep only the non-obvious conventions (flat docs/, the README<->doc parity requirement). Also fix the Vocabulary guard examples (InputGuard/OutputGuard, not the nonexistent InputGuardrail/ CostGuard). * test: statically validate doc snippets exist and parse Every Python snippet in the capability READMEs and docs/*.md pages is now checked for the two failures a reader hits immediately: it does not parse (syntax), or it imports a pydantic_ai_harness symbol that does not exist (stale module path or renamed name -- the class of bug behind the experimental.* import drift). Static only: no model/network execution, so it needs no mocking. The four illustrative API-signature blocks opt out with a {test="skip"} fence (read by pytest-examples, stripped-safe for the unified-docs render). * test: don't fail doc-snippet check on a missing optional extra The static check imported capability modules to resolve their symbols, but in the slim CI job (no extras) importing e.g. pydantic_ai_harness.experimental.acp raises ModuleNotFoundError for the absent third-party 'acp' package -- the harness module exists, its extra just isn't installed. Distinguish a genuinely missing harness module (fail) from a missing extra (skip) by the ImportError's module name.
Logfire-backed capabilities
Drive agent configuration from Logfire managed variables, so you can iterate on it from the Logfire UI -- versioned, labelled, and rolled out -- without redeploying.
Install the extra:
pip install 'pydantic-ai-harness[logfire]'
ManagedPrompt
Back an agent's instructions with a Logfire-managed Prompt.
A broader, first-party
Managedcapability is in flight in pydantic-ai#5107 and will eventually be importable aspydantic_ai.managed.logfire.Managed-- covering instructions, model settings, and whole-spec variables. Until then,ManagedPromptis the supported path for backing instructions with a Logfire-managed prompt.
The problem
Prompts are critical to agent behavior, but iterating on them through the normal edit -> review -> deploy loop is slow, and you can't easily A/B test a change or roll it back the moment it misbehaves in production.
The solution
ManagedPrompt declares the backing managed variable for you and resolves it once per
run, feeding the value into the agent's instructions. The resolution happens inside the
run's wrap_run hook using the
ResolvedVariable
as a context manager that stays open for the whole run -- so the selected label and version
are attached as baggage to every child span of the agent run. You get a direct correlation
between a run's behavior and the exact prompt version that produced it, plus instant
iteration and rollback from the Logfire UI.
Usage
Pass the prompt name and a default value. The name support_agent is declared as the managed
variable prompt__support_agent -- the naming Logfire's Prompt management uses (hyphens in a
name become underscores). The default keeps the agent working until a remote value is published.
import logfire
from pydantic_ai import Agent
from pydantic_ai_harness.logfire import ManagedPrompt
logfire.configure()
agent = Agent(
'openai:gpt-5',
capabilities=[
ManagedPrompt(
'support_agent',
default='You are a helpful customer support agent. Be friendly and concise.',
label='production',
)
],
)
result = agent.run_sync('My order never arrived.')
print(result.output)
Targeting
For deterministic A/B assignment (the same user always sees the same label), pass a
targeting_key. It can be a static string or a callable that derives the key from the
RunContext -- handy
when the key lives in your agent's deps:
from dataclasses import dataclass
from pydantic_ai import Agent
from pydantic_ai_harness.logfire import ManagedPrompt
@dataclass
class Deps:
user_id: str
agent = Agent(
'openai:gpt-5',
deps_type=Deps,
capabilities=[
ManagedPrompt(
'support_agent',
default='You are a helpful customer support agent.',
targeting_key=lambda ctx: ctx.deps.user_id,
),
],
)
Pass attributes (or a callable returning them) for condition-based targeting rules.
When label is omitted, the variable's rollout and targeting rules pick the label;
when both targeting_key and attributes are omitted, Logfire falls back to its own
targeting context and then to the active trace id.
Templating with deps
By default the resolved prompt is used verbatim. Pass render_template=True to render it as a
Handlebars template against the agent's deps -- the same mechanism as
TemplateStr -- so {{field}} is filled
from deps:
from dataclasses import dataclass
from pydantic_ai import Agent
from pydantic_ai_harness.logfire import ManagedPrompt
@dataclass
class Deps:
customer_name: str
agent = Agent(
'openai:gpt-5',
deps_type=Deps,
capabilities=[
ManagedPrompt(
'support_agent',
default='You are helping {{customer_name}}. Be friendly and concise.',
render_template=True,
),
],
)
Rendering requires pydantic-handlebars (install pydantic-ai-slim[spec]). It is off by default.
Prompt-cache trade-off
The resolved value lands in the agent's system instructions. Provider prompt caches (Anthropic,
OpenAI, etc.) key strictly by prefix -- tools -> system -> messages -- so any change to the system
block invalidates the cached prefix for the affected runs.
| Mode | Cache impact |
|---|---|
Pinned label='production', no rollout split |
Cache-stable. The value only changes on a deliberate prompt rollout, which is the same cost as a redeploy. |
Percentage rollout across labels (no label=) |
Different runs land on different labels -> splits the cache into one lane per label. |
targeting_key per user/tenant with multiple labels in play |
Cache lanes per assigned label; deterministic per key but still N lanes overall. |
| Mid-traffic label flip in the Logfire UI | One-shot cold-invalidation for everyone on that label. |
In short: pinning a label keeps the cache hot; using ManagedPrompt as an A/B platform is opt-in
cache cost. If you don't need rollouts, label='production' is the recommended default.
Using your own variable
Declaring the same name more than once is fine -- each ManagedPrompt builds its own backing
variable, so sharing a prompt across several agents just works. Pass an existing
logfire.variables.Variable
as the first argument instead of a name when you want to declare the variable yourself --
for example a template_var, or one registered for variables_push:
import logfire
from pydantic_ai import Agent
from pydantic_ai_harness.logfire import ManagedPrompt
logfire.configure()
support_prompt = logfire.var(
name='prompt__support_agent',
type=str,
default='You are a helpful customer support agent. Be friendly and concise.',
)
agent = Agent('openai:gpt-5', capabilities=[ManagedPrompt(support_prompt, label='production')])
When name is a prompt name, pass logfire_instance= to declare the variable on a specific
Logfire instance instead of the module-level default.
Notes
- The prompt resolves to a
str. By default it's used verbatim; setrender_template=Trueto render{{...}}againstdeps(see Templating with deps). - Resolution is isolated per run via a context variable, so a single capability instance is safe to share across concurrent runs.
ManagedPrompt.resolvedexposes the active run'sResolvedVariable(value, label, version, reason) for inspection -- e.g. from inside a tool.- The capability runs outermost (wrapping
Instrumentation) so the resolved variable's baggage covers the agent run span as well as its children. On recent Logfire versions both the selected label and the version are propagated as separate baggage attributes. - Resolution happens once per run. A label flip or rollout change that lands in Logfire mid-run is not picked up until the next run starts -- the trade-off for run-stable instructions and a single baggage scope across all child spans.
- For Logfire-side targeting that lives outside the agent (e.g. set once per request handler),
use Logfire's
targeting_contextin an outer scope;ManagedPromptonly needstargeting_key/attributeswhen the key comes from the agent'sRunContext.