Files
David SFandGitHub 3ba9e2f9a5 docs: capability pages for the unified docs site + README/doc parity gate (#329)
* docs: publish capability docs to the unified site + add README/doc parity gate

Every capability shipped only a README (kept for GitHub/PyPI). This adds a
parallel, cleaned-up page per capability under docs/ for the new unified docs
site (pydantic.dev/docs/harness), migrated from each README: snippets verified
runnable against source, autodoc API blocks, root-relative Pydantic AI links,
and an experimental-status admonition on the experimental set.

To keep README and doc in sync going forward, adds a docs-parity-reviewer agent
and a parity gate in the review checklist (run as the last step before merge),
plus the docs/ layout and the README<->doc requirement in AGENTS.md and the
capability-authoring guide.

* docs: fix README<->doc<->source inconsistencies across capabilities

A parity audit against source found drift, mostly in the capability READMEs
(staler than the migrated docs). All fixes verified against source:

- Correctness: the "approval/deferred tools are excluded from the sandbox" claim
  (code_mode README + doc) was false -- those tools are sandboxed like any
  other; corrected in both. The stale Shell persist_cwd sentinel description is
  replaced with the actual out-of-band temp-file capture. filesystem protected
  default `.git/` -> `.git/*` (the bare form never matched).
- Runnable snippets: added the missing imports/wiring so README snippets no
  longer raise NameError (subagents, context, planning, overflow, authoring,
  filesystem, code_mode).
- Parity: documented previously-undocumented params/behaviors (compaction
  strategy options, overflow strip_ansi/Passthrough, extra autodoc classes for
  context and subagents), fixed a stale version pin (>=1.95.1 -> >=2.1.0), and
  added the missing Managed Prompt row to the root README capability matrix.
- Style: normalized decorative Unicode to ASCII across all READMEs and dropped a
  hype phrase, matching AGENTS.md writing style and the docs.

* docs: add nav.json to drive the unified-docs harness sidebar

The unified docs mount the harness docs under /docs/ai/harness (fed live from
this repo via the pydantic-ai 'Pydantic AI Harness' section). This nav.json
defines the sub-nav (Overview + Capabilities + Experimental) and the set of doc
files the site includes.

* docs: migrate "What goes where?" explainer into harness overview

Adds the core-vs-harness boundary section (anchor #what-goes-where) to the
canonical harness overview, so the pydantic-ai docs that link to it can point
here after the duplicated in-repo stub is removed.

* docs: address CodeRabbit review -- runnable snippets, accuracy, multi-class autodoc

* docs: flatten harness nav and align with graduated capabilities

Following the experimental-graduation refactor (#347), restructure the
unified-docs harness pages:

- Flatten docs/ (drop capabilities/ and experimental/ subdirs); the sidebar
  is now Overview + one flat list per Douwe's request.
- Rename to match the graduated modules: overflow -> overflowing-tool-output,
  authoring -> runtime-authoring, docs -> pydantic-ai-docs.
- Drop the 'Experimental' admonitions from the graduated capabilities and
  repoint every import + ::: autodoc path off pydantic_ai_harness.experimental.
- Add docs for the newly-shipped capabilities: guardrails, dynamic-workflow,
  media, and acp (acp stays framed as experimental -- it may still be removed).
- Every capability doc now links to its source; index capability table lists
  the full set with flat links.

* docs: apply team-sync authoring rules + enforce them in CI

From the 2026-07-10 docs review on #329:

- Purpose-first leads: drop hook names (before_model_request,
  after_tool_execute) from the opening paragraphs of compaction and
  overflowing-tool-output (doc + README); mechanism moves lower.
- Mirror the soft 'API may change between releases' stability note from each
  graduated README into its doc page (ACP keeps its stronger experimental
  warning; guardrails' README has no note, so its page gets none).
- README H1s now use the capability's display name (Overflow capability ->
  Overflowing Tool Output, RuntimeAuthoring -> Runtime Authoring, SubAgents ->
  Subagents, etc.).
- Extend tests/test_docs_parity.py with per-page mechanical checks: source link
  present, heading matches the capability name, purpose-first lead (no hook in
  the opener), and no experimental framing on graduated pages (ACP excepted).
- Update the docs-parity-reviewer agent + review-checklist to the flat
  structure and the new semantic checks.

* docs: add the stability note to guardrails (parity with sibling capabilities)

guardrails was the one graduated capability whose README and doc page lacked
the shared 'API may change between releases' note. Add it to both.

* fix: restore uv.lock to match pyproject (bad text-merge dropped 8 lines)

Merging origin/main did a git text-merge of the generated uv.lock, leaving it
inconsistent with pyproject.toml -- every CI job failed at 'uv sync --locked'.
pyproject.toml is identical to main here, so the correct lock is main's.

* docs: address CodeRabbit review on #329

Findings that failed to post inline (GitHub error) but were real:
- context/README.md, planning/README.md: two nested examples still imported
  from pydantic_ai_harness.experimental.* -- repoint to the graduated modules.
- guardrails/README.md: replace em dashes with '--' (repo style) and add the
  source-module link.
- docs/media.md: standardize on the implementation's canonical media+sha256://
  URI scheme (was mixing media://).
- tests/test_docs_parity.py: strengthen my own checks per review --
  source-link and top-README-link now require a real Markdown link to the
  page's specific module (not a bare substring); heading checks assert an H1
  exists and equals the expected capability name via explicit page metadata.

* fix: restore uv.lock [options.exclude-newer-package] block

The lock lost its [options.exclude-newer-package] manifest (pydantic-ai-slim
= false, ...) -- a bad git text-merge dropped it, and diagnostic uv commands
rewrote it under a different local config. Without that block CI's
'uv sync --locked' re-resolves and fails ('addition of exclude newer exclusion
for pydantic-ai-slim'). Restore origin/main's exact lock.

* fix: restore uv.lock [options.exclude-newer-package] block

A pre-commit hook was rewriting uv.lock under the local uv config, stripping
the [options.exclude-newer-package] manifest (pydantic-ai-slim = false, ...).
Without it CI's 'uv sync --locked' re-resolves and fails. Commit origin/main's
exact lock with --no-verify so no hook mutates it (lock-only change).

* test: cover the docs-parity helper edge cases (100% coverage)

The strengthened helpers added defensive branches (missing frontmatter close,
fenced code before the lead, missing/forbidden/ClassName H1, lead running to
EOF) that no real doc exercises. Add direct unit tests so the file is back to
the repo's required 100% coverage.

* docs: link every capability README to its source module + enforce it

CodeRabbit re-flagged planning/README.md for a missing source link. Only
guardrails had one, so add the source-module link to all 15 remaining
capability READMEs (matching the doc pages) and add a parity test so the
requirement is mechanical and cannot silently regress.

* docs(agents): drop stale folder tree; fix flat docs path + guard names

AGENTS.md's File-structure ASCII tree and capability-authoring's doc paths
still showed docs/capabilities// docs/experimental/ (flattened in this PR) and
the old /docs/harness URL. Delete the tree rather than redraw it -- the layout
is discoverable by listing the repo; keep only the non-obvious conventions
(flat docs/, the README<->doc parity requirement). Also fix the Vocabulary
guard examples (InputGuard/OutputGuard, not the nonexistent InputGuardrail/
CostGuard).

* test: statically validate doc snippets exist and parse

Every Python snippet in the capability READMEs and docs/*.md pages is now
checked for the two failures a reader hits immediately: it does not parse
(syntax), or it imports a pydantic_ai_harness symbol that does not exist (stale
module path or renamed name -- the class of bug behind the experimental.* import
drift). Static only: no model/network execution, so it needs no mocking. The
four illustrative API-signature blocks opt out with a {test="skip"} fence
(read by pytest-examples, stripped-safe for the unified-docs render).

* test: don't fail doc-snippet check on a missing optional extra

The static check imported capability modules to resolve their symbols, but in
the slim CI job (no extras) importing e.g. pydantic_ai_harness.experimental.acp
raises ModuleNotFoundError for the absent third-party 'acp' package -- the
harness module exists, its extra just isn't installed. Distinguish a genuinely
missing harness module (fail) from a missing extra (skip) by the ImportError's
module name.
2026-07-13 11:00:12 -05:00
..

Shell

Give an agent the ability to run shell commands, with allow/deny controls and managed background processes.

Source

The problem

Agents frequently need to run a build, a test suite, a linter, or a quick grep. Wiring up subprocess handling -- streaming output, timeouts, truncation, killing runaway processes, and cleaning up background jobs at the end of a run -- is fiddly boilerplate that every agent reinvents.

The solution

Shell exposes command-execution tools rooted at a working directory, with configurable allow/deny lists and automatic cleanup of background processes when the agent run ends.

from pydantic_ai import Agent
from pydantic_ai_harness import Shell

agent = Agent(
    'anthropic:claude-sonnet-4-6',
    capabilities=[Shell(cwd='./workspace', allowed_commands=['ls', 'cat', 'rg'])],
)

result = agent.run_sync('List the Python files and summarize the largest one.')
print(result.output)

Tools

Tool Purpose
run_command Run a command synchronously and return labelled stdout/stderr plus exit code. Honors a per-call or default timeout.
start_command Launch a long-running command (server, watcher) in the background; returns an ID.
check_command Report the status and accumulated output of a background command.
stop_command Terminate a background command and return its final output.

Output is labelled with [stdout] / [stderr] markers and an [exit code: N] line on non-zero exit. When it exceeds max_output_chars the tail is kept (the head is dropped), so errors, stack traces, and the [stderr] section -- which all land at the end -- survive truncation.

Command controls

Field Effect
allowed_commands If non-empty, only these executables may run (allowlist).
denied_commands These executables are always rejected (denylist).
denied_operators Shell operators (e.g. >, >>, `
allow_interactive If False (default), commands that expect a TTY (vi, sudo, ssh, ...) are blocked.

allowed_commands and denied_commands are mutually exclusive -- set one, not both. denied_commands defaults to a list of destructive commands (rm, rmdir, mkfs, dd, format, shutdown, reboot, halt, poweroff, init); pass an empty list to disable. The executable name is extracted with shlex, so arguments don't bypass the check.

A denied or blocked command surfaces to the model as a ModelRetry (the model can retry with an allowed command) rather than aborting the run.

These checks are best-effort, not a security boundary. A sufficiently motivated agent can defeat them (e.g. bash -c '...', env-var indirection). For hard guarantees, run the agent inside OS-level isolation -- a container or sandbox.

Environment control

By default a spawned command inherits the agent process's full environment. In a sandbox that holds LLM API keys, tokens, or other secrets, a command the model writes can read them. Two fields control what the subprocess sees:

Field Effect
env Explicit environment that replaces inheritance entirely. The subprocess sees exactly these variables and nothing else.
denied_env_patterns Glob patterns (fnmatch) for variable names stripped from the base environment. Mirrors denied_commands.

env is a hard boundary for inherited environment variables: set it and inherited secrets cannot reach the subprocess at all (you supply PATH and anything else the command needs). denied_env_patterns is a denylist over the inherited environment -- lighter to configure when you only need to drop a few known-sensitive names. The two compose: when both are set, patterns also filter the explicit env. Leaving both unset preserves the inherit-everything default.

from pydantic_ai_harness import Shell
from pydantic_ai_harness.shell import LLM_API_KEY_ENV_PATTERNS

# Strip provider credentials from the inherited environment.
Shell(cwd='./repo', denied_env_patterns=LLM_API_KEY_ENV_PATTERNS)

# Or hand the subprocess a fixed environment, inheriting nothing.
import os
Shell(cwd='./repo', env={'PATH': os.environ['PATH'], 'HOME': os.environ['HOME']})

LLM_API_KEY_ENV_PATTERNS covers common provider prefixes (ANTHROPIC_*, OPENAI_*, OPENROUTER_*, GOOGLE_*, GEMINI_*, GATEWAY_*) plus PYDANTIC_AI_GATEWAY_API_KEY. It targets LLM credentials only -- it does not cover other host secrets (a LOGFIRE_TOKEN, a GitHub token, cloud credentials), and its prefixes are coarse, so GOOGLE_* also strips non-credential vars like GOOGLE_APPLICATION_CREDENTIALS. Treat it as a starting point and add your own patterns. It is not the default: stripping environment variables silently would break agents that rely on inherited credentials, so it is opt-in.

env is enforced at spawn, not applied as a post-hoc filter on a running process: the subprocess starts with exactly the resolved environment (your env, minus anything denied_env_patterns removes from it). That makes it a real boundary for inherited environment variables, unlike the best-effort command denylist. It is not a full security boundary: a command running under the same OS identity can still read host files -- use OS-level isolation for that. The flip side is that a pattern broad enough to strip PATH or HOME, or an env that omits them, can break command resolution. External commands may still run via the shell's built-in default PATH on some systems, but don't rely on it -- set PATH explicitly when you replace the environment.

Background processes

start_command writes stdout/stderr to temp files and returns a short ID. Use check_command(command_id) to poll and stop_command(command_id) to terminate and collect final output. Processes are launched in their own session (start_new_session) so the whole process group can be signalled -- SIGTERM, escalating to SIGKILL after a grace period.

On run end, the toolset's __aexit__ terminates every still-running background process and deletes its temp files. The agent runtime enters toolsets via an AsyncExitStack, so this cleanup runs whether the run succeeds or raises -- an agent that forgets to call stop_command won't leak processes.

Working directory

By default each command runs in cwd and cd has no lasting effect. Set persist_cwd=True to make cd sticky across calls: each command is wrapped so that after it runs, its final working directory is recorded to a private temp file, and that directory is carried into subsequent calls. The path is only updated when the command exits 0, and the record is written out-of-band (not to stdout) so command output can never spoof the tracked directory.

Configuration

Shell(
    cwd='.',                       # str | Path -- working directory
    allowed_commands=[],           # allowlist (mutually exclusive with denied)
    denied_commands=[...],         # denylist (defaults to destructive commands)
    denied_operators=[],           # blocked shell operators
    default_timeout=30.0,          # seconds, per run_command
    max_output_chars=50_000,       # output cap returned to the model
    persist_cwd=False,             # make cd sticky across calls
    allow_interactive=False,       # allow TTY-style commands
    env=None,                      # explicit env, replacing inheritance (None = inherit)
    denied_env_patterns=[],        # glob patterns stripped from the inherited env
)

Agent spec (YAML/JSON)

Shell works with Pydantic AI's agent spec:

# agent.yaml
model: anthropic:claude-sonnet-4-6
capabilities:
  - Shell:
      cwd: ./workspace
      allowed_commands: ['ls', 'cat', 'rg', 'pytest']
from pydantic_ai import Agent
from pydantic_ai_harness import Shell

agent = Agent.from_file('agent.yaml', custom_capability_types=[Shell])

Pass custom_capability_types so the spec loader knows how to instantiate Shell.

Further reading