Temporary review material, drop this commit before merge: diff fixtures with a patch of uncommitted edits, review notes and screenshots, and a local color tuner that writes overrides into the diff view through diff-color-tuning.ts.
Use v2 green and red tokens for changed rows, gutters, inline highlights, change bars and line numbers in light and dark modes. Changed-row fills are final translucent colors instead of being re-mixed by Pierre, emphasized text overrides syntax colors so it stays legible on its highlight, and unchanged line numbers and fold labels use the faint text token.
Remove the attention.enabled master switch. Notifications and sounds are
now enabled independently and both default to off. Existing enabled keys
are ignored; there is no migration.
Add project and worktree navigation with restored search and selection, workspace-preserving targets, and optional worktree naming. Keep creation in the Ctrl+N footer and defer filesystem browsing.
Move files deployments to Wrangler, make CLI and desktop own their publishing destinations, add direct-download update metadata and desktop feeds, and refresh installation docs.
Reuse the shared dependency-aware plugin loader for Core hot reloads. Scope watcher setup callbacks and subscriptions, recover missing external helpers, and preserve explicit local entrypoint invalidation without reloading package dependencies. Add regression coverage and portable RPC fixtures.
Reload statically discovered local plugin dependencies through native module caches while preserving shared package identities and best-effort last-good registrations.
Show grouped server plugin failure notices with an Open plugins action, retain the home failure count, and reveal failed built-ins in the plugin dialog. Keep unchanged failures quiet across inventory refreshes and reconnects.
Carry unfinished model work across consecutive Location handoffs without resetting the logical step allowance. Keep idle moves and queued prompt admission unchanged. Cover steered and queued second moves, preserved tool history, and durable event ordering.
Capture child stdout and stderr eagerly before lazy Effect readers attach. Preserve the bounded post-exit drain and process cleanup policies, with delayed-consumption and backpressure regressions.
Update the Core compile options and Server replacement values to the current LayerNode API. Preserve test expectations, replacement targets, and layer lifetimes.
Replace exposed layer graph assembly with opaque declarations, checked substitutions, and lifetime-aware compilation. Preserve deep replacement, ordered startup, and Effect-owned resource lifetimes; migrate callers and verify source and published package contracts.
Keep the title and workspace label above scrollable sidebar details. Disable the unused horizontal scrollbar and place the automatic vertical scrollbar in the reserved gutter so tab changes do not flash or shift the sidebar.
Configure custom Markdown renderers before assigning content and share a reactive message-position index across assistant footers. Preserve completion ordering and historical footer metrics, with regression coverage for prepend, same-length refresh, and revert.
Move standalone skill activation into ID-bound Session handles and delegate from the public service. Preserve current-placement lookup, raw skill content, ambient publication context, validation order, and host-scoped detached resume behavior. Cover ownership and lifecycle contracts with focused regressions.
Move existing Session list and message queries into SessionStore. Preserve public response wrapping, Session existence checks, pagination, ordering, and typed message decoding errors.
Wait for the diff-base search input to receive focus before typing. Preserve existing assertions and timeouts while removing the render-versus-focus test race.
Separate ID-bound Session policy from host routing. Bind Inbox and Location preparation dependencies at construction, preserve admission and execution semantics, and cover the extracted ownership contracts directly.
Reuse a file-local fixture for renderer, HTTP server, event stream, app startup, and teardown across lifecycle tests. Preserve scenario-specific handlers, configuration, deferred responses, and assertions while removing 187 lines of repeated setup.
Name the tool-owned pre-spawn preparation boundary while preserving hook edits, permission ordering, directory validation, and effective timeout reporting. Strengthen the existing regression assertions.
Run user shell commands concurrently with model execution, retain their output in one shell entry, and admit non-waking completion messages. Share result capture and notification formatting across shell callers.
Open the recent-session and project picker synchronously with selectable cached rows and independent refreshes. Reconcile committed moves and deletions without restoring stale rows, preserve dismissal and selection through delayed reads, and keep filtered selections visible after asynchronous results arrive.
Buffer bulk history privately and publish once with a bounded head window. Preserve pending anchor compensation while ensuring newer navigation cancels obsolete Home work. Add atomic-loading and real-App navigation regressions.
Render queued compaction feedback before model setup, coalesce repeated gestures, and reconcile canonical admissions without restoring consumed rows. Serialize prompt and compaction preparation through the existing per-session admission chain.
const commentAge = now - new Date(complianceComment.created_at).getTime();
if (commentAge < twoHours) {
core.info(`${kind} #${item.number} still within 2-hour window (${Math.round(commentAge / 60000)}m elapsed)`);
if (commentAge < seventyTwoHours) {
core.info(`${kind} #${item.number} still within 72-hour window (${Math.round(commentAge / 60000)}m elapsed)`);
continue;
}
const closeMessage = isPR
? 'This pull request has been automatically closed because it was not updated to meet our [contributing guidelines](../blob/dev/CONTRIBUTING.md) within the 2-hour window.\n\nFeel free to open a new pull request that follows our guidelines.'
: 'This issue has been automatically closed because it was not updated to meet our [contributing guidelines](../blob/dev/CONTRIBUTING.md) within the 2-hour window.\n\nFeel free to open a new issue that follows our issue templates.';
? 'This pull request has been automatically closed because it was not updated to meet our [contributing guidelines](../blob/dev/CONTRIBUTING.md) within the 72-hour window.\n\nFeel free to open a new pull request that follows our guidelines.'
: 'This issue has been automatically closed because it was not updated to meet our [contributing guidelines](../blob/dev/CONTRIBUTING.md) within the 72-hour window.\n\nFeel free to open a new issue that follows our issue templates.';
await github.rest.issues.createComment({
owner: context.repo.owner,
@@ -129,5 +129,5 @@ jobs:
});
}
core.info(`Closed non-compliant ${kind} #${item.number} after 2-hour window`);
core.info(`Closed non-compliant ${kind} #${item.number} after 72-hour window`);
If the issue is NOT compliant and the author association is not OWNER or MEMBER, start the comment with:
<!-- issue-compliance -->
Then explain what needs to be fixed and that they have 2 hours to edit the issue before it is automatically closed. Also add the label needs:compliance to the issue using: gh issue edit ${{ github.event.issue.number }} --add-label needs:compliance
Then explain what needs to be fixed and that they have 72 hours to edit the issue before it is automatically closed. Also add the label needs:compliance to the issue using: gh issue edit ${{ github.event.issue.number }} --add-label needs:compliance
If duplicates were found, include a section about potential duplicates with links.
@@ -124,7 +124,7 @@ jobs:
**What needs to be fixed:**
- [specific reasons]
Please edit this issue to address the above within **2 hours**, or it will be automatically closed.
Please edit this issue to address the above within **72 hours**, or it will be automatically closed.
- After changing the public Protocol or Server `HttpApi`, run `bun run generate` from `packages/client`. Do not edit generated client files directly.
- Keep runtime dependencies directed from Schema to Core and Protocol, then from Core and Protocol to Server. Client runtime code may depend on Schema and Protocol but never Core or Server; `sdk` composes Client, Core, and Server.
- Current implementation changes belong in `packages/core`, `packages/cli`, `packages/server`, `packages/protocol`, `packages/schema`, and related generated client surfaces when required.
- This repository does not use Changesets. Do not add `.changeset` files; follow the existing release workflow instead.
- The default branch in this repo is `v2`.
-Base all new branches and worktrees on`v2`, or `origin/v2` when the local `v2` ref is unavailable. Do not base them on `dev`.
-Default new branches and worktrees to `v2`, or `origin/v2` when the local `v2` ref is unavailable, and default pull requests to target `v2`. Use another base or target branch when the requester explicitly instructs it.
- Local `main` ref may not exist; use `v2` or `origin/v2` for diffs.
## Live V2 TUI Testing
- Run `bun run dev:live` from a development worktree to test its TUI against the currently elected `opencode2` background server and live sessions.
- Run `bun run dev:live` from a development worktree to test its TUI against the currently elected `opencode` background server and live sessions.
- Pass a directory after the script when needed, for example `bun run dev:live /path/to/project`.
- The script discovers the server with `opencode2 service status`, injects its private local credential from `opencode2 service get password`, and uses the `dev` TUI storage channel so tabs and other client-local state match the installed client.
- The script discovers the server with `opencode service status`, injects its private local credential from `opencode service get password`, and uses the `dev` TUI storage channel so tabs and other client-local state match the installed client.
- Prefer `dev:live` over plain `bun run dev` for this workflow. An implicit managed-service connection may replace the live server when the worktree client version differs; explicit `--server` warns and continues without replacing it.
- Keep things in one function unless composable or reusable
- Validate unknown values once at the boundary that owns them. Pass typed values inward instead of repeating `typeof value === "object"` and property-existence checks. Do not defensively revalidate values already guaranteed by a schema, constructor, or internal type.
- Do not extract single-use helpers preemptively. Inline the logic at the call site unless the helper is reused, hides a genuinely complex boundary, or has a clear independent name that improves the caller.
- Before adding complexity for a speculative or vanishingly unlikely race or security edge case, explain the concrete failure mode, likelihood, and complexity cost to the user and get their buy-in. Do not silently expand scope for theoretical robustness.
- Avoid `try`/`catch` where possible
@@ -82,9 +84,9 @@ const { a, b } = obj
### Imports
- Never alias imports. Do not use `import { foo as bar } from "..."` or renamed imports like `resolve as pathResolve`.
- Never use type-position `import("...")` references such as `Schema.declare<import("@opencode-ai/plugin/effect/plugin").Plugin["effect"]>`. Only when two imports genuinely collide on a name and no other option exists, an aliased type import (`import type { Plugin as PluginDefinition } from "..."`) is permitted as a last resort — still strongly preferred not to.
- Never use type-position `import("...")` references such as `Schema.declare<import("@opencode/plugin/effect/plugin").Plugin["effect"]>`. Only when two imports genuinely collide on a name and no other option exists, an aliased type import (`import type { Plugin as PluginDefinition } from "..."`) is permitted as a last resort — still strongly preferred not to.
- Never use star imports. Do not use `import * as Foo from "..."` or `import type * as Foo from "..."`.
- If a namespace-style value is needed, import the module's own exported namespace by name, for example `import { Project } from "@opencode-ai/core/project"`, then reference `Project.ID`.
- If a namespace-style value is needed, import the module's own exported namespace by name, for example `import { Project } from "@opencode/core/project"`, then reference `Project.ID`.
- Prefer dynamic imports for heavy modules that are only needed in selected code paths, especially in startup-sensitive entrypoints. Destructure dynamic import bindings near the top of the narrowest scope that needs them so they read like normal imports. Avoid inline chains such as `await import("./module").then((mod) => mod.value())` or `(await import("./module")).value()`. Keep branch-specific imports inside the branch that needs them to preserve lazy loading.
- Keep `SessionRunner`, model resolution, tool registry, permissions, and filesystem Location-scoped. Omitted `Location.workspaceID` means implicit-local placement; explicit workspace identity remains reserved for future placement semantics.
- Preserve one explicit `llm.stream(request)` call per Physical Attempt and reload projected history before durable continuation. A logical Step may use generic pre-output retries, one full-context retry after continuation rejection, incomplete-stream continuation, or one overflow-compaction rebuild. Generic retries retain the logical step number and do not consume another agent-step allowance. Do not delegate orchestration to an in-memory tool loop.
- Keep local Session drains process-local until clustering is implemented. `SessionRunCoordinator` joins explicit same-Session resumes, coalesces prompt wakeups, and allows different Sessions to run concurrently. A write-ahead execution claim marks a process-local busy period for restart recovery: terminal completion, failure, or user interruption releases it, while shutdown interruption and process death preserve it. Startup recovery resumes claimed top-level Sessions with durable per-execution attempt accounting. The claim is a recovery marker, not clustered ownership, fencing, or an exactly-once guarantee.
- Keep delivery vocabulary explicit. Prompts steer by default. Steers deliver in enqueue order at safe step boundaries, stopping before compaction or move control items. At an idle boundary, steers take priority; otherwise exactly one queued item delivers before the runner reevaluates continuation. Inbox items may be cancelled or changed between queue and steer before delivery. Promoting new user input resets the selected agent's step allowance; a batch of steers resets it once.
- Keep provider-specific native compaction mechanisms in `@opencode/ai` behind `LLMClient.compact`. `SessionCompaction` chooses a summary or native compaction from the model's `compaction` setting and owns route provenance, request shrinking, the retry policy, interruption, usage accounting, and checkpoint persistence.
- Keep delivery vocabulary explicit. Prompts steer by default. At safe step boundaries, steered compaction takes priority up to the first steered move control; other steers retain enqueue order. At an idle boundary, steers take priority; otherwise exactly one queued item delivers before the runner reevaluates continuation. Inbox items may be cancelled or changed between queue and steer before delivery. Promoting new user input resets the selected agent's step allowance; a batch of steers resets it once.
- One step is one logical LLM call; its durable record covers only the model-visible span. Do not write "provider turn", and do not use bare "turn" for a single call: "turn" is reserved for the future assistant-turn unit containing all steps from prompt promotion until the session would go idle.
- Keep event replay ownership separate from clustered Session execution ownership.
- Keep the Instructions algebra and built-ins in `src/instructions`; keep instruction producers with their observed domains, and keep Session History selection plus `InstructionState` and `InstructionEntry` persistence Session-owned. `InstructionDiscovery` observes ambient global and upward-project instructions. The runner composes built-ins, discovery, guidance, and entries explicitly in `loadInstructions`; there is no instruction registry.
Review endpoints in document order. For each endpoint, select one disposition and capture rationale or follow-up work in Notes. Mark **Reviewed** only after the disposition is agreed.
### Review criteria
- Resource and operation naming
- HTTP method and idempotency
- Request parameters and location scope
- Response shape and error taxonomy
- Authentication and authorization
- Current production consumers
- Stability level: public, experimental, or internal
- Whether the generated client API is intuitive
### Disposition legend
- **Keep:** ship unchanged as a supported V2 API
- **Change:** retain after a defined contract change
- **Remove:** exclude from the official V2 API
- **Experimental-only:** retain outside the stable API commitment
## Progress
- [x] Group 1: Foundation and placement (4)
- [x] Group 2: Configuration and capability catalogs (16)
- [x] Group 3: Credentials, integrations, MCP, and web search (22)
- [x] Group 4: Session lifecycle (12)
- [x] Group 5: Session execution and inputs (11)
- [x] Group 6: Session history and recovery (13)
- [x] Group 7: Inbox, permissions, and forms (19)
- [x] Group 8: Filesystem, worktrees, and VCS (12)
- [x] Group 9: PTYs, persistent terminals, and shells (24)
- [x] Group 10: Events, RPC, and experimental operations (6)
## Resolved during audit
### [x] `POST /api/plugin/await-activation`
- **Decision:** Remove
- **Notes:** Activation timing is an internal server concern. Catalog reads remain non-blocking.
- **Notes:** Full project metadata remains available from `GET /api/location`; no consumers used it from wrapped responses.
### [x] `GET /api/health` and `GET /api/server`
- **Decision:** Merge and rename
- **Replacement:** `GET /api/info` with operation ID `server.info`.
- **Notes:** Returns `version`, `pid`, and connection `urls`; readiness is conveyed by HTTP status.
### [x] `GET /api/project/current`
- **Decision:** Remove
- **Replacement:** `GET /api/location`, using `project` from the response.
- **Notes:** The endpoint duplicated `Location.Info.project`; production callers were migrated.
### [x] `POST /api/workspace` and `DELETE /api/workspace/{workspaceID}`
- **Decision:** Remove
- **Notes:** Provider-backed workspaces are not part of the V2 HTTP contract and can be introduced later. Core and the embedded SDK retain internal workspace support.
| [x] 063 | `POST` | `/api/session/{sessionID}/shell` | `session.shell` | Change | Caller ID is now the optimistic shell message ID; server derives its event ID. |
| [x] 090 | `GET` | `/api/session/{sessionID}/form/{formID}` | `session.form.get` | Change | Form definition and lifecycle state are now returned together. |
| [x] 091 | — | — | — | Remove | State is included by `session.form.get`. |
exportconstlongLine="Start unchanged | The old checkout flow waits for manual confirmation before showing the receipt | Middle unchanged with punctuation: brackets [one, two], braces {three}, quotes 'four', slash /five/ | The old final instruction asks the reviewer to close the window | End unchanged"
Unchanged prose should remain quiet, not compete with changed rows.
Replace one contiguous phrase: the original checkout experience stays here.
Several little edits: red apple, cold tea, slow train.
One character: ticket 7.
before — the middle and the end remain unchanged.
The start remains — before — the end remains.
The start and the middle remain — before
Punctuation only: hello, world!
Whitespace only: one two three
Trailing whitespace only.
Delete this sentence without a replacement.
This unchanged sentence separates a deletion from an addition.
Long prose: The original checkout flow asks the reader to review the old confirmation message before proceeding through the receipt screen, while the unchanged middle of this deliberately long paragraph checks whether horizontal scrolling or line wrapping keeps small inline edits visible at a narrow viewport; the final old phrase appears near the far right edge.
**Syntax-like Markdown** competes with `inline code`, [links](https://example.com/old), and _emphasis_.
Inventory for this worktree's V2 web/desktop file viewer. Source inspection only; no colors or behavior changed. Includes diff-specific colors, syntax colors, and the selection/search/comment colors used inside the viewer—not every unrelated global app variable. Unified and split use the same color system.
## Main controls and their wiring
| Visual role | Renderer variable | OpenCode source |
|---|---|---|
| Neutral file background | `--diffs-bg` | `--opencode-diffs-bg`, falling back to `--color-background-stronger` (alias of `--background-stronger`) |
| Default text | `--fg` / `--diffs-fg` | Registered theme foreground: `--text-base`; syntax spans override it |
| Search match / current match | CSS highlight backgrounds | Alpha of `--surface-warning-base` / `--surface-warning-strong` |
**Important:**`--surface-diff-add-*`, `--surface-diff-delete-*`, and `--surface-diff-hidden-*` exist in the global theme, but are not the direct row/inline/fold controls in the current Pierre rendering path. The older/custom separator CSS in `components/file.css` does use `--surface-diff-hidden-base` and `--surface-diff-hidden-strong`.
Each name below is a CSS custom property. Theme JSON override keys omit the leading `--`. Tailwind's `--color-…` aliases are defined in `packages/ui/src/styles/tailwind/colors.css` (for example `--color-surface-diff-add-base` → `--surface-diff-add-base`); these are aliases, not separate color decisions.
### Surfaces
```text
--surface-diff-unchanged-base
--surface-diff-skip-base
--surface-diff-hidden-base
--surface-diff-hidden-weak
--surface-diff-hidden-weaker
--surface-diff-hidden-strong
--surface-diff-hidden-stronger
--surface-diff-add-base
--surface-diff-add-weak
--surface-diff-add-weaker
--surface-diff-add-strong
--surface-diff-add-stronger
--surface-diff-delete-base
--surface-diff-delete-weak
--surface-diff-delete-weaker
--surface-diff-delete-strong
--surface-diff-delete-stronger
```
### Text, icons, and highlighter diff seeds
```text
--text-diff-add-base
--text-diff-add-strong
--text-diff-delete-base
--text-diff-delete-strong
--icon-diff-add-base
--icon-diff-add-hover
--icon-diff-add-active
--icon-diff-delete-base
--icon-diff-delete-hover
--icon-diff-modified-base
--syntax-diff-add
--syntax-diff-delete
--syntax-diff-unknown
```
The built-in OpenCode theme maps `--syntax-diff-add` and `--text-diff-add-base` to `--v2-state-fg-success`, and their delete counterparts to `--v2-state-fg-danger`. Optional palette seeds `diffAdd` and `diffDelete` also exist for theme resolution; they are theme inputs, not CSS variable names, and built-in overrides can supersede generated results.
## Syntax foreground tokens
```text
--syntax-comment
--syntax-regexp
--syntax-string
--syntax-keyword
--syntax-primitive
--syntax-operator
--syntax-variable
--syntax-property
--syntax-type
--syntax-constant
--syntax-punctuation
--syntax-object
--syntax-success
--syntax-warning
--syntax-critical
--syntax-info
--syntax-unknown
```
`--syntax-unknown` is referenced by the registered highlighter, but no definition was found in the current UI theme sources: treat it as an unresolved reference, not an existing resolved theme token. `--syntax-success` is available globally, although not directly used by that registered theme's current token-color rules.
Built-in syntax overrides additionally reference `--v2-text-text-muted`, `--v2-pink-800`, `--v2-green-800`, `--v2-orange-800`, `--v2-purple-800`, and `--v2-red-800` (with separate dark/light resolutions). Other syntax tokens use the ordinary text tokens or theme-resolved colors.
## Pierre renderer color variables — complete relevant inventory
These come from the installed `@pierre/diffs`**1.5.1** stylesheet plus OpenCode's injected CSS. `*-override` variables are hooks; other variables include derived outputs and internal implementation details, not independent OpenCode theme tokens. Some optional hooks are unset by default.
| Group | Variables |
|---|---|
| Base | `--diffs-bg`, `--diffs-fg`, `--diffs-mixer`, `--opencode-diffs-bg`, `--bg`, `--fg` |
| Conflict backgrounds (library support; not exercised by the fixture) | `--conflict-bg-current-header-override`, `--conflict-bg-current-number-override`, `--conflict-bg-current-override`, `--conflict-bg-incoming-header-override`, `--conflict-bg-incoming-number-override`, `--conflict-bg-incoming-override` |
Related mix controls are **percentages, not colors**: `--mix-light`, `--mix-dark`, `--mix-deco-light`, `--mix-deco-dark`, `--mix-selection-light`, `--mix-selection-dark`, and `--diffs-editor-active-line-source-mix`.
- Custom fold UI: `--surface-diff-hidden-base`, `--surface-diff-hidden-strong`, `--icon-strong-base`; `--text-mix-blend-mode` affects compositing but is not a color.
-`diff-color-fixture/README.md` is tracked and unchanged.
- Only fixtures changed. No renderer, token, behavior, formatter, push, PR, or deployment changes.
## Actual UI and how to open
The captures use the real source-backed OpenCode web app from this worktree, connected to the existing V2 background server (2.0.21). No mocked data or copied renderer. The app dev server remains available at `http://127.0.0.1:4444` and was started with `VITE_OPENCODE_SERVER_PORT=49374 bun run dev -- --port 4444` from `packages/app`. Existing app/server processes were not restarted.
Alternatively open this session in OpenCode Desktop; its location is now this worktree. Toggle Review, choose **Git changes**, and select a fixture file. This UI calls the working-tree source “Git changes,” not “Uncommitted changes.” Use the file-tree toggle if the list is hidden. Unified/Split controls are in the review toolbar.
Captures: 1440 × 1000 CSS pixels, default OpenCode (`oc-2`) theme, system light/dark media preference. Chat pane was resized to its minimum and the file tree hidden to give the diff about 964 pixels. Wrapping was enabled by the production viewer. Close-ups below are unchanged crops of those screenshots, not re-rendered mockups.
Additional unified captures are in this same directory.
## Case notes
| Case | Unchanged sections | Red/green scan | Inline emphasis | Text readability |
|---|---|---|---|---|
| Single additions/deletions/replacements, adjacent rows (`01`, `03`) | Neutral background recedes; unchanged syntax can still draw the eye | Light is clearer; dark relies more on gutter bars and line numbers | Distinct but restrained | TS strings stay green even on red deletion rows, competing with change meaning |
| Contiguous versus separated edits (`01` lines 11–12; long line) | Common words inside an inline span do **not** recede | Row direction is still clear from gutter and tint | Current production `word-line` rendering joins separated edits into one broad span; multiple independent highlight islands were not available in this state | Broad wrapped emphasis adds visual weight without improving precision |
| One-character, start/middle/end edits (`01` lines 13–16; `03`) | Surrounding text remains readable | Large words are easy to find; one-character changes need deliberate attention | One-character patch is visible but easy to miss at normal scan speed | No observed loss of string readability |
| Punctuation and spaces/indent/trailing spaces (`01` lines 17–20; `03`) | Quiet surroundings | Row markers expose that something changed | Small semicolon/space blocks are visible on close inspection; blank blocks give little explanation without visible whitespace glyphs | Readable; punctuation has less salience than colored keywords |
| Isolated hunks and collapsed sections (`02`) | Neutral context recedes; three bars visible simultaneously in split | Distinct isolated rows, but dark tints are weak | Small changed-word blocks don't dominate rows | Fold-label text is unusually faint, especially dark; context comments are clearer than the fold labels |
| TS/TSX types, keywords, comments and JSX (`01`) | Syntax in unchanged rows remains fairly prominent | Pink keywords/green strings compete with the red/green layer | Emphasis is generally subordinate; the long merged span is the exception | Comments/strings/types readable overall; muted JSX words like “old,” “new,” and button text are weak on tinted/highlighted backgrounds in dark mode |
| Markdown and wrapped prose (`03`) | Plain prose is visually quieter than code | Light tints are clearer than dark | Long spans emphasize unchanged middle words as well as changes | Plain prose remains readable but looks muted; bold/heading syntax draws attention more strongly than small edit spans |
| Long/wrapped lines (`01`, `03`) | Common middle text is swallowed by merged emphasis | Tinted multi-line blocks are apparent | Broad highlights repeat across wraps and visually dominate small changes | Wrapping works in unified and split; side-by-side requires more vertical scanning |
| File beginning/end (`01`, `02`, `03`) | Neutral surrounding rows recede | Changes shown at first/last lines | Same inline behavior as middle edits | Readable |
| Dense larger file (`04`: 63 lines, 40 replacements / 80 changed rows) | **Not visually checked** | Not checked | Not checked | Fixture ready for review |
| Added/deleted whole files (`05`, `06`) | Viewer lists both with D/A badges | **File contents not visually checked** | Not checked | Fixture ready for review |
| Narrow viewport (planned 800 × 1000) | **Not checked** | Not checked | Not checked | Stopped expanding capture scope at user request |
## Specific problems to carry into a later design pass
1. Dark row backgrounds have weak separation from neutral context; syntax hue is often more salient than change direction.
2. Green string syntax appears on deleted rows too, weakening red/green semantic scanning. Pink keywords similarly attract attention independently of change status.
3. Collapsed-section labels recede too far: their subdued text is harder to read than surrounding context comments.
4. One-character, punctuation, and whitespace highlights are easy to miss without focused inspection. Their issue is subtlety/size, not overpowering brightness.
5. Muted JSX text on dark inline backgrounds is less readable than the surrounding syntax. Broad inline spans across wrapped lines add disproportionate visual mass.
The merged-span behavior is a renderer observation, not a proposed color fix. No final values selected, no fixes implemented, and no contrast-ratio compliance claim made.
## GitHub Desktop comparison boundary
The referenced discussion/image was not included in this session. A precise comparison to that direction is therefore **not verified**. As a provisional hierarchy comparison only: quiet context → recognizable changed row → localized stronger inline emphasis is the useful target. OpenCode already keeps ordinary inline backgrounds subordinate to rows, but dark row separation, syntax competition, faint fold controls, and merged broad spans weaken that hierarchy. This is not a claim that the specific GitHub Desktop reference was inspected.
Scope was intentionally stopped after the user's request to do less. No full lint/typecheck was run: these are standalone visual examples with deliberate whitespace/formatting changes, not product code.
-export const longLine = "Start unchanged | The old checkout flow waits for manual confirmation before showing the receipt | Middle unchanged with punctuation: brackets [one, two], braces {three}, quotes 'four', slash /five/ | The old final instruction asks the reviewer to close the window | End unchanged"
+export const longLine = "Start unchanged | The new payment flow waits for automatic confirmation before showing the receipt | Middle unchanged with punctuation: brackets [one, two], braces {three}, quotes 'four', slash /five/ | The new final instruction asks the reviewer to keep the window | End unchanged"
Unchanged prose should remain quiet, not compete with changed rows.
-Replace one contiguous phrase: the original checkout experience stays here.
-Several little edits: red apple, cold tea, slow train.
-One character: ticket 7.
-before — the middle and the end remain unchanged.
-The start remains — before — the end remains.
-The start and the middle remain — before
-Punctuation only: hello, world!
-Whitespace only: one two three
-Trailing whitespace only.
-Delete this sentence without a replacement.
+Replace one contiguous phrase: the redesigned payment summary stays here.
+Several little edits: green apple, warm tea, fast train.
+One character: ticket 8.
+after — the middle and the end remain unchanged.
+The start remains — after — the end remains.
+The start and the middle remain — after
+Punctuation only: hello; world?
+Whitespace only: one two three
+Trailing whitespace only.
This unchanged sentence separates a deletion from an addition.
+Add this sentence without a corresponding deletion.
-Long prose: The original checkout flow asks the reader to review the old confirmation message before proceeding through the receipt screen, while the unchanged middle of this deliberately long paragraph checks whether horizontal scrolling or line wrapping keeps small inline edits visible at a narrow viewport; the final old phrase appears near the far right edge.
+Long prose: The redesigned payment flow asks the reader to review the new confirmation message before proceeding through the receipt screen, while the unchanged middle of this deliberately long paragraph checks whether horizontal scrolling or line wrapping keeps small inline edits visible at a narrow viewport; the final new phrase appears near the far right edge.
-**Syntax-like Markdown** competes with `inline code`, [links](https://example.com/old), and _emphasis_.
+**Syntax-rich Markdown** competes with `inline code`, [links](https://example.com/new), and _emphasis_.
+This entire file is newly added, not a replacement for the deleted example.
+Only addition tint should be present.
+Plain text: keep the words readable against the changed-row background.
+A deliberately long added line checks the horizontal extent of the green row while the reviewer compares the quiet surrounding application surface against the color used for new content in both dark and light modes.
{id:"bg",label:"File background",help:"--diffs-bg. Base for context rows and Pierre's derived mixes.",rule:hostVar("--diffs-bg"),light:"v2-grey-50",dark:"v2-grey-1100"},
{id:"buffer-bg",label:"Split empty side",help:"--diffs-bg-context-gutter-override (blank side of split rows).",rule:hostVar("--diffs-bg-context-gutter-override"),light:"v2-grey-100",dark:"v2-grey-1000"},
{id:"buffer-hatch",label:"Split empty side hatch",help:"--diffs-bg-buffer-override (Pierre stripes; OpenCode currently hides them).",rule:hostVar("--diffs-bg-buffer-override"),light:"v2-grey-200",dark:"v2-grey-900"},
{id:"hover",label:"Hover tint",help:"--diffs-bg-hover-override. Mixed in at ~3% light / ~9% dark, not exact.",rule:hostVar("--diffs-bg-hover-override"),light:"v2-grey-1200",dark:"v2-grey-50"},
{id:"selection",label:"Selected lines",help:"--diffs-selection-base (row and number tints derive from it).",rule:scopedVar("--diffs-selection-base"),light:"v2-background-bg-accent",dark:"v2-background-bg-accent"},
{id:"add-row",label:"Added row",help:"Exact fill of added code rows.",rule:exactRow(`${ADD}${ROW}`),light:"v2-green-100",dark:"v2-green-1200"},
{id:"add-gutter",label:"Added gutter",help:"Exact fill of added line-number cells.",rule:exactRow(`${ADD}${GUTTER}`),light:"v2-green-100",dark:"v2-green-1200"},
{id:"add-inline",label:"Added inline highlight",help:"Word/char emphasis span, painted over the row.",rule:direct(`${ADD} [data-diff-span]`,(v)=>`background-color: ${v};`),light:"v2-green-200",dark:"v2-green-1100"},
{id:"add-inline-text",label:"Added inline text",help:"Text color inside the highlight. Replaces syntax colors there.",rule:inlineText(ADD),light:"v2-green-1000",dark:"v2-green-300"},
{id:"add-seed",label:"Addition seed",help:"--diffs-addition-color-override. Feeds anything above left on Current.",rule:hostVar("--diffs-addition-color-override"),light:"v2-state-fg-success",dark:"v2-state-fg-success"},
]],
["Deletions",[
{id:"del-row",label:"Deleted row",help:"Exact fill of deleted code rows.",rule:exactRow(`${DEL}${ROW}`),light:"v2-red-100",dark:"v2-red-1200"},
{id:"del-gutter",label:"Deleted gutter",help:"Exact fill of deleted line-number cells.",rule:exactRow(`${DEL}${GUTTER}`),light:"v2-red-100",dark:"v2-red-1200"},
{id:"del-inline",label:"Deleted inline highlight",help:"Word/char emphasis span, painted over the row.",rule:direct(`${DEL} [data-diff-span]`,(v)=>`background-color: ${v};`),light:"v2-red-200",dark:"v2-red-1100"},
{id:"del-inline-text",label:"Deleted inline text",help:"Text color inside the highlight. Replaces syntax colors there.",rule:inlineText(DEL),light:"v2-red-1000",dark:"v2-red-300"},
{id:"del-seed",label:"Deletion seed",help:"--diffs-deletion-color-override. Feeds anything above left on Current.",rule:hostVar("--diffs-deletion-color-override"),light:"v2-state-fg-danger",dark:"v2-state-fg-danger"},
]],
["Text",[
{id:"number",label:"Unchanged line-number text",help:"--diffs-fg-number-override. Also fold labels unless set below.",rule:hostVar("--diffs-fg-number-override"),light:"v2-text-text-faint",dark:"v2-text-text-faint"},
{id:"fold-text",label:"Fold label and expand icon",help:"“N unmodified lines” text and expand button.",rule:direct(":is([data-separator-content], [data-expand-button])",(v)=>`color: ${v};`),light:"v2-text-text-muted",dark:"v2-text-text-muted"},
$("#reset").onclick=()=>{if(!confirm("Reset every slot to Current?"))return;state={...state,enabled:true,disabledGroups:{},slots:{}};render();changed(true)}
// Shared CSS carries the exact settings in a header so pasting restores token choices, opacity and category toggles.
return`// Temporary diff color tuning output. Generated by diff-color-review-artifacts/tuner; delete with the tuner.\nexport const diffColorTuningCSS = \`${escaped}\`\n`
"dev":"bun run --cwd packages/cli --conditions=browser src/index.ts",
"dev:live":"OPENCODE_TUI_CHANNEL=dev OPENCODE_PASSWORD=\"$(opencode2 service get password)\" bun run dev --server \"$(opencode2 service status)\"",
"dev":"bun run --cwd packages/cli src/index.ts",
"dev:live":"sh -c 'OPENCODE_TUI_CHANNEL=dev OPENCODE_PASSWORD=\"$(opencode service get password)\" exec bun run dev \"$@\" --server \"$(opencode service status)\"' --",
"dev:vite":"bun run --cwd packages/cli --conditions=browser dev/vite.ts",
"dev:vite:live":"sh -c 'OPENCODE_TUI_CHANNEL=dev OPENCODE_PASSWORD=\"$(opencode service get password)\" exec bun run dev:vite \"$@\" --server \"$(opencode service status)\"' --",
"dev:desktop":"bun --cwd packages/desktop dev",
"dev:web":"bun --cwd packages/app dev",
"dev:console":"ulimit -n 10240 2>/dev/null; bun run --cwd packages/console/app dev",
"dev:stats":"bun sst shell --stage=production -- bun run --cwd packages/stats/app dev",
Per-type constructors live on the type, not as top-level re-exports. Use `Message.system(...)`, `Message.user(...)`, `Message.assistant(...)`, `Message.tool(...)`, `LanguageModel.make(...)`, `ToolDefinition.make(...)`, `ToolCallPart.make(...)`, `ToolResultPart.make(...)`, `ToolChoice.make(...)`, `ToolChoice.named(...)`, `SystemPart.make(...)`, and `GenerationOptions.make(...)` directly. The top-level `LLM` namespace is reserved for request-shaped call APIs: `LLM.request`, `LLM.generate`, `LLM.stream`, and `LLM.generateObject`. Use `LLMRequest.update(...)` when deriving canonical request data; do not add a duplicate `LLM.updateRequest(...)` path. Two ways to construct the same thing is one too many.
Per-type constructors live on the type, not as top-level re-exports. Use `Message.system(...)`, `Message.user(...)`, `Message.assistant(...)`, `Message.tool(...)`,`Message.media(...)`,`LanguageModel.make(...)`, `ToolDefinition.make(...)`, `ToolCallPart.make(...)`, `ToolResultPart.make(...)`, `ToolChoice.make(...)`, `ToolChoice.named(...)`, `SystemPart.make(...)`, and `GenerationOptions.make(...)` directly. The top-level `LLM` namespace is reserved for request-shaped call APIs: `LLM.request`, `LLM.generate`, `LLM.stream`, and `LLM.generateObject`.`LLM.generate`/`LLM.stream` and Promise `ai.llm.generate`/`ai.llm.stream` accept ergonomic input or an `LLMRequest`; both paths use the same canonical request. Core still builds, logs, replays, and updates that durable `LLMRequest` boundary. Use `LLMRequest.update(...)` when deriving canonical request data; do not add a duplicate `LLM.updateRequest(...)` path.
- Keep provider-defined string enums forward-compatible. Expose known values for autocomplete while accepting future values with `Known | (string & {})`; use `Schema.String` at runtime unless rejecting unknown values is required for correctness.
Modality namespaces mirror `LLM` exactly: `Image.request`, `Image.generate`, `Image.stream`, and the same for `Video`, `Speech`, and `Transcription`. Common request fields (`images`, `mask`, `n`, `size`, `aspectRatio`, `seed`, `format`) lower natively or fail with a typed `AIError`; provider-native controls always live under `providerOptions`, never under a modality-specific `options` key.
Media payloads are always `Media.Asset` (`src/media.ts`). Construct them with `Media.bytes`, `Media.base64`, `Media.url`, `Media.ref`, `Media.fromDataUrl`, or `Media.file`; never introduce a parallel `data: string | Uint8Array` shape. `MediaPart.media`, `ImageRequest.images`/`mask`, `ImageResponse.images`, and the `media``LLMEvent` all share it. Protocols branch on `asset.source.type` and `asset.kind` and use `ProviderShared.requireInlineMedia` / `inlineRequired` / `mediaUrl` / `mediaReference` and `MediaInput.inlineBytes` / `refID` rather than re-deriving base64 or URL handling.
`schema/messages.ts → media.ts → route/executor-service.ts` is an accepted runtime dependency from the schema layer on the executor service tag: `Media.Asset.bytes()` must be able to download `url` sources, and the tag lives in that leaf module precisely so the schema barrel never imports the executor implementation (which imports the schema barrel back). Do not move the tag into `route/executor.ts` or import `route/executor.ts` from `src/schema/*` or `src/media.ts`.
Nothing in `src/*` except `src/promise.ts` may know about Promises. `@opencode/ai/promise` (`AI.make({ layer? })`, default `ai`) is the single Promise/`AsyncIterable` surface for LLM and media; it runs the Effect APIs in one `ManagedRuntime` and rethrows `AIError` unchanged.
- Prefer forward compatibility for provider-defined options that OpenCode only passes through. For pass-through string enums, expose known values for autocomplete while accepting future values with `Known | (string & {})`, and accept any string at runtime. Closed literals are appropriate when OpenCode branches on a value, transforms its associated structure, or otherwise cannot correctly handle an unknown variant. New options whose shape or behavior requires implementation remain unsupported until they are handled; do not blindly forward unknown structures.
- Order reasoning-effort values from lowest to highest: `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. Provider-specific subsets follow the same relative order in types, schemas, option lists, and tests.
`LLM.request(...)` builds an `LLMRequest`. `LLMClient.generate(...)` reads the executable route carried by `request.model.route`, builds the provider-native body, asks the route's transport for a real `HttpClientRequest.HttpClientRequest`, sends it through `RequestExecutor.Service`, parses the provider stream into common `LLMEvent`s, and finally returns an `LLMResponse`.
Route defaults are request-shaping defaults such as `headers`, `limits`, `generation`, `providerOptions`, and `http`. Endpoint host/query belongs on the route endpoint. Selected `LanguageModel` values carry only model id, provider id, and the configured route value. Model capability/catalog metadata lives outside this package; protocol support is enforced by request lowering and typed `AIError`s.
The four-axis decomposition is the reason DeepSeek, TogetherAI, Cerebras, Baseten, Fireworks, and DeepInfra all reuse `OpenAIChat.protocol` verbatim — each provider deployment is a 5-15 line`Route.make(...)` call instead of a 300-400 line route clone. Bug fixes in one protocol propagate to every consumer of that protocol in a single commit.
The four-axis decomposition is the reason DeepSeek, TogetherAI, Cerebras, Baseten, Fireworks, and DeepInfra all reuse `OpenAIChat.protocol` verbatim — each provider owns a small`Route.make(...)` composition instead of a protocol clone. Bug fixes in one protocol propagate to every consumer of that protocol in a single commit.
When a provider supports multiple physical transports, selection remains execution policy below its semantic route. `OpenResponsesChannel.transport(...)` owns the provider-neutral Responses WebSocket concept: it prepares one final request, executes HTTP by default, strips WebSocket-disallowed fields, and passes a generic channel exchange to a per-call `WebSocketChannelExecutor` when supplied. Provider-specific Responses routes opt in with handshake and connection-age policy. `Route.streamPrepared` owns decoding and acknowledges channel completion only after successful full consumption.
### Media Routes
Media does not fit the SSE-frames-to-event-state-machine LLM route. `MediaRoute.inline(...)` / `queued(...)` / `stream(...)` (`src/route/media.ts`) compose a `MediaProtocol` kind with `Endpoint` and `Auth` and own the transport plumbing: `http` option merging, URL/query rendering, auth headers, JSON vs multipart encoding, and handing the response back to the protocol. `MediaProtocol.inline` (`src/route/media-protocol.ts`) is `body.from(request)` plus `response.decode(response, context)`; each protocol declares `const route = MediaProtocol.identity({ id, name, provider })` once and decodes through `route.decodeJson` / `route.text` / `route.decodeStarted` so decode failures retain the raw body and HTTP context, raising `route.unsupported(operation, message)` for requests it cannot lower, and passes `route` as the first argument to `MediaProtocol.inline` / `queued` / `stream`. `Generation` (`src/generation.ts`) is the provider-neutral handle for a queued generation over a `GenerationRoute` (`status`, `result`, `cancel`). Image protocol files follow the same section order as LLM protocols and declare unsupported common fields once through the protocol's `unsupported` list.
`MediaProtocol.queued` is the submit-then-poll kind every video route uses: `start` (body + decode into `{ token, snapshot }`), `status`, `result`, and optional `cancel` (with `activeOnly` when the provider's cancel endpoint deletes finished work, as Runway's does: the route refreshes status first and skips terminal generations), each addressed by a route-owned `token` whose `Schema.Codec` makes it serializable. `MediaRoute.inline` and `MediaRoute.queued` compose the two kinds with `Endpoint` and `Auth`; the queued route decodes the token once at the boundary (`start` output or `resume` input) and closes over it in a token-free `GenerationRoute` (`status`/`result`/`cancel` are plain Effects), so `Generation` never sees the token's shape and only carries the encoded JSON for persistence. Polls reuse the route's auth and deployment headers plus the request's `http` overlay after `start`, and resolve relative paths against the route base URL (provider-issued absolute URLs such as fal's `status_url` pass through). `result` is always its own GET even when the provider returns output inside the status document, so `Generation.await` behaves the same after `start` and after `resume`. `PollContext.auth` carries only what `Auth` added or changed so protocols can hand download credentials to output assets as transient `Media.Asset.headers` (Veo) — never part of `source` or JSON. Status strings map through a per-protocol `STATUS` table via `MediaProtocol.status`; terminal generations without output fail through `output.ended` / `output.contentPolicy` with the provider document on `reason.body`; a `failed` generation maps the provider's error code through a per-protocol `FAILURE` table via `MediaProtocol.failure` so rejected inputs are not reported as retryable `ProviderInternal`. `GenerationAwaitOptions` (`AwaitOptions` in `src/generation.ts`, `{ poll?: Poll }`) is the one options type for `await`, `events`, `Video.generate`, and `Video.stream`.
`MediaProtocol.stream` is the incremental kind every speech route uses, with the same discipline as LLM protocols. `MediaRoute.stream` submits the caller's request as `MediaProtocol.Addressed<Request>` (`{ ...request, mode }`, `mode: "generate" | "stream"`), so one provider stays one protocol: `body.from`, the endpoint path, and `frames` read `request.mode` to pick the body, path, and framing. `frames(bytes, context)` returns frames — `Framing.sse`, `Framing.lines`, `Framing.document` (a single-document response shaped like a streamed record), or the raw `bytes` for chunked audio. `initial()` is fresh per-response parser state; `step` folds each frame into it and emits modality events; `finish(state, context)` runs once after the last frame with the request, body, and observed `http` (header-only usage lives there) and emits exactly one terminal event or fails with `route.incomplete()`. Keep parser state to real accumulators and derive anything the request or body determines in `finish`. `generate` runs the same stream and folds it with the modality's `collect`. Request-derived URL parameters go on the body's `query` (array values repeat the parameter), applied before route and caller `http.query`. Decode frames with `route.decodeFrame` and raise stream-time failures with `route.frameError` (the frame stays on `reason.body`); protocols never thread HTTP context, because the route fills `reason.http` on stream errors that lack it. Speech protocols share `protocols/utils/speech-stream.ts` for deltas, timestamps, voice ids, PCM and container descriptions, and the terminal asset.
Every modality route is the inline | stream | queued union (transcription uses all three: OpenAI and Gemini stream, Deepgram and ElevenLabs are inline, AssemblyAI is queued), every client is `MediaClient.make(Service, { modality, responseEvents })` (`src/media-client.ts`), which dispatches on the route's `kind`, and every model composes through `composeRoute`. fal queue protocols come from `protocols/utils/fal-queue.ts`, bodies are `json`, `multipart`, or `binary` (a raw upload), and a queued protocol that must upload media before submitting implements `start.prepare` (`MediaProtocol.Prepare`; AssemblyAI `/v2/upload`).
### URL Construction
`Endpoint` owns `{ baseURL, path, query }`. Each protocol route includes a canonical endpoint when the provider has one (e.g. `https://api.openai.com/v1`); provider helpers override endpoint fields by configuring the route before selecting a model. Generic OpenAI-compatible routes have no canonical URL and require configuration before execution.
@@ -93,11 +112,12 @@ For providers where the URL is derived from typed inputs (Azure resource name, B
### Provider Facades
Provider-facing APIs are configured facades over route values. Endpoint/auth/resource/API-version setup happens before model selection, and model selectors accept only a model or deployment id:
Provider-facing APIs are configured facades over route values. Endpoint/auth/resource/API-version setup happens before model selection, and model selectors accept only a model or deployment id. Media models use per-modality selectors on the same facade (`openai.image(id)`, `.speech(id)`, `.transcription(id)`, `google.video(id)`) that mirror `openai.responses(id)`; the one-word overlap with the request namespace is accepted over a second construction path:
@@ -115,18 +135,20 @@ Keep provider facades small and explicit:
- Prefer `apiKey` as provider-specific sugar and `auth` as the explicit override; keep them mutually exclusive in provider option types with `ProviderAuthOption`.
- Resolve `apiKey` → `Auth` with `AuthOptions.bearer(options, "<PROVIDER>_API_KEY")` (it honors an explicit `auth` override and falls back to `Auth.config(envVar)` so missing keys surface a typed `Authentication` error rather than a runtime crash).
- Use separate top-level facades for products with different required setup, such as `CloudflareAIGateway` and `CloudflareWorkersAI`.
- Give every named provider its own file and top-level export. Keep its endpoint, auth defaults, and route setup in that file. Compose shared protocols directly; do not nest named provider presets under generic compatible facades or keep their endpoints in a shared provider profile registry.
`Provider.make(...)` remains available for simple static provider definitions, but new built-in providers should prefer plain configured facades unless a helper removes real duplication without adding runtime behavior.
### Provider Package Entrypoints
Catalog-selected native providers use package-like export paths from `@opencode-ai/ai`. They are internal entrypoints in one npm package, not separately published provider packages. Every entrypoint implements `ProviderPackage.Definition` and exposes `model(modelID, settings)`, where settings are serializable provider configuration plus common `headers`,`body`, and `limits` overlays.
Catalog-selected native providers use package-like export paths from `@opencode/ai`. They are internal entrypoints in one npm package, not separately published provider packages. Every entrypoint implements `ProviderPackage.Definition` and exposes `model(modelID, settings)`, where settings are one flat serializable object: the connection keys the entrypoint declares (`apiKey`, `baseURL`, `region`, …), the common `headers` and`body` overlays, and the protocol's request options (`reasoningEffort`, `thinking`, …) side by side. Each entrypoint destructures its own connection keys and passes the rest to the route as `providerOptions`; there is no nested `providerOptions` at the entrypoint.
@@ -163,6 +185,10 @@ Native chronological system messages are route/model-specific. Open Responses lo
The wrapped-user fallback preserves ordering while visibly lowering authority. Never silently pass a raw chronological `role: "system"` through a route that might reject it. Do not insert raw retrieved documents, tool output, or web content into privileged chronological system updates; keep untrusted content in ordinary user/tool channels.
### Effort Updates
`Message.effort({ effort, previous })` is a chronological "reasoning effort changed here" marker (`undefined` means the model default). Changing a top-level effort invalidates the whole provider prompt cache, so protocols with a native per-message update (`Protocol.supportsEffortUpdates`) keep the top-level effort at the first marker's `previous` and lower each marker in place: Anthropic Messages emits an empty `role: "system"` message with `output_config.effort` plus the `mid-conversation-output-config-2026-07-01` beta, and OpenAI Responses emits `configuration_update` items. `applyEffortUpdates` runs in `prepareRequest` and strips the markers for every other route, so a protocol without support keeps today's plain top-level behaviour. When the last marker disagrees with the effort the request asks for (reverted or forked history), `resolveEffortUpdates` strips the markers and falls back to a plain top-level change.
### Tools
Tool loops are represented in common messages and events:
@@ -249,6 +275,7 @@ Use this order for every protocol module:
### Rules
- Keep protocol files focused on the protocol. Move provider-specific projection, signing, media normalization, or other bulky transformations into `src/protocols/utils/*`.
- Send `tool.inputSchema` as given. `prepareRequest` applies the tool schema rules (`ToolSchemaProjection.tools`) once per request, including tools in namespaces. A protocol whose API needs a model family's rules for every model declares `sanitizer` instead of transforming schemas itself.
- Use `Effect.fn("Provider.fromRequest")` for request body construction entrypoints. Use `Effect.fn(...)` for event handlers that yield effects; keep purely synchronous handlers as plain functions returning a `StepResult` that the dispatcher lifts via `Effect.succeed(...)`.
- Parser state owns terminal information. The state machine records finish reason, usage, and pending tool calls; emit one terminal `finish` event (or `provider-error`) for each completed response. If a provider splits reason and usage across events, merge them in parser state before flushing.
- Emit exactly one terminal `finish` event for a completed response, normally after a matching `step-finish`. Use `stream.terminal` to stop reading when the provider has a completion sentinel; use `stream.onHalt` when the final event must be flushed after the framed stream ends.
`@opencode/ai` becomes the one package you reach for to generate anything: text, images, video, speech, transcripts, and later music and realtime. The LLM surface already exists and is shaped by three constraints: Effect-first, used by OpenCode Core, usable externally. Media has a different priority order: **external DX first**, Effect and Promise as peers, Core as one consumer among many.
The design below is derived from a survey of the raw provider APIs (OpenAI, Gemini/Veo/Imagen, xAI, Stability, BFL, fal, Replicate, Runway, Luma, Kling, MiniMax, ElevenLabs, Deepgram, Cartesia, AssemblyAI, Lyria) and of existing multi-provider SDKs.
## What the survey forces
1.**Three execution shapes, everywhere.** Inline sync (OpenAI images, all TTS, Gemini), async job with polling or webhook (every video provider, BFL, fal, Replicate, AssemblyAI), and bidirectional streams (ElevenLabs/Cartesia/Deepgram WS, realtime). Video has no sync provider at all.
2.**Output is never just bytes.** base64, signed URLs with TTLs from 10 minutes (BFL, so its route downloads before returning via `PollContext.materialize`) to 2 days (Veo), URLs that need auth plus redirect (Veo), separate download endpoints (Sora `/content?variant=`), raw bodies (Stability, TTS). Multi-output is the norm.
3.**Inputs have roles.** First/last frame, mask, style/subject reference, source video for edit/extend, reference audio, prior generation id, provider-side file handles (`file_id`, `gs://`, `runway://`, `mm_file://`).
4.**Partial streaming is modality-specific.** Images: a few whole partial frames. Audio: ordered chunks plus timestamp events. Jobs: status/progress/logs. Video: none.
5.**Usage is a union**: tokens, seconds, characters (often only in headers), credits, compute time.
6.**Moderation can be partial success** (Veo strips audio but returns video). Deprecations are constant (Sora API shuts down 2026-09-24; Imagen is shut down on the Gemini API and past its 2026-06-30 discontinuation date on Vertex).
## Where existing SDKs are weak and we should not be
- No streaming TTS.
- Video handles are experimental start/status pairs; the polling loop lives inside the generate call.
- Unsupported inputs become silent warnings arrays, so a request can succeed while dropping your mask.
-`n` is fanned out into hidden parallel calls, which obscures cost and idempotency.
- Each modality has its own bespoke result type; the file abstraction is a lazy base64/bytes pair with no URL, expiry, or provider ref.
- Effect's own `unstable/ai` has no media generation. Nothing in the Effect ecosystem owns this.
## Design principles
- **Same shape as LLM.** `X.request(...)` → Schema class; `X.generate(request)` / `X.stream(request)`; `XClient.Service` + `layer`; typed `AIError`. If you know `LLM`, you know `Video`.
- **Execution shape is route policy, not API shape.** `Image.generate` returns an image whether the provider is inline or queued. Job control is available uniformly when you want it.
- **Errors, not warnings.** Unsupported common fields fail at the protocol boundary with a typed `AIError`, as the LLM routes do today. Provider-side partial results (filtered audio, moderated sample) surface as `notices` on the response, never as silent drops.
- **One asset type in, one asset type out**, shared with LLM messages and tool results.
- **Typed per-model options**, no hidden fan-out, no implicit retries that spend money.
- **Promise API is one mechanism for the whole package**, not a media-only wrapper.
- **One construction path per model.** Media models come from per-modality selectors on the configured facade (`openai.image("gpt-image-2")`), the same shape as `openai.responses("gpt-5")`.
## Public API
### Model selection
A model value is built as `OpenAI.configure({ apiKey }).responses("gpt-5")` or `.image("gpt-image-2")`: `configure` fixes credentials, endpoint, and defaults; the selector fixes which of the provider's APIs to hit and binds the typed `providerOptions` generic. Media follows the same shape with one selector per modality — `.image(id)`, `.video(id)`, `.speech(id)`, `.transcription(id)` on the facades that offer each — mirroring `openai.responses(id)`. `Image.request` accepts `ImageModel` only, exactly as `LLM.request` accepts `LanguageModel`.
The request namespace and the selector share one word (`Image.request` + `.image(...)`). That redundancy is accepted: a callable facade returning a lazily resolved ref would be a second way to construct the same model, and the type machinery to infer `providerOptions` through it is not worth one word. Where a provider has two routes for one modality, the selectors stay explicit (`openai.chat`, `stability.image` inline vs `stability.upscale()` queued), and one default per modality per provider is part of the facade definition (OpenAI image → Images API, Google image → Gemini-native; Imagen is shut down, so there is no `google.imagen`). The facade selector (`openai.image(id)`) is the public path for media models. Modality-specific package entrypoints (`model(modelID, settings)` beside today's LLM paths such as `@opencode/ai/providers/openai/responses`) are deferred until Core has a modality-aware model resolver; Core's resolver accepts only `LanguageModel` today.
### `Media` — the asset type
Replaces `MediaPart.data: string | Uint8Array`, `ImageInput`, `GeneratedImage`, and aligns `Tool.FileContent`.
`size` and `aspectRatio` are not interchangeable; each route rejects fields it cannot lower — see the portability table
in the README's Image generation section.
Editing is not a separate function; `images`/`mask` on the request select the edit path in the route (OpenAI `/images/edits`, Gemini multimodal parts, xAI `/images/edits`). Routes that cannot honor `mask` fail with `Unsupported`.
`ImageRoute` is the inline | stream | queued union, dispatched on `route.kind`, like every modality route. `Image.stream` on a streaming route emits `image-partial` previews before each `image`; on a queued route it emits `generation-queued` / `generation-progress` observations, then the result's `image` and `finish` events.
Execution is `MediaProtocol.stream` for every provider: one request whose body is framed and folded by a `step`
state machine, with `generate` running the same stream and collecting it. The route submits the request with its
`mode` (`"generate" | "stream"`), which lets one provider stay one protocol — OpenAI adds `stream_format: "sse"` (except `tts-1`/`tts-1-hd`, which stream raw bytes), ElevenLabs appends
`/stream`, Cartesia switches `/tts/bytes` to `/tts/sse`, Gemini switches `generateContent` to
`streamGenerateContent`. The terminal `finish` event carries the assembled asset (every provider's stream is
concatenable chunks), so stream consumers also get the whole file and `generate` is just "take `finish`, gather
`timestamps`". The cost is memory: a stream holds every chunk until `finish`, so even a consumer that only plays deltas
keeps the whole clip in memory. That is bounded by the providers' input text limits (a few minutes of audio); a
long-form or session API would need an opt-out.
**Voice.**`voice?: string | { id: string }`. A string is passed through as the provider's native identifier — a
name on OpenAI and Gemini, a voice id on ElevenLabs (path segment) and Cartesia. `{ id }` selects an OpenAI custom
voice and is treated as the plain string on routes that do not distinguish custom from built-in. Deepgram's voice is
the model id (`aura-2-thalia-en`), so `voice` is `unsupported` there. There is no cross-provider voice catalog or
name→id resolution. Multi-speaker (Gemini `speechConfig.multiSpeakerVoiceConfig`) and per-voice settings
(ElevenLabs `voice_settings`) go through `providerOptions`.
**Format and PCM.**`format` is container-level; provider sample rates and bitrates live under `providerOptions`
events(options?: GenerationAwaitOptions):Stream<GenerationEvent,AIError>// fails with Timeout past poll.timeout, checked per observation
}
GenerationAwaitOptions={poll?: Poll}
Poll={interval?: Duration;timeout?: Duration}
```
`Generation` is not video-specific. Image routes on BFL, fal, Replicate, and Stability `upscale()` are queued; `Image.start` exists for them. A route declares itself `inline` or `queued`; `generate` on a queued route is `start` then `await`.
Status polls and result reads retry transient failures (rate limits, provider 5xx, and transport errors, classified by the same `isRetryable` the Session runner uses) inside `MediaRoute.queued`. Only the HTTP exchange retries, never the decoded document: a terminal `failed` generation also surfaces as `ProviderInternal` and must not be re-read. Gaps grow exponentially from 1s with jitter, up to 30s each, honoring a provider `retry-after` up to that cap, for at most 8 retries. `await`, `events`, and `Video.stream` cut retries off at `poll.timeout` and fail with `Timeout`, so retries never extend the caller's deadline; a direct `result()` or `resume` read is bounded by the retry cap alone. `start` and `cancel` never retry: a repeated submit can start and bill a second job. The policy is internal; there is no option for it.
Interrupting `await`, `events`, or `Video.stream` (or aborting the promise API's `signal`) stops waiting only. The provider job keeps running and billing; call `cancel()` explicitly to stop it.
### Usage
```ts
Usage=
|{type:"tokens";input;output;total;details?}
|{type:"seconds";seconds}
|{type:"characters";characters}
|{type:"credits";credits}
|{type:"compute";seconds}
```
Header-only usage (ElevenLabs `character-cost`, Deepgram `dg-char-count`) is lifted into `usage` by the route.
### Promise API — `@opencode/ai/promise`
Mirrors the `packages/plugin/src/effect` and `packages/plugin/src/promise` split that already exists in this repo. One mechanism for LLM and media.
```ts
import{AI}from"@opencode/ai/promise"
constai=AI.make()// ManagedRuntime over RequestExecutor.fetchLayer + all clients
// AI.make({ layer }) to inject a custom executor / recorder / middleware
constimage=awaitai.image.generate({model,prompt})
awaitai.bytes(image.image)// also ai.base64, ai.materialize, ai.write(asset, path)
constresumed=awaitai.video.resume(model,JSON.parse(saved))// persist provider + model ID with the token
constrequest=ai.llm.request({model,prompt})
consttext=awaitai.llm.generate(request)
forawait(consteventofai.llm.stream(request)){…}
awaitai.dispose()
```
Streams become `AsyncIterable` via `Stream.toAsyncIterable`. `AIError` is thrown as-is. Aborting an `AbortSignal` interrupts the work and, like `fetch`, rejects the Promise or throws from the stream with `signal.reason` instead of ending the stream as if complete. Nothing in `src/*` except this entrypoint knows about promises.
### Providers
Existing facades gain per-modality selectors; the modality routes each facade provides (*italics* are not
implemented):
| Facade | llm | image | video | speech | transcription | other |
New facades follow the existing one-file-per-provider rule. The facade selector is the public path for media models; modality-specific package entrypoints (for example `@opencode/ai/providers/openai/images`) are deferred until Core has a modality-aware model resolver.
`ImageModel<Options>` gives typed `providerOptions` per model; `VideoModel`, `SpeechModel`, and `TranscriptionModel` follow the same generic. As with `LanguageModel`, the route type does not carry `Options`, so `ImageModel<OpenAIImageOptions>` is an `ImageModel` and client methods take plain `ImageRequestFor`. They share an internal `MediaModel` base class (ids, route, `http` overlays) that is not part of the public exports; `Generation` and the promise client work with the concrete modality models.
### Routes and protocols
Media does not fit the LLM four-axis route (SSE frames → event state machine) except for streaming TTS/STT. Reuse `Endpoint`, `Auth`, `Framing`, `RequestExecutor`, and add media protocol kinds:
-`MediaProtocol.inline` — `body.from(request)` (JSON, multipart, or query), `response.decode(response)` (JSON, or binary body → `Media.Asset`).
-`MediaProtocol.queued` — `start` (body + decode to `{ token, snapshot }`), `status`, `result`, optional `cancel`, and a `token` codec. `result` is always a separate GET (against the status document for Veo/xAI/Runway, fal's `response_url` otherwise) so `await` after `start` and after `resume` share one path. `PollContext.auth` hands the auth headers the route sent to the protocol for output URLs that need them (Veo downloads); they become transient `Media.Asset.headers`, never part of `source`. There is no separate `download` step: `Media.Asset.bytes()` downloads through the executor with those headers. `MediaRoute.inline(...)` / `MediaRoute.queued(...)` compose each kind with endpoint and auth; the queued route decodes the token once and hands `Generation` a token-free `{ status, result, cancel? }`.
-`MediaProtocol.stream` — `body.from(request)` over the request plus its `mode`, `frames` (a function that picks the framing for the call: `Framing.sse`, `lines`, `document`, or the raw bytes), fresh per-response `initial()` state, `step` emitting modality events, and `finish(state, context)` — with the observed response for header-only usage — emitting exactly one terminal event or failing as an incomplete stream. The route fills `reason.http` on stream errors. `MediaRoute.stream(...)` exposes `stream` and `generate` (the same stream folded by the modality's `collect`).
`MediaRoute.inline` / `MediaRoute.queued` / `MediaRoute.stream` compose one protocol kind with endpoint/auth and tag the route with its `kind`; `ImageModel`/`VideoModel`/`SpeechModel`/`TranscriptionModel` share the `MediaModel` base (`src/media-model.ts`).
### LLM integration
-`MediaPart` becomes `{ type: "media"; media: Media.Asset; … }` so protocols branch on `kind` and can pass `url`/`ref` sources through natively (OpenAI `image_url`, Gemini `fileData`).
- New `LLMEvent`s: `media { media: Media.Asset }` so Gemini inline image output is first-class instead of dropped. OpenAI Responses `image_generation_call` keeps its single carrier — the provider-executed `tool-result` with `file` content — because Core consumes hosted tool-result content today and has no `media` event handling yet; it switches to the `media` carrier when Core adopts the event, so the image is never emitted twice.
-`Message.assistant([...])` accepts media parts; Gemini multi-turn image editing replays them.
-`Tool.FileContent` aligns with `Media.Source`.
## Decisions
All settled:
1.**Per-modality selectors** (`openai.image(id)`, `.video`, `.speech`, `.transcription`) name media models, mirroring `openai.responses(id)`. The one-word overlap with the request namespace is accepted over a callable-facade `ModelRef` as a second construction path.
2.**`providerOptions` everywhere** (rename current `Image.options`) for consistency with LLM.
3.**No hidden `n` fan-out.**`n` lowers natively; routes that cannot do `n > 1` fail typed. Callers use `Effect.all` / `Promise.all` explicitly.
4.**Errors over warnings** for unsupported common fields; `notices` for provider-side partial results only.
5.**`Media.Asset` is a class** (lazy bytes, cached) with `Media.Source` as the serializable Schema for wire/persistence. `Asset.from(source)` / `asset.source` round-trip losslessly. Same pattern as `LanguageModel` today.
6.**Promise entrypoint**: `@opencode/ai/promise` exporting `AI.make(options?: { layer? })` plus a module-level default `ai` for scripts, covering LLM too.
7.**Modality set for v1**: `Image`, `Video`, `Speech`, `Transcription`. `Music`/`SoundEffect` and `session` (bidirectional WS, realtime) are designed-for but deferred.
8.**Sora is skipped** (API shuts down 2026-09-24). Video launches with Veo, xAI, fal, Runway.
## Build order
Foundation + Image ship together as the reference implementation, serially. Video, Speech, and Transcription then proceed in parallel on separate branches. Image jobs and partial streaming come last, after Video has hardened `Generation`.
## Phasing
1.**Foundation** — per-modality selectors, `Media`, `Generation`, `Poll`, `Usage` union, `MediaProtocol` kinds, `@opencode/ai/promise` with `llm` + `image`. Port the five existing image protocols onto it. Unify `MediaPart` and add the `media` LLM event (fixes Gemini image output being dropped).
3.**Speech + Transcription** — ✅ Speech: OpenAI, Gemini TTS, ElevenLabs, Cartesia, Deepgram shipped (`MediaProtocol.stream`, `Speech.generate/stream`, promise `ai.speech`). ✅ Transcription: OpenAI, Gemini, Deepgram, ElevenLabs Scribe, AssemblyAI shipped across all three route kinds (`Transcription.generate/stream/start/resume`, promise `ai.transcription`). Deferred: `Speech.session` and `Transcription.session` (WebSocket streaming).
4.**Image queued routes and partials** — ✅ BFL, fal, Replicate, and Stability creative upscale queued; Stability generate inline; OpenAI `partial_images` streaming (`image-partial` restored). Imagen dropped: shut down on the Gemini API and discontinued on Vertex (2026-06-30). Deferred: Stability's synchronous edit and fast/conservative upscale endpoints.
Loaded 100 of 4748 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.