Compare commits

...
Author SHA1 Message Date
Aiden Cline 8f3cf164bd fix(core): share affinity in provider session headers 2026-09-28 14:37:07 -05:00
James Long 46e53e3f2b fix(ai): expose evaluation confidence (#51930) 2026-09-28 15:24:21 -04:00
Aiden Cline 6cf442b545 fix(core): share child session affinity headers (#51923) 2026-09-28 13:31:17 -05:00
opencode-agent[bot]andrekram1-node f20f5b68ee chore(release): announce V2 releases in Discord (#51880)
Co-authored-by: rekram1-node <rekram1-node@users.noreply.github.com>
2026-09-28 12:30:55 -05:00
SebastianandOpenCode Agent 96dd9f77a9 fix(core): attribute one-shot generation requests (#48358)
Co-authored-by: OpenCode Agent <opencode-agent[bot]@users.noreply.github.com>
2026-09-28 11:55:14 -05:00
opencode-agent[bot] dd786c62af chore(core): refresh bundled models.dev snapshot 2026-09-28 12:23:50 +00:00
Frank 87c402a124 docs(go): simplify v2 usage limit explanation 2026-09-28 07:39:42 -04:00
Frank 7076a878a4 docs(www): document Go Plus (#51834) 2026-09-28 07:25:37 -04:00
Frank 45b91eed82 docs(go): sync v2 model list with v1 (#51837) 2026-09-28 11:23:40 +00:00
Niels Kootstra 39e1ce55bc fix(core): reject relative path segments in repository hosts (#51577) 2026-09-28 10:14:35 +05:30
opencode-agent[bot]andrekram1-node d9f54392ba fix(tui): distinguish background shell from interrupted command (#51769)
Co-authored-by: rekram1-node <rekram1-node@users.noreply.github.com>
2026-09-27 23:01:00 -05:00
Aiden Cline d73396ab3d fix(ai): preserve Gemini thought signatures on OpenAI Chat tool calls (#51768) 2026-09-27 22:57:22 -05:00
Jérôme BenoitandTest User 0caae608a2 chore(nix): update nixpkgs for Bun 1.4 (#50221)
Co-authored-by: Test User <test@test.com>
2026-09-27 21:34:52 -05:00
Kit Langton 96f23508be refactor(core): remove unused project discovery option (#51729) 2026-09-27 15:05:46 -07:00
Kit Langton 28bb0a7158 refactor(core): drop forwarding shim modules (#51670) 2026-09-27 14:47:19 -07:00
DS 3d109828ff fix(tui): truncate btw question preview (#51713) 2026-09-27 21:05:20 +02:00
Shoubhit Dash c0d49f101c feat(ai): retry transient failures on queued generation reads (#51635) 2026-09-27 19:47:25 +05:30
Shoubhit Dash be2446e188 feat(ai): add ElevenLabs Scribe transcription route (#51641) 2026-09-27 19:34:36 +05:30
Shoubhit Dash 107966eddd fix(ai): reject and throw with signal.reason on abort (#51633) 2026-09-27 19:24:45 +05:30
Kit Langton f5e580cde1 chore(core): remove dead modules and exports (#51667) 2026-09-27 06:38:43 -07:00
Shoubhit Dash 4428a77acd fix(ai): classify terminal generation failures by provider error code (#51632) 2026-09-27 18:44:21 +05:30
91 changed files with 1895 additions and 807 deletions
-5
View File
@@ -1,5 +0,0 @@
---
"@opencode/core": patch
---
Correct directory page headings when the read offset is zero.
+20
View File
@@ -24,6 +24,10 @@ on:
description: "Override version (optional)"
required: false
type: string
release_notes:
description: "Reviewed V2 release notes for the Discord announcement (optional)"
required: false
type: string
concurrency: ${{ github.workflow }}-${{ github.ref }}-${{ (github.ref_name == 'v2' && (inputs.version || inputs.bump) && 'release') || inputs.version || inputs.bump }}
@@ -653,3 +657,19 @@ jobs:
OPENCODE_DESKTOP_DIST: /tmp/desktop
CLOUDFLARE_ACCOUNT_ID: 15d29c8639fd3733b1b5486a2acfd968
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
notify-discord-v2:
needs:
- version
- publish
if: github.repository == 'anomalyco/opencode' && github.ref_name == 'v2' && needs.version.outputs.release && needs.publish.result == 'success'
runs-on: blacksmith-4vcpu-ubuntu-2404
steps:
# Unlike dev, V2 publishes a tag rather than a GitHub Release event.
- name: Announce V2 release in Discord
uses: SethCohen/github-releases-to-discord@24d166886aee4646d448c8a389ff9e1ebcab3682 # v1.20.0
with:
webhook_url: ${{ secrets.DISCORD_WEBHOOK }}
release_name: OpenCode V2 ${{ needs.version.outputs.tag }}
release_body: ${{ inputs.release_notes }}
release_html_url: https://github.com/${{ github.repository }}/tree/${{ needs.version.outputs.tag }}
Generated
+3 -3
View File
@@ -2,11 +2,11 @@
"nodes": {
"nixpkgs": {
"locked": {
"lastModified": 1776683584,
"narHash": "sha256-NuTLMrr10Tng72hurYG8jYQ4XKK8wnpJmOGcPiis96g=",
"lastModified": 1790510107,
"narHash": "sha256-EVMNYv7hYDDD9TGVT/hIyTYgpiXA8y3m5xIEIxuGNU0=",
"owner": "NixOS",
"repo": "nixpkgs",
"rev": "9dd5558b06dbdacbf635a3dd36dce1b1a7ee3a89",
"rev": "3181085bfd08663b6b9e60bc7a8395c2aaa741bd",
"type": "github"
},
"original": {
+2 -2
View File
@@ -98,11 +98,11 @@ When a provider supports multiple physical transports, selection remains executi
Media does not fit the SSE-frames-to-event-state-machine LLM route. `MediaRoute.inline(...)` / `queued(...)` / `stream(...)` (`src/route/media.ts`) compose a `MediaProtocol` kind with `Endpoint` and `Auth` and own the transport plumbing: `http` option merging, URL/query rendering, auth headers, JSON vs multipart encoding, and handing the response back to the protocol. `MediaProtocol.inline` (`src/route/media-protocol.ts`) is `body.from(request)` plus `response.decode(response, context)`; each protocol declares `const route = MediaProtocol.identity({ id, name, provider })` once and decodes through `route.decodeJson` / `route.text` / `route.decodeStarted` so decode failures retain the raw body and HTTP context, raising `route.unsupported(operation, message)` for requests it cannot lower, and passes `route` as the first argument to `MediaProtocol.inline` / `queued` / `stream`. `Generation` (`src/generation.ts`) is the provider-neutral handle for a queued generation over a `GenerationRoute` (`status`, `result`, `cancel`). Image protocol files follow the same section order as LLM protocols and declare unsupported common fields once through the protocol's `unsupported` list.
`MediaProtocol.queued` is the submit-then-poll kind every video route uses: `start` (body + decode into `{ token, snapshot }`), `status`, `result`, and optional `cancel` (with `activeOnly` when the provider's cancel endpoint deletes finished work, as Runway's does: the route refreshes status first and skips terminal generations), each addressed by a route-owned `token` whose `Schema.Codec` makes it serializable. `MediaRoute.inline` and `MediaRoute.queued` compose the two kinds with `Endpoint` and `Auth`; the queued route decodes the token once at the boundary (`start` output or `resume` input) and closes over it in a token-free `GenerationRoute` (`status`/`result`/`cancel` are plain Effects), so `Generation` never sees the token's shape and only carries the encoded JSON for persistence. Polls reuse the route's auth and deployment headers plus the request's `http` overlay after `start`, and resolve relative paths against the route base URL (provider-issued absolute URLs such as fal's `status_url` pass through). `result` is always its own GET even when the provider returns output inside the status document, so `Generation.await` behaves the same after `start` and after `resume`. `PollContext.auth` carries only what `Auth` added or changed so protocols can hand download credentials to output assets as transient `Media.Asset.headers` (Veo) — never part of `source` or JSON. Status strings map through a per-protocol `STATUS` table via `MediaProtocol.status`; terminal generations without output fail through `output.ended` / `output.contentPolicy` with the provider document on `reason.body`. `GenerationAwaitOptions` (`AwaitOptions` in `src/generation.ts`, `{ poll?: Poll }`) is the one options type for `await`, `events`, `Video.generate`, and `Video.stream`.
`MediaProtocol.queued` is the submit-then-poll kind every video route uses: `start` (body + decode into `{ token, snapshot }`), `status`, `result`, and optional `cancel` (with `activeOnly` when the provider's cancel endpoint deletes finished work, as Runway's does: the route refreshes status first and skips terminal generations), each addressed by a route-owned `token` whose `Schema.Codec` makes it serializable. `MediaRoute.inline` and `MediaRoute.queued` compose the two kinds with `Endpoint` and `Auth`; the queued route decodes the token once at the boundary (`start` output or `resume` input) and closes over it in a token-free `GenerationRoute` (`status`/`result`/`cancel` are plain Effects), so `Generation` never sees the token's shape and only carries the encoded JSON for persistence. Polls reuse the route's auth and deployment headers plus the request's `http` overlay after `start`, and resolve relative paths against the route base URL (provider-issued absolute URLs such as fal's `status_url` pass through). `result` is always its own GET even when the provider returns output inside the status document, so `Generation.await` behaves the same after `start` and after `resume`. `PollContext.auth` carries only what `Auth` added or changed so protocols can hand download credentials to output assets as transient `Media.Asset.headers` (Veo) — never part of `source` or JSON. Status strings map through a per-protocol `STATUS` table via `MediaProtocol.status`; terminal generations without output fail through `output.ended` / `output.contentPolicy` with the provider document on `reason.body`; a `failed` generation maps the provider's error code through a per-protocol `FAILURE` table via `MediaProtocol.failure` so rejected inputs are not reported as retryable `ProviderInternal`. `GenerationAwaitOptions` (`AwaitOptions` in `src/generation.ts`, `{ poll?: Poll }`) is the one options type for `await`, `events`, `Video.generate`, and `Video.stream`.
`MediaProtocol.stream` is the incremental kind every speech route uses, with the same discipline as LLM protocols. `MediaRoute.stream` submits the caller's request as `MediaProtocol.Addressed<Request>` (`{ ...request, mode }`, `mode: "generate" | "stream"`), so one provider stays one protocol: `body.from`, the endpoint path, and `frames` read `request.mode` to pick the body, path, and framing. `frames(bytes, context)` returns frames — `Framing.sse`, `Framing.lines`, `Framing.document` (a single-document response shaped like a streamed record), or the raw `bytes` for chunked audio. `initial()` is fresh per-response parser state; `step` folds each frame into it and emits modality events; `finish(state, context)` runs once after the last frame with the request, body, and observed `http` (header-only usage lives there) and emits exactly one terminal event or fails with `route.incomplete()`. Keep parser state to real accumulators and derive anything the request or body determines in `finish`. `generate` runs the same stream and folds it with the modality's `collect`. Request-derived URL parameters go on the body's `query` (array values repeat the parameter), applied before route and caller `http.query`. Decode frames with `route.decodeFrame` and raise stream-time failures with `route.frameError` (the frame stays on `reason.body`); protocols never thread HTTP context, because the route fills `reason.http` on stream errors that lack it. Speech protocols share `protocols/utils/speech-stream.ts` for deltas, timestamps, voice ids, PCM and container descriptions, and the terminal asset.
Every modality route is the inline | stream | queued union (transcription uses all three: OpenAI and Gemini stream, Deepgram is inline, AssemblyAI is queued), every client is `MediaClient.make(Service, { modality, responseEvents })` (`src/media-client.ts`), which dispatches on the route's `kind`, and every model composes through `composeRoute`. fal queue protocols come from `protocols/utils/fal-queue.ts`, bodies are `json`, `multipart`, or `binary` (a raw upload), and a queued protocol that must upload media before submitting implements `start.prepare` (`MediaProtocol.Prepare`; AssemblyAI `/v2/upload`).
Every modality route is the inline | stream | queued union (transcription uses all three: OpenAI and Gemini stream, Deepgram and ElevenLabs are inline, AssemblyAI is queued), every client is `MediaClient.make(Service, { modality, responseEvents })` (`src/media-client.ts`), which dispatches on the route's `kind`, and every model composes through `composeRoute`. fal queue protocols come from `protocols/utils/fal-queue.ts`, bodies are `json`, `multipart`, or `binary` (a raw upload), and a queued protocol that must upload media before submitting implements `start.prepare` (`MediaProtocol.Prepare`; AssemblyAI `/v2/upload`).
### URL Construction
+24 -12
View File
@@ -129,8 +129,9 @@ VercelAIGateway.configure().experimental.evaluation("typesafe-ai/jev")
OpenRouter reads `OPENROUTER_API_KEY`. Vercel reads `AI_GATEWAY_API_KEY`, then `VERCEL_OIDC_TOKEN`.
The common API uses `boolean`; System One routes lower it to native `noul`.
Choice and score confidence plus score legends remain available in provider metadata, and the
provider's rounded probabilities are returned unchanged.
Choice and score answers include `confidence` when the provider returns it, such as
`response.answers.department.confidence`. Score legends remain available in provider metadata, and
the provider's rounded probabilities are returned unchanged.
## Alibaba Cloud Model Studio
@@ -752,7 +753,10 @@ const events = Video.stream({ model: Runway.configure({ apiKey }).video("gen4.5"
Status polls, result fetches, cancels, and asset downloads all run through the same request executor with the route's
auth. `Generation.await` and `Generation.events` fail with a
`Timeout` reason when `poll.timeout` (default 10 minutes) elapses. Failed,
`Timeout` reason when `poll.timeout` (default 10 minutes) elapses. Status polls and result fetches retry transient
failures (rate limits, provider 5xx, network errors) with backoff that honors `retry-after`, always within
`poll.timeout`; submits and cancels never retry. Interrupting a wait (or aborting its `signal`) does not cancel the
provider job, which keeps running and billing: call `cancel()` to stop it. Failed,
cancelled, and expired generations fail typed with the provider's terminal document on `reason.body`; moderation
outcomes (Veo `raiMediaFilteredReasons`, xAI `respect_moderation`, Runway `SAFETY.*` codes) surface as `notices` when
a video is still returned and as a `ContentPolicy` reason when nothing is.
@@ -773,7 +777,9 @@ Provider notes:
The promise client exposes the same surface: `ai.video.start(...)` resolves to a handle with `await`, `events`,
`result`, `refresh`, `cancel`, and `token`; `ai.video.generate`, `ai.video.resume(model, token)`, and
`ai.video.stream` mirror the Effect API. The handle's `status` and `progress` are a snapshot from when it was
created; `refresh()` resolves to a new handle.
created; `refresh()` resolves to a new handle. Every promise method and stream accepts `{ signal }`: like `fetch`,
aborting rejects the Promise or throws from the `for await` loop with `signal.reason` (an `AbortError` `DOMException`
unless `abort(reason)` passed one), while `break` stops a stream without throwing.
```ts
import { ai } from "@opencode/ai/promise"
@@ -871,11 +877,12 @@ for await (const event of ai.speech.stream({ model, text: "Hello from OpenCode."
## Transcription
Transcription (speech-to-text) is the one modality whose providers use every route kind: OpenAI and Gemini stream,
Deepgram answers inline, and AssemblyAI is queued. `Transcription.generate` and `Transcription.stream` work on all of
them; `Transcription.start` / `resume` return a `Generation` on queued routes and fail with `UnsupportedOperation`
elsewhere. Models come from `.transcription(...)` selectors on the `OpenAI`, `Google`, `Deepgram`, and `AssemblyAI`
facades. Common fields (`language`, `prompt`, `timestamps: "none" | "segment" | "word"`, `diarize`, `speakers`) lower
natively or fail with a typed `AIError` before any network call; a route may return more than asked.
Deepgram and ElevenLabs answer inline, and AssemblyAI is queued. `Transcription.generate` and `Transcription.stream`
work on all of them; `Transcription.start` / `resume` return a `Generation` on queued routes and fail with
`UnsupportedOperation` elsewhere. Models come from `.transcription(...)` selectors on the `OpenAI`, `Google`,
`Deepgram`, `ElevenLabs`, and `AssemblyAI` facades. Common fields (`language`, `prompt`,
`timestamps: "none" | "segment" | "word"`, `diarize`, `speakers`) lower natively or fail with a typed `AIError` before
any network call; a route may return more than asked.
```ts
import { Console, Effect, Stream } from "effect"
@@ -887,7 +894,7 @@ const openai = OpenAI.configure({ apiKey: process.env.OPENAI_API_KEY })
const program = Effect.gen(function* () {
const audio = yield* Media.file("./call.mp3")
// Speaker-labelled segments; labels are provider-native strings ("A", "0", "spk:0").
// Speaker-labelled segments; labels are provider-native strings ("A", "0", "spk:0", "speaker_0").
const response = yield* Transcription.generate({
model: Deepgram.configure({ apiKey }).transcription("nova-3"),
audio,
@@ -897,7 +904,7 @@ const program = Effect.gen(function* () {
response.text // "Hello from OpenCode."
response.segments // [{ text, startSeconds, endSeconds, speaker: "0" }]
response.words // [{ text, startSeconds, endSeconds, speaker, confidence }]
response.language // the provider's own value, lowercased ("en", "english", "en_us")
response.language // the provider's own value, lowercased ("en", "eng", "english", "en_us")
// Text deltas as the model transcribes, then one finish carrying the whole transcript.
yield* Transcription.stream({ model: openai.transcription("gpt-4o-mini-transcribe"), audio }).pipe(
@@ -921,7 +928,12 @@ Provider notes:
- **OpenAI** takes inline audio only; `diarize` needs `gpt-4o-transcribe-diarize`, timestamps need `whisper-1`, and `whisper-1` does not stream.
- **Gemini** needs a transcribe model (`gemini-3.5-transcribe`); `prompt` and `speakers` fail typed.
- **Deepgram** detects the language unless `language` is set; vocabulary goes in `providerOptions.keyterm`.
- **AssemblyAI** uploads inline audio before submitting and is the only route that accepts `speakers`.
- **ElevenLabs** (`scribe_v2`) uploads inline audio as the multipart `file` and sends a URL as `source_url`. Words
always carry timestamps, and segments are speaker turns, so `diarize`, `timestamps: "segment"`, or `speakers` turns
on diarization. `speakers` is an upper bound (`num_speakers`); `prompt` fails typed (vocabulary goes in
`providerOptions.keyterms`), as do webhook delivery and per-channel output (`use_multi_channel` without
`multichannel_output_style: "combined"`).
- **AssemblyAI** uploads inline audio before submitting and treats `speakers` as the exact speaker count.
The promise client mirrors the Effect API:
+27 -17
View File
@@ -1,7 +1,6 @@
# Media generation in `@opencode/ai` — public API direction
Status: phases 1–4 implemented (through Image queued routes and partial images; ElevenLabs Scribe transcription
pending); phase 5 proposal.
Status: phases 1–4 implemented (through Image queued routes and partial images); phase 5 proposal.
## Goal
@@ -271,8 +270,8 @@ Deferred: `Speech.session(...)` — input-streaming TTS where text arrives incre
#### Transcription (STT)
Shipped as the second half of phase 3 (`src/transcription.ts`, `src/transcription-client.ts`, protocols
`openai-transcription`, `google-transcription`, `deepgram-transcription`, `assemblyai-transcription`; new `AssemblyAI`
facade).
`openai-transcription`, `google-transcription`, `deepgram-transcription`, `elevenlabs-transcription`,
`assemblyai-transcription`; new `AssemblyAI` facade).
```ts
const request = Transcription.request({
@@ -281,7 +280,7 @@ const request = Transcription.request({
language: "en", // provider-native passthrough
timestamps: "segment", // none | segment | word
diarize: true,
speakers: 2, // exact speaker count (AssemblyAI only)
speakers: 2, // speaker count (AssemblyAI exact, ElevenLabs maximum)
providerOptions: { known_speaker_names: ["agent"] },
})
@@ -309,17 +308,23 @@ upload); `packages/ai/AGENTS.md` (Media Routes) describes both.
Settled rules:
- **Timestamps.** A granularity the selected route or model cannot produce fails as `UnsupportedOperation`
(`media.timestamps`), following Speech; a route that returns more than asked (Deepgram and AssemblyAI always return
words) is not stripped. Segments always carry start and end times: Gemini times each transcription part from its
(`media.timestamps`), following Speech; a route that returns more than asked (Deepgram, ElevenLabs, and AssemblyAI
always return words) is not stripped. Segments always carry start and end times: Gemini times each transcription part from its
word offsets, so segment timestamps and diarization also request word offsets there.
- **Diarization.** `diarize` means segments (and words, where the provider labels them) carry `speaker`. Labels are
provider-native strings — OpenAI `A` or a known speaker name, Deepgram `0`, Gemini `spk:0`, AssemblyAI `A` — with no
cross-provider speaker model. `speakers` is the exact number of speakers to label, which AssemblyAI (`speakers_expected`, the only route that
accepts it) treats as a constraint rather than a hint.
provider-native strings — OpenAI `A` or a known speaker name, Deepgram `0`, Gemini `spk:0`, AssemblyAI `A`,
ElevenLabs `speaker_0` — with no cross-provider speaker model. `speakers` is the number of speakers to label:
AssemblyAI (`speakers_expected`) treats it as an exact constraint rather than a hint, and ElevenLabs
(`num_speakers`) as the maximum. Both turn on diarization for it; the other routes reject it.
- **Segments from words.** ElevenLabs returns only a token list (`word`, `spacing`, `audio_event`), so its segments
are speaker turns: consecutive words and spacing with one `speaker_id`, text joined from the provider's own spacing
tokens. `words` drops spacing and audio events. Segments therefore need diarization, which `timestamps: "segment"`
turns on, as AssemblyAI's utterances need speaker labels.
- **Language** is passed through (`language`, OpenAI `gpt-transcribe` `languages[]`, Gemini `languageCodes`,
AssemblyAI `language_code`). `response.language` is the provider's own value, lowercased but not normalized: an
ISO code on most routes (AssemblyAI's detection returns `en`), `english` from whisper-1. Deepgram and AssemblyAI
assume English unless asked to detect, so a missing `language` enables their detection.
AssemblyAI and ElevenLabs `language_code`). `response.language` is the provider's own value, lowercased but not
normalized: an ISO code on most routes (AssemblyAI's detection returns `en`, ElevenLabs ISO 639-3 `eng`), `english`
from whisper-1. Deepgram and AssemblyAI assume English unless asked to detect, so a missing `language` enables their
detection.
- **Gemini** requires a transcribe model; other model ids fail with `UnsupportedOperation` before the call, because
general models ignore `audioTranscriptionConfig` and answer conversationally. Streamed chunks carry whole speaker
turns (one part per turn), which join with a space.
@@ -332,11 +337,12 @@ Settled rules:
| OpenAI | stream (`stream: true` in `stream` mode; `whisper-1` ignores `stream`, so it emits only `finish`) | multipart `file` (inline only) | `whisper-1` (`verbose_json`); diarize model: `segment` | `gpt-4o-transcribe-diarize` (`diarized_json`) | `speakers`; `prompt` on the diarize model | `tokens` or `seconds` |
| Gemini | stream (`generateContent` / `streamGenerateContent`) | `inlineData` or Gemini Files `fileData` | `audioTranscriptionConfig.wordTimestamp` | `audioTranscriptionConfig.diarization` | `prompt`, `speakers` | `tokens` |
| Deepgram | inline | raw body, or JSON `{ url }` | words always; `segment` → `utterances` | `diarize_model=latest` + `utterances` | `prompt`, `speakers` | `seconds` (`metadata.duration`) |
| ElevenLabs | inline | multipart `file`, or `source_url` | words always; `segment` → `diarize` (speaker turns) | `diarize` | `prompt`; `webhook`, per-channel `use_multi_channel` | `seconds` (`audio_duration_secs`) |
| AssemblyAI | queued (upload → submit → poll) | `/v2/upload` then `audio_url`, or a URL | words always; `segment` → `speaker_labels` | `speaker_labels` | — | `seconds` (`audio_duration`) |
Deferred: `Transcription.session(...)` — realtime STT over WebSocket (Deepgram live, AssemblyAI streaming, ElevenLabs
realtime, OpenAI realtime transcription) — is the same future scoped `session` shape as input-streaming TTS and ships
with the realtime work in phase 5. ElevenLabs Scribe is not implemented yet.
with the realtime work in phase 5.
### `Generation` — shared async execution
@@ -361,6 +367,10 @@ Poll = { interval?: Duration; timeout?: Duration }
`Generation` is not video-specific. Image routes on BFL, fal, Replicate, and Stability `upscale()` are queued; `Image.start` exists for them. A route declares itself `inline` or `queued`; `generate` on a queued route is `start` then `await`.
Status polls and result reads retry transient failures (rate limits, provider 5xx, and transport errors, classified by the same `isRetryable` the Session runner uses) inside `MediaRoute.queued`. Only the HTTP exchange retries, never the decoded document: a terminal `failed` generation also surfaces as `ProviderInternal` and must not be re-read. Gaps grow exponentially from 1s with jitter, up to 30s each, honoring a provider `retry-after` up to that cap, for at most 8 retries. `await`, `events`, and `Video.stream` cut retries off at `poll.timeout` and fail with `Timeout`, so retries never extend the caller's deadline; a direct `result()` or `resume` read is bounded by the retry cap alone. `start` and `cancel` never retry: a repeated submit can start and bill a second job. The policy is internal; there is no option for it.
Interrupting `await`, `events`, or `Video.stream` (or aborting the promise API's `signal`) stops waiting only. The provider job keeps running and billing; call `cancel()` explicitly to stop it.
### Usage
```ts
@@ -402,7 +412,7 @@ for await (const event of ai.llm.stream(request)) { … }
await ai.dispose()
```
Streams become `AsyncIterable` via `Stream.toAsyncIterable`. `AIError` is thrown as-is. `AbortSignal` maps to interruption. Nothing in `src/*` except this entrypoint knows about promises.
Streams become `AsyncIterable` via `Stream.toAsyncIterable`. `AIError` is thrown as-is. Aborting an `AbortSignal` interrupts the work and, like `fetch`, rejects the Promise or throws from the stream with `signal.reason` instead of ending the stream as if complete. Nothing in `src/*` except this entrypoint knows about promises.
### Providers
@@ -414,7 +424,7 @@ implemented):
| `OpenAI` | responses (default), chat | Images API (stream) | *Sora skipped (decision 8)* | ✓ | ✓ | |
| `Google` | Gemini | Gemini-native | Veo | Gemini TTS | `gemini-3.5-transcribe` | |
| `XAI` | ✓ | ✓ | ✓ | | | |
| `ElevenLabs` | | | | ✓ | *Scribe (pending)* | *soundEffect, music (phase 5)* |
| `ElevenLabs` | | | | ✓ | Scribe | *soundEffect, music (phase 5)* |
| `Cartesia` | | | | ✓ | | |
| `Deepgram` | | | | Aura | ✓ | |
| `Fal` | | ✓ (queued) | ✓ | | | |
@@ -467,7 +477,7 @@ Foundation + Image ship together as the reference implementation, serially. Vide
1. **Foundation** — per-modality selectors, `Media`, `Generation`, `Poll`, `Usage` union, `MediaProtocol` kinds, `@opencode/ai/promise` with `llm` + `image`. Port the five existing image protocols onto it. Unify `MediaPart` and add the `media` LLM event (fixes Gemini image output being dropped).
2. **Video** — ✅ Veo, xAI, fal, Runway shipped (`MediaProtocol.queued`, `Video.start/generate/resume/stream`, promise `ai.video`). Deferred: `Video.complete` (webhooks), Luma, Kling, MiniMax, Replicate.
3. **Speech + Transcription** — ✅ Speech: OpenAI, Gemini TTS, ElevenLabs, Cartesia, Deepgram shipped (`MediaProtocol.stream`, `Speech.generate/stream`, promise `ai.speech`). ✅ Transcription: OpenAI, Gemini, Deepgram, AssemblyAI shipped across all three route kinds (`Transcription.generate/stream/start/resume`, promise `ai.transcription`). Pending: ElevenLabs Scribe. Deferred: `Speech.session` and `Transcription.session` (WebSocket streaming).
3. **Speech + Transcription** — ✅ Speech: OpenAI, Gemini TTS, ElevenLabs, Cartesia, Deepgram shipped (`MediaProtocol.stream`, `Speech.generate/stream`, promise `ai.speech`). ✅ Transcription: OpenAI, Gemini, Deepgram, ElevenLabs Scribe, AssemblyAI shipped across all three route kinds (`Transcription.generate/stream/start/resume`, promise `ai.transcription`). Deferred: `Speech.session` and `Transcription.session` (WebSocket streaming).
4. **Image queued routes and partials** — ✅ BFL, fal, Replicate, and Stability creative upscale queued; Stability generate inline; OpenAI `partial_images` streaming (`image-partial` restored). Imagen dropped: shut down on the Gemini API and discontinued on Vertex (2026-06-30). Deferred: Stability's synchronous edit and fast/conservative upscale endpoints.
5. **Later** — ElevenLabs music/SFX, Lyria, `Speech.session` / `Transcription.session`, realtime.
@@ -63,6 +63,7 @@ export const ChoiceAnswer = Schema.Struct({
type: Schema.Literal("choice"),
choice: Schema.String,
probabilities: Schema.optional(Schema.Record(Schema.String, Probability)),
confidence: Schema.optional(Probability),
})
export type ChoiceAnswer = Schema.Schema.Type<typeof ChoiceAnswer>
@@ -70,6 +71,7 @@ export const ScoreAnswer = Schema.Struct({
type: Schema.Literal("score"),
score: Schema.Number,
probabilities: Schema.optional(Schema.Record(Schema.String, Probability)),
confidence: Schema.optional(Probability),
})
export type ScoreAnswer = Schema.Schema.Type<typeof ScoreAnswer>
@@ -92,6 +94,7 @@ export type AnswerFor<Question extends EvaluationQuestion> = Question extends {
readonly type: "choice"
readonly choice: Extract<keyof Criteria, string>
readonly probabilities?: Readonly<Record<Extract<keyof Criteria, string>, number>>
readonly confidence?: number
}
: Question extends { readonly type: "score" }
? ScoreAnswer
+10 -5
View File
@@ -142,32 +142,37 @@ export const model = <Options extends EvaluationOptions = EvaluationOptions>(cfg
Effect.mapError((cause) => fail("System One returned an invalid response", cause, text)),
)
const confidence: Record<string, number> = {}
const legend: Record<string, Record<string, Schema.Json>> = {}
const answers = Object.fromEntries(
Object.entries(data.answers).map(([id, answer]): [string, EvaluationAnswer] => {
if (answer.type === "noul") return [id, { type: "boolean", probability: answer.noul }]
if (answer.type === "choice") {
if (answer.confidence !== undefined) confidence[id] = answer.confidence
return [
id,
{
type: "choice",
choice: answer.choice,
probabilities: answer.probabilities,
...(answer.confidence === undefined ? {} : { confidence: answer.confidence }),
},
]
}
if (answer.confidence !== undefined) confidence[id] = answer.confidence
if (answer.legend !== undefined) legend[id] = answer.legend
return [id, { type: "score", score: answer.score, probabilities: answer.probabilities }]
return [
id,
{
type: "score",
score: answer.score,
probabilities: answer.probabilities,
...(answer.confidence === undefined ? {} : { confidence: answer.confidence }),
},
]
}),
)
const meta = {
...(data.id === undefined ? {} : { responseId: data.id }),
...(data.provider === undefined ? {} : { provider: data.provider }),
...data.provider_metadata?.[cfg.providerMetadataKey],
...(Object.keys(confidence).length === 0 ? {} : { confidence }),
...(Object.keys(legend).length === 0 ? {} : { legend }),
}
return new EvaluationResponse({
+47 -28
View File
@@ -102,7 +102,7 @@ export class Generation<Response> {
return settled.pipe(
// Non-completed terminal states also go through `result` so the route can surface its provider failure body.
Effect.flatMap((generation) => generation.result()),
Effect.timeoutOrElse({ duration: timeout, orElse: () => this.timeoutError(timeout) }),
Effect.timeoutOrElse({ duration: timeout, orElse: () => timeoutError(this.id, timeout) }),
)
}
@@ -123,20 +123,7 @@ export class Generation<Response> {
Clock.currentTimeMillis.pipe(
Effect.map((start) => {
const deadline = start + Duration.toMillis(timeout)
// Fail before polling once the deadline has passed: a fast status request could otherwise win the zero-budget
// race and schedule another zero-delay poll.
const refresh = Clock.currentTimeMillis.pipe(
Effect.flatMap((now) =>
now >= deadline
? this.timeoutError(timeout)
: this.refresh().pipe(
Effect.timeoutOrElse({
duration: Duration.millis(deadline - now),
orElse: () => this.timeoutError(timeout),
}),
),
),
)
const refresh = within(this.refresh(), this.id, timeout, deadline)
const schedule = this.schedule(options?.poll).pipe(
Schedule.modifyDelay((meta) =>
Effect.succeed(Duration.min(meta.duration, Duration.millis(Math.max(0, deadline - meta.now)))),
@@ -157,15 +144,6 @@ export class Generation<Response> {
return { type: "generation-progress", id: this.id, progress: this.progress }
}
private timeoutError(timeout: Duration.Duration) {
return new AIError({
reason: new TimeoutError({
message: `Generation ${this.id} did not finish within ${Duration.format(timeout)}`,
timeoutMs: Duration.toMillis(timeout),
}),
})
}
private poll(poll: Poll | undefined) {
return this.refresh().pipe(
Effect.repeat({ schedule: this.schedule(poll), until: (generation) => generation.terminal }),
@@ -177,12 +155,53 @@ export class Generation<Response> {
}
}
/** `events` followed by the expanded result, with the result fetch bounded by the same `poll.timeout` deadline. */
export const resultEvents = <Response, A>(
generation: Generation<Response>,
expand: (response: Response) => ReadonlyArray<A>,
options?: AwaitOptions,
): Stream.Stream<Observation | A, AIError> =>
generation.events(options).pipe(
Stream.filter((event): event is Observation => event.type !== "generation-finished"),
Stream.concat(Stream.fromIterableEffect(Effect.map(generation.result(), expand))),
): Stream.Stream<Observation | A, AIError> => {
const timeout = Duration.fromInputUnsafe(options?.poll?.timeout ?? DEFAULT_POLL_TIMEOUT)
return Stream.unwrap(
Clock.currentTimeMillis.pipe(
Effect.map((start) =>
generation.events(options).pipe(
Stream.filter((event): event is Observation => event.type !== "generation-finished"),
Stream.concat(
Stream.fromIterableEffect(
within(generation.result(), generation.id, timeout, start + Duration.toMillis(timeout)).pipe(
Effect.map(expand),
),
),
),
),
),
),
)
}
/**
* Run `effect` within the time left until `deadline`. Fails before starting once the deadline has passed: a fast
* request could otherwise win the zero-budget race and schedule another zero-delay poll.
*/
const within = <A>(effect: Effect.Effect<A, AIError>, id: string, timeout: Duration.Duration, deadline: number) =>
Clock.currentTimeMillis.pipe(
Effect.flatMap((now) =>
now >= deadline
? Effect.fail(timeoutError(id, timeout))
: effect.pipe(
Effect.timeoutOrElse({
duration: Duration.millis(deadline - now),
orElse: () => Effect.fail(timeoutError(id, timeout)),
}),
),
),
)
const timeoutError = (id: string, timeout: Duration.Duration) =>
new AIError({
reason: new TimeoutError({
message: `Generation ${id} did not finish within ${Duration.format(timeout)}`,
timeoutMs: Duration.toMillis(timeout),
}),
})
+1 -1
View File
@@ -4,7 +4,7 @@ export { ImageClient } from "./image-client.js"
export { Auth } from "./route/auth.js"
export { Provider } from "./provider.js"
export { ProviderPackage } from "./provider-package.js"
export { isContextOverflow, isContextOverflowFailure } from "./provider-error.js"
export { isContextOverflow, isContextOverflowFailure, isRetryable } from "./provider-error.js"
export type {
RouteLanguageModelInput,
RouteRoutedLanguageModelInput,
+7 -6
View File
@@ -42,7 +42,7 @@ export type GenerationHandle<Response> = Snapshot & {
/** Serializable JSON; pass it back to `resume` from another process. */
readonly token: unknown
readonly await: (options?: AwaitOptions & RunOptions) => Promise<Response>
/** Status observations until the first terminal one, polling like `await`; abort ends iteration without throwing. */
/** Status observations until the first terminal one, polling like `await`; abort throws `signal.reason`. */
readonly events: (options?: AwaitOptions & RunOptions) => AsyncIterable<Event>
/** The result without polling; fails when the generation has not completed. */
readonly result: (options?: RunOptions) => Promise<Response>
@@ -50,15 +50,16 @@ export type GenerationHandle<Response> = Snapshot & {
readonly cancel: (options?: RunOptions) => Promise<void>
}
// Fails with `signal.reason` so aborted calls reject and aborted streams throw like `fetch`: an `AbortError` by default.
const abortEffect = (signal: AbortSignal | undefined) =>
signal === undefined
? Effect.never
: Effect.callback<void>((resume) => {
: Effect.callback<never, unknown>((resume) => {
if (signal.aborted) {
resume(Effect.void)
resume(Effect.fail(signal.reason))
return
}
const onAbort = () => resume(Effect.void)
const onAbort = () => resume(Effect.fail(signal.reason))
signal.addEventListener("abort", onAbort, { once: true })
return Effect.sync(() => signal.removeEventListener("abort", onAbort))
})
@@ -68,14 +69,14 @@ export const make = (options: Options = {}) => {
/** Run any package Effect (for example `LLMClient.compact(...)`) inside this runtime. */
const run = <A, E>(effect: Effect.Effect<A, E, Services>, options?: RunOptions) =>
runtime.runPromise(effect, { signal: options?.signal })
runtime.runPromise(Effect.raceFirst(effect, abortEffect(options?.signal)))
const iterate = <A, E>(stream: Stream.Stream<A, E, Services>, options?: RunOptions): AsyncIterable<A> =>
Stream.toAsyncIterable(
Stream.unwrap(
runtime.contextEffect.pipe(
Effect.map(
(context): Stream.Stream<A, E> =>
(context): Stream.Stream<A, unknown> =>
stream.pipe(Stream.interruptWhen(abortEffect(options?.signal)), Stream.provideContext(context)),
),
),
@@ -6,6 +6,7 @@ import { mergeJsonRecords, type OpenString } from "../schema/index.js"
import { TranscriptionModel, TranscriptionResponse, type TranscriptionRequestFor } from "../transcription.js"
import { ProviderShared } from "./shared.js"
import { MediaInput } from "./utils/media-input.js"
import { SpeakerTurns } from "./utils/speaker-turns.js"
const route = MediaProtocol.identity({ id: "deepgram-transcription", name: "Deepgram", provider: "deepgram" })
export const DEFAULT_BASE_URL = "https://api.deepgram.com"
@@ -115,16 +116,6 @@ const speaker = (value: number | undefined) => (value === undefined ? undefined
const wordText = (word: typeof Word.Type) => word.punctuated_word ?? word.word
// Utterances split on pauses, not speakers: the v2 diarizer labels a whole utterance with one speaker even when its
// words change speaker, so segments split each utterance at speaker changes.
const speakerTurns = (words: ReadonlyArray<typeof Word.Type>) =>
words.reduce<Array<Array<typeof Word.Type>>>((turns, word) => {
const last = turns.at(-1)
if (last === undefined || last[0].speaker !== word.speaker) return [...turns, [word]]
last.push(word)
return turns
}, [])
const decodeResponse = Effect.fn("DeepgramTranscription.decodeResponse")(function* (
response: HttpClientResponse.HttpClientResponse,
) {
@@ -136,6 +127,8 @@ const decodeResponse = Effect.fn("DeepgramTranscription.decodeResponse")(functio
const requestID = output.value.metadata?.request_id
return new TranscriptionResponse({
text: alternative.transcript,
// Utterances split on pauses, not speakers: the v2 diarizer labels a whole utterance with one speaker even when
// its words change speaker, so segments split each utterance at speaker changes.
segments: output.value.results.utterances?.flatMap((utterance) =>
utterance.words === undefined || utterance.words.length === 0
? [
@@ -146,7 +139,7 @@ const decodeResponse = Effect.fn("DeepgramTranscription.decodeResponse")(functio
speaker: speaker(utterance.speaker),
},
]
: speakerTurns(utterance.words).map((turn) => ({
: SpeakerTurns.group(utterance.words, (word) => word.speaker).map((turn) => ({
text: turn.map(wordText).join(" "),
startSeconds: turn[0].start,
endSeconds: turn[turn.length - 1].end,
@@ -0,0 +1,211 @@
import { Effect, Schema } from "effect"
import type { HttpClientResponse } from "effect/unstable/http"
import { MediaProtocol } from "../route/media-protocol.js"
import { MediaRoute } from "../route/media.js"
import { mergeJsonRecords, type OpenString } from "../schema/index.js"
import { TranscriptionModel, TranscriptionResponse, type TranscriptionRequestFor } from "../transcription.js"
import { mediaTypeExtension } from "../utils/media-type.js"
import { ProviderShared, optionalNull } from "./shared.js"
import { MediaInput } from "./utils/media-input.js"
import { SpeakerTurns } from "./utils/speaker-turns.js"
const route = MediaProtocol.identity({
id: "elevenlabs-transcription",
name: "ElevenLabs Transcription",
provider: "elevenlabs",
})
export const DEFAULT_BASE_URL = "https://api.elevenlabs.io"
export const PATH = "/v1/speech-to-text"
// ---------------------------------------------------------------------------
// 1. Public model input
// ---------------------------------------------------------------------------
export type ElevenLabsTranscriptionOptions = {
readonly tag_audio_events?: boolean
readonly timestamps_granularity?: OpenString<"none" | "word" | "character">
readonly diarization_threshold?: number
readonly file_format?: OpenString<"pcm_s16le_16" | "other">
readonly temperature?: number
readonly seed?: number
readonly keyterms?: ReadonlyArray<string>
readonly no_verbatim?: boolean
readonly detect_speaker_roles?: boolean
readonly use_speaker_library?: boolean
readonly entity_detection?: string | ReadonlyArray<string>
readonly entity_redaction?: string | ReadonlyArray<string>
readonly entity_redaction_mode?: OpenString<"redacted" | "entity_type" | "enumerated_entity_type">
} & Record<string, unknown>
export type Request = TranscriptionRequestFor<ElevenLabsTranscriptionOptions>
// ---------------------------------------------------------------------------
// 2. Response schema
// ---------------------------------------------------------------------------
/** `type` is `word`, `spacing` (the whitespace between words), or `audio_event` (`(laughter)`). */
const Token = Schema.Struct({
text: Schema.String,
type: Schema.String,
start: optionalNull(Schema.Number),
end: optionalNull(Schema.Number),
speaker_id: optionalNull(Schema.String),
logprob: optionalNull(Schema.Number),
})
type Token = Schema.Schema.Type<typeof Token>
const Transcript = Schema.Struct({
language_code: optionalNull(Schema.String),
text: Schema.String,
words: optionalNull(Schema.Array(Token)),
transcription_id: optionalNull(Schema.String),
audio_duration_secs: optionalNull(Schema.Number),
})
// ---------------------------------------------------------------------------
// 5. Request body construction
// ---------------------------------------------------------------------------
/** Speaker turns are the only segments ElevenLabs can produce, and `num_speakers` only applies to diarization. */
const diarizes = (request: Request) =>
request.diarize === true || request.timestamps === "segment" || request.speakers !== undefined
const RESERVED_FORM_FIELDS = new Set([
"file",
"cloud_storage_url",
"source_url",
"model_id",
"language_code",
"diarize",
"num_speakers",
])
const validate = (request: Request, overlay: Record<string, unknown>) => {
// Webhook requests return 202 with no transcript; the result arrives at a configured webhook instead.
if (overlay.webhook === true)
return Effect.fail(route.unsupported("transcription.webhook", `${route.name} does not deliver to webhooks`))
// Separate multichannel output replaces the transcript with one transcript per channel.
if (overlay.use_multi_channel === true && overlay.multichannel_output_style !== "combined")
return Effect.fail(
route.unsupported(
"transcription.multichannel",
`${route.name} returns a single transcript; set multichannel_output_style: "combined" to merge channels`,
),
)
if (overlay.timestamps_granularity === "none" && (request.timestamps === "word" || diarizes(request)))
return Effect.fail(
route.unsupported(
"media.timestamps",
`${route.name} cannot return word timestamps or speaker turns with timestamps_granularity: "none"`,
),
)
return Effect.void
}
const fromRequest = Effect.fn("ElevenLabsTranscription.fromRequest")(function* (request: Request) {
const overlay = mergeJsonRecords(request.providerOptions, request.http?.body) ?? {}
yield* validate(request, overlay)
const form = new FormData()
const url = ProviderShared.mediaUrl(request.audio)
if (url === undefined) {
const extension = mediaTypeExtension(request.audio.mediaType)
const audio = yield* MediaInput.inlineBytes(route.id, request.audio)
form.append(
"file",
MediaInput.blob(audio, request.audio.mediaType),
extension === undefined ? "audio" : `audio.${extension}`,
)
}
MediaInput.appendFields(
form,
{
model_id: request.model.id,
// `cloud_storage_url` is deprecated in favor of `source_url`, which accepts any hosted audio or video URL.
source_url: url,
language_code: request.language,
diarize: diarizes(request) ? true : undefined,
num_speakers: request.speakers,
},
{ overlay, reserved: RESERVED_FORM_FIELDS, repeatArrays: "key" },
)
return MediaProtocol.multipart(form)
})
// ---------------------------------------------------------------------------
// 6. Response decoding
// ---------------------------------------------------------------------------
const decodeTranscript = route.decodeJson(Transcript)
type TimedWord = Token & { readonly start: number; readonly end: number }
const isTimedWord = (token: Token): token is TimedWord =>
token.type === "word" && typeof token.start === "number" && typeof token.end === "number"
/** Turn text keeps the provider's own spacing tokens, so languages written without spaces are not re-spaced. */
const speakerTurns = (tokens: ReadonlyArray<Token>) =>
SpeakerTurns.group(
tokens.filter((token) => token.type === "word" || token.type === "spacing"),
(token) => token.speaker_id,
).flatMap((turn) => {
const words = turn.filter(isTimedWord)
if (words.length === 0) return []
return [
{
text: turn
.map((token) => token.text)
.join("")
.trim(),
startSeconds: words[0].start,
endSeconds: words[words.length - 1].end,
speaker: turn[0].speaker_id ?? undefined,
},
]
})
const decodeResponse = Effect.fn("ElevenLabsTranscription.decodeResponse")(function* (
response: HttpClientResponse.HttpClientResponse,
context: MediaProtocol.DecodeContext<Request>,
) {
const output = yield* decodeTranscript(response)
const transcript = output.value
const tokens = transcript.words ?? []
const duration = transcript.audio_duration_secs ?? undefined
const transcriptionID = transcript.transcription_id ?? undefined
return new TranscriptionResponse({
text: transcript.text,
segments: diarizes(context.request) ? speakerTurns(tokens) : undefined,
words: tokens.filter(isTimedWord).map((word) => ({
text: word.text,
startSeconds: word.start,
endSeconds: word.end,
speaker: word.speaker_id ?? undefined,
confidence: typeof word.logprob === "number" ? Math.exp(word.logprob) : undefined,
})),
language: transcript.language_code?.toLowerCase(),
durationSeconds: duration,
usage: duration === undefined ? undefined : { type: "seconds", seconds: duration },
providerMetadata: transcriptionID === undefined ? undefined : { elevenlabs: { transcriptionId: transcriptionID } },
})
})
// ---------------------------------------------------------------------------
// 7. Protocol and route
// ---------------------------------------------------------------------------
export const protocol = MediaProtocol.inline<Request, TranscriptionResponse>(route, {
unsupported: ["prompt"],
body: { from: fromRequest },
response: { decode: decodeResponse },
})
export const model = (input: MediaRoute.ModelInput) =>
TranscriptionModel.fromRoute<ElevenLabsTranscriptionOptions>(
{ protocol, baseURL: DEFAULT_BASE_URL, path: PATH },
input,
)
export const ElevenLabsTranscription = {
protocol,
model,
} as const
+14 -1
View File
@@ -36,7 +36,9 @@ const StartResponse = Schema.Struct({ name: Schema.String })
const Operation = Schema.Struct({
done: Schema.optional(Schema.Boolean),
error: Schema.optional(Schema.Struct({ message: Schema.optional(Schema.String) })),
error: Schema.optional(
Schema.Struct({ code: Schema.optional(Schema.Number), message: Schema.optional(Schema.String) }),
),
response: Schema.optional(
Schema.Struct({
generateVideoResponse: Schema.optional(
@@ -60,6 +62,16 @@ const Operation = Schema.Struct({
metadata: Schema.optional(Schema.Unknown),
})
// Operation errors are `google.rpc.Status`; unlisted codes (INTERNAL, UNAVAILABLE, ...) are provider-side.
const FAILURE = {
3: "InvalidRequest", // INVALID_ARGUMENT
7: "Authentication", // PERMISSION_DENIED
8: "RateLimit", // RESOURCE_EXHAUSTED
9: "InvalidRequest", // FAILED_PRECONDITION
11: "InvalidRequest", // OUT_OF_RANGE
16: "Authentication", // UNAUTHENTICATED
} as const satisfies Record<number, MediaProtocol.Failure>
// ---------------------------------------------------------------------------
// 5. Request body construction
// ---------------------------------------------------------------------------
@@ -154,6 +166,7 @@ const decodeResult = Effect.fn("GoogleVideo.decodeResult")(function* (
return yield* output.ended(
"failed",
`${route.name} operation failed${operation.error?.message === undefined ? "" : `: ${operation.error.message}`}`,
MediaProtocol.failure(FAILURE, operation.error?.code),
)
const generated = operation.response?.generateVideoResponse
// Downloads require the same API key as the poll; the asset carries it transiently and follows the redirect.
+26 -9
View File
@@ -64,6 +64,14 @@ const OpenAIChatTool = Schema.Struct({
})
type OpenAIChatTool = Schema.Schema.Type<typeof OpenAIChatTool>
// Gemini's OpenAI-compatible surface carries thought signatures in tool call
// `extra_content` and rejects replayed parallel calls without them:
// https://ai.google.dev/gemini-api/docs/thinking#signatures
const ExtraContent = Schema.Struct({
google: Schema.Struct({ thought_signature: Schema.String }),
})
const decodeExtraContent = (value: unknown) => Option.getOrUndefined(Schema.decodeUnknownOption(ExtraContent)(value))
const OpenAIChatAssistantToolCall = Schema.Struct({
id: Schema.String,
type: Schema.tag("function"),
@@ -71,6 +79,7 @@ const OpenAIChatAssistantToolCall = Schema.Struct({
name: Schema.String,
arguments: Schema.String,
}),
extra_content: Schema.optional(ExtraContent),
})
type OpenAIChatAssistantToolCall = Schema.Schema.Type<typeof OpenAIChatAssistantToolCall>
@@ -112,12 +121,6 @@ const decodeReasoningDetail = Schema.decodeUnknownOption(ReasoningDetail)
const knownReasoningDetails = (details: ReadonlyArray<unknown>) =>
details.flatMap((detail) => Option.toArray(decodeReasoningDetail(detail)))
// Intentionally omit Gemini's provider-specific `extra_content.google.thought_signature`
// extension until direct Google OpenAI-compatible routing is supported here:
// https://github.com/vercel/ai/issues/11590
// https://github.com/vercel/ai/pull/11745
// https://ai.google.dev/gemini-api/docs/thought-signatures#openai
const OpenAIChatUserContent = Schema.Union([
Schema.Struct({
type: Schema.Literal("text"),
@@ -242,6 +245,7 @@ const OpenAIChatToolCallDelta = Schema.Struct({
index: optionalNull(Schema.Number),
id: optionalNull(Schema.String),
function: optionalNull(OpenAIChatToolCallDeltaFunction),
extra_content: optionalNull(Schema.Unknown),
})
type OpenAIChatToolCallDelta = Schema.Schema.Type<typeof OpenAIChatToolCallDelta>
@@ -294,6 +298,7 @@ interface PendingToolDelta {
readonly id?: string
readonly name?: string
readonly input: string
readonly extraContent?: Schema.Schema.Type<typeof ExtraContent>
}
export interface ParserState {
@@ -347,13 +352,17 @@ const lowerToolChoice = (toolChoice: NonNullable<LLMRequest["toolChoice"]>) =>
tool: (name) => ({ type: "function" as const, function: { name } }),
})
const lowerToolCall = (part: ToolCallPart, options: LoweringOptions): OpenAIChatAssistantToolCall => ({
const lowerToolCall = (
part: ToolCallPart,
options: LoweringOptions & { readonly providerMetadataKey: string },
): OpenAIChatAssistantToolCall => ({
id: options.toolCallID?.(part.id) ?? part.id,
type: "function",
function: {
name: part.name,
arguments: ProviderShared.encodeJson(part.input === undefined ? {} : part.input),
},
extra_content: decodeExtraContent(part.providerMetadata?.[options.providerMetadataKey]?.extraContent),
})
const lowerMedia = Effect.fn("OpenAIChat.lowerMedia")(function* (part: MediaPart) {
@@ -721,7 +730,9 @@ const detectSupportsStore = (provider: string, baseURL: string | undefined): boo
p === "vercel-ai-gateway" || url.includes("ai-gateway.vercel.sh") || url.includes("vercel.sh")
const isAntLing = p === "ant-ling" || url.includes("api.ant-ling.com")
const isOpencode = p === "opencode" || url.includes("opencode.ai")
const isGemini = url.includes("generativelanguage.googleapis.com")
const isNonStandard =
isGemini ||
isNvidia ||
isCerebras ||
isXai ||
@@ -1114,12 +1125,13 @@ const step = (state: ParserState, event: OpenAIChatEvent) =>
const id = current?.id ?? pending?.id ?? (tool.id || undefined)
const name = current?.name ?? pending?.name ?? (tool.function?.name || undefined)
const text = `${pending?.input ?? ""}${tool.function?.arguments ?? ""}`
const extraContent = pending?.extraContent ?? decodeExtraContent(tool.extra_content)
latestToolIndex = index
nextToolIndex = Math.max(nextToolIndex, index + 1)
if (!current && (!id || !name)) {
pendingTools = {
...pendingTools,
[index]: { id: id || undefined, name: name || undefined, input: text },
[index]: { id: id || undefined, name: name || undefined, input: text, extraContent },
}
continue
}
@@ -1131,7 +1143,12 @@ const step = (state: ParserState, event: OpenAIChatEvent) =>
ADAPTER,
tools,
index,
{ id: id || undefined, name: name || undefined, text },
{
id: id || undefined,
name: name || undefined,
text,
providerMetadata: extraContent && { [state.providerMetadataKey]: { extraContent } },
},
"OpenAI Chat tool call delta is missing id or name",
)
if (ToolStream.isError(result))
@@ -193,7 +193,7 @@ const fromRequest = Effect.fn("OpenAITranscription.fromRequest")(function* (requ
{
overlay: mergeJsonRecords(request.providerOptions, request.http?.body),
reserved: RESERVED_FORM_FIELDS,
repeatArrays: true,
repeatArrays: "key[]",
},
)
return MediaProtocol.multipart(form)
+6 -1
View File
@@ -137,7 +137,12 @@ const decodeResult = Effect.fn("RunwayVideo.decodeResult")(function* (
const message = `${route.name} task failed${code === undefined ? "" : ` (${code})`}${task.failure ? `: ${task.failure}` : ""}`
// Runway failure codes are dotted paths; every moderation outcome carries a SAFETY segment.
if (code !== undefined && /(^|\.)SAFETY(\.|$)/.test(code)) return yield* output.contentPolicy(message)
return yield* output.ended("failed", message)
// ASSET.INVALID rejects the caller's input media; Runway documents it as not retryable.
return yield* output.ended(
"failed",
message,
code !== undefined && /^ASSET\.INVALID(\.|$)/.test(code) ? "InvalidRequest" : "ProviderInternal",
)
}
if (status === "cancelled")
return yield* output.ended("cancelled", `${route.name} task ${context.token.taskID} was cancelled`)
@@ -71,8 +71,9 @@ export const imageOutput = (
}
/**
* Append multipart text fields: strings as-is, other values as JSON, or arrays as repeated `key[]` parts with
* `repeatArrays`. `overlay` keys in `reserved` are dropped so `http.body` cannot replace route-owned fields.
* Append multipart text fields: strings as-is, other values as JSON, or scalar arrays as one part per item with
* `repeatArrays`, named `key[]` or `key`. `overlay` keys in `reserved` are dropped so `http.body` cannot replace
* route-owned fields.
*/
export const appendFields = (
form: FormData,
@@ -80,13 +81,13 @@ export const appendFields = (
options: {
readonly overlay?: Record<string, unknown>
readonly reserved: ReadonlySet<string>
readonly repeatArrays?: true
readonly repeatArrays?: "key[]" | "key"
},
) => {
const overlay = Object.entries(options.overlay ?? {}).filter(([key]) => !options.reserved.has(key))
Object.entries(mergeJsonRecords(fields, Object.fromEntries(overlay)) ?? {}).forEach(([key, value]) => {
if (Array.isArray(value) && options.repeatArrays)
return value.forEach((item) => form.append(`${key}[]`, String(item)))
if (Array.isArray(value) && value.every(isScalar) && options.repeatArrays !== undefined)
return value.forEach((item) => form.append(options.repeatArrays === "key[]" ? `${key}[]` : key, String(item)))
form.append(key, typeof value === "string" ? value : encodeJson(value))
})
}
@@ -0,0 +1,10 @@
/** Split an ordered token list into runs of consecutive tokens with the same speaker. */
export const group = <Item>(items: ReadonlyArray<Item>, speaker: (item: Item) => unknown) =>
items.reduce<Array<Array<Item>>>((turns, item) => {
const last = turns.at(-1)
if (last === undefined || speaker(last[0]) !== speaker(item)) return [...turns, [item]]
last.push(item)
return turns
}, [])
export * as SpeakerTurns from "./speaker-turns.js"
@@ -147,7 +147,12 @@ export const appendOrStart = <K extends StreamKey>(
route: string,
tools: State<K>,
key: K,
delta: { readonly id?: string; readonly name?: string; readonly text: string },
delta: {
readonly id?: string
readonly name?: string
readonly text: string
readonly providerMetadata?: ProviderMetadata
},
missingToolMessage: string,
): AppendOutcome<K> | AIError => {
const current = tools[key]
@@ -161,7 +166,7 @@ export const appendOrStart = <K extends StreamKey>(
namespace: current?.namespace,
input: `${current?.input ?? ""}${delta.text}`,
providerExecuted: current?.providerExecuted,
providerMetadata: current?.providerMetadata,
providerMetadata: current?.providerMetadata ?? delta.providerMetadata,
}
if (current && delta.text.length === 0 && current.id === id && current.name === name)
return { tools, tool: current, events: [] }
+8
View File
@@ -65,6 +65,13 @@ const STATUS = {
expired: "expired",
} as const satisfies Record<string, Status>
// Documented video error codes; `service_unavailable`, `internal_error`, and unknown codes are provider-side.
const FAILURE = {
invalid_argument: "InvalidRequest",
failed_precondition: "InvalidRequest",
permission_denied: "Authentication",
} as const satisfies Record<string, MediaProtocol.Failure>
// ---------------------------------------------------------------------------
// 5. Request body construction
// ---------------------------------------------------------------------------
@@ -143,6 +150,7 @@ const decodeResult = Effect.fn("XAIVideo.decodeResult")(function* (
return yield* output.ended(
"failed",
`${route.name} generation failed${code === undefined ? "" : ` (${code})`}${message === undefined ? "" : `: ${message}`}`,
MediaProtocol.failure(FAILURE, code),
)
}
if (status !== "completed")
+41
View File
@@ -58,6 +58,47 @@ export const isContextOverflowFailure = (failure: unknown) =>
? failure.reason._tag === "InvalidRequest" && failure.reason.classification === "context-overflow"
: Schema.is(ProviderErrorEvent)(failure) && failure.classification === "context-overflow"
/**
* Whether a failed call may succeed when sent again: rate limits, provider-side failures, transport failures that did
* not deliver an accepted write, and unrecognized failures. Callers decide which calls are safe to repeat.
*/
export const isRetryable = (error: AIError) => {
const override = error.reason.http?.headers["x-should-retry"]
if (override === "true") return true
if (override === "false") return false
switch (error.reason._tag) {
case "RateLimit":
case "ProviderInternal":
return true
// A WebSocket acknowledgment marks delivery accepted before model output may exist.
// Read failures can still recover; the caller chooses retry versus continuation from durable output.
case "Transport":
return (
error.reason.delivery !== "rejected" &&
(error.reason.delivery !== "accepted" || error.reason.operation === "read")
)
case "InvalidProviderOutput":
return error.reason.classification === "incomplete-stream"
// Unrecognized failures retry: classification records affirmative
// deterministic evidence, and transient failures are exactly the ones
// that arrive in shapes no classifier anticipates.
case "UnknownProvider":
return true
case "Authentication":
case "QuotaExceeded":
case "ContentPolicy":
case "InvalidRequest":
case "UnsupportedOperation":
case "NoRoute":
case "Timeout":
return false
default: {
const exhaustive: never = error.reason
return exhaustive
}
}
}
const decodeJson = Schema.decodeUnknownOption(Schema.fromJsonString(Schema.Unknown))
// OpenCode Zen reports account caps as typed 429/402 errors that are not throttles.
const QUOTA_CODES = new Set([
+5
View File
@@ -3,8 +3,10 @@ import type { ProviderAuthOption } from "../route/auth-options.js"
import { MediaRoute } from "../route/media.js"
import { type HttpOptions, ProviderID, type ModelID } from "../schema/index.js"
import { ElevenLabsSpeech } from "../protocols/elevenlabs-speech.js"
import { ElevenLabsTranscription } from "../protocols/elevenlabs-transcription.js"
export type { ElevenLabsOutputFormat, ElevenLabsSpeechOptions } from "../protocols/elevenlabs-speech.js"
export type { ElevenLabsTranscriptionOptions } from "../protocols/elevenlabs-transcription.js"
export const id = ProviderID.make("elevenlabs")
@@ -24,12 +26,15 @@ const auth = (options: ProviderAuthOption<"optional">) => {
export const configure = (input: Config = {}) => {
const media = MediaRoute.deployment(input, auth(input))
const speech = (modelID: string | ModelID) => ElevenLabsSpeech.model({ ...media, id: modelID })
const transcription = (modelID: string | ModelID) => ElevenLabsTranscription.model({ ...media, id: modelID })
return {
id,
speech,
transcription,
configure,
}
}
export const provider = configure()
export const speech = provider.speech
export const transcription = provider.transcription
+26 -5
View File
@@ -5,12 +5,14 @@ import { Media } from "../media.js"
import type { AuthInput } from "./auth.js"
import {
AIError,
AuthenticationError,
ContentPolicyError,
HttpContext,
InvalidProviderOutputError,
InvalidRequestError,
ProviderID,
ProviderInternalError,
RateLimitError,
UnsupportedOperationError,
} from "../schema/index.js"
@@ -188,6 +190,16 @@ export const stream = <Request, Event, Frame, State>(
// Response helpers
// ---------------------------------------------------------------------------
/** Reasons a provider can report for a `failed` generation; anything it does not classify is `ProviderInternal`. */
const FAILURES = {
InvalidRequest: InvalidRequestError,
Authentication: AuthenticationError,
RateLimit: RateLimitError,
ProviderInternal: ProviderInternalError,
}
export type Failure = keyof typeof FAILURES
const context = (response: HttpClientResponse.HttpClientResponse) =>
new HttpContext({ url: response.request.url, status: response.status, headers: response.headers })
@@ -199,9 +211,10 @@ export const identity = (input: { readonly id: string; readonly name: string; re
/**
* Read a text body while retaining the original payload and HTTP context on every downstream error. `invalid` is a
* malformed provider document; `ended` is a generation that reached a terminal status without output (`failed` is
* provider-side, `cancelled`/`expired` mean the result will never exist); `pending` is a `result()` read before the
* generation finished, which is caller misuse; `contentPolicy` is a moderated result.
* malformed provider document; `ended` is a generation that reached a terminal status without output (`failed`
* carries the provider's classification, defaulting to `ProviderInternal`; `cancelled`/`expired` mean the result
* will never exist); `pending` is a `result()` read before the generation finished, which is caller misuse;
* `contentPolicy` is a moderated result.
*/
const text = Effect.fn("MediaProtocol.text")(function* (response: HttpClientResponse.HttpClientResponse) {
const http = context(response)
@@ -223,11 +236,15 @@ export const identity = (input: { readonly id: string; readonly name: string; re
http,
invalid: (message: string, cause?: unknown) =>
new AIError({ reason: new InvalidProviderOutputError({ route: input.id, message, body, http, cause }) }),
ended: (status: Exclude<Status, "queued" | "running" | "completed">, message: string) =>
ended: (
status: Exclude<Status, "queued" | "running" | "completed">,
message: string,
failure: Failure = "ProviderInternal",
) =>
new AIError({
reason:
status === "failed"
? new ProviderInternalError({ message, body, http })
? new FAILURES[failure]({ message, body, http })
: new InvalidRequestError({ message, body, http }),
}),
pending: (id: string) =>
@@ -303,6 +320,10 @@ export const status = <Table extends Record<string, Status>>(
return Effect.succeed(table[raw])
}
/** Map a provider error code through the protocol's table; missing or unmapped codes are `ProviderInternal`. */
export const failure = (table: Readonly<Record<string, Failure>>, code: string | number | undefined): Failure =>
code !== undefined && Object.hasOwn(table, code) ? table[code] : "ProviderInternal"
/** A `url` asset whose provider-declared retention window starts now. */
export const expiringUrl = (url: string, retention: Duration.Duration, options?: Parameters<typeof Media.url>[1]) =>
Clock.currentTimeMillis.pipe(
+34 -4
View File
@@ -1,4 +1,4 @@
import { Effect, Schema, Stream } from "effect"
import { Duration, Effect, Schedule, Schema, Stream } from "effect"
import { Headers, HttpClientRequest, type HttpClientResponse } from "effect/unstable/http"
import { Auth, type AuthInput } from "./auth.js"
import { Endpoint } from "./endpoint.js"
@@ -7,6 +7,7 @@ import { RequestExecutor } from "./executor.js"
import { MediaProtocol } from "./media-protocol.js"
import { Generation, isTerminal } from "../generation.js"
import type { Media } from "../media.js"
import { isRetryable } from "../provider-error.js"
import {
AIError,
AIErrorReason,
@@ -137,6 +138,32 @@ export const inline = <Request extends MediaRequest, Response>(
}
}
const READ_RETRY_MAX_DELAY = Duration.seconds(30)
/**
* Status and result reads retry transient failures; `start` and `cancel` never do. Gaps grow exponentially from 1s,
* jittered, up to 30s each, for at most 8 retries (about two minutes when every attempt fails), so a direct
* `Generation.result()` stays bounded; `await` and `events` also cut retries off at `poll.timeout`. A provider
* `retryAfterMs` raises the gap, still capped at 30s.
*/
const READ_RETRY = Schedule.max([
Schedule.min([Schedule.exponential("1 second"), Schedule.spaced(READ_RETRY_MAX_DELAY)]),
Schedule.recurs(8),
]).pipe(
Schedule.jittered,
Schedule.setInputType<AIError>(),
Schedule.modifyDelay(({ input, duration }) =>
Effect.succeed(
Duration.min(
input.reason._tag === "RateLimit" || input.reason._tag === "ProviderInternal"
? Duration.max(duration, Duration.millis(input.reason.retryAfterMs ?? 0))
: duration,
READ_RETRY_MAX_DELAY,
),
),
),
)
/**
* Compose a queued media protocol the same way, adding `start`/`resume` handles whose polls reuse the route's auth,
* deployment headers, and (for `start`) the request's `http` overlay. The token is decoded once at the boundary and
@@ -154,6 +181,8 @@ export const queued = <Request extends MediaRequest, Response, Token>(
const generationRoute = (token: Token, http: HttpOptions | undefined, execute: Execute) => {
const materialize = (asset: Media.Asset) =>
asset.materialize().pipe(Effect.provideService(RequestExecutorService, { execute }))
// Only the GET exchange retries: a decoded terminal failure (`output.ended`) can be a `ProviderInternal` too, and
// re-reading it would spin until the caller's deadline.
const poll = <A>(operation: {
readonly path: (token: Token) => string
readonly decode: (
@@ -161,9 +190,10 @@ export const queued = <Request extends MediaRequest, Response, Token>(
context: MediaProtocol.PollContext<Token>,
) => Effect.Effect<A, AIError>
}) =>
transport
.call("GET", operation.path(token), http, execute)
.pipe(Effect.flatMap((sent) => operation.decode(sent.response, { token, auth: sent.auth, materialize })))
transport.call("GET", operation.path(token), http, execute).pipe(
Effect.retry({ schedule: READ_RETRY, while: isRetryable }),
Effect.flatMap((sent) => operation.decode(sent.response, { token, auth: sent.auth, materialize })),
)
const status = poll(protocol.status)
const cancel = protocol.cancel
const send =
+1 -1
View File
@@ -103,7 +103,7 @@ export type TranscriptionRequestInput<Model extends TranscriptionModel = Transcr
// Response and events
// ---------------------------------------------------------------------------
/** Speaker labels are provider-native (`A`, `0`, `spk:0`, or a known speaker name). */
/** Speaker labels are provider-native (`A`, `0`, `spk:0`, `speaker_0`, or a known speaker name). */
export const TranscriptionSegment = Schema.Struct({
text: Schema.String,
startSeconds: Schema.Number,
+2 -1
View File
@@ -39,17 +39,18 @@ describe("experimental Evaluation", () => {
type: "choice",
choice: "billing",
probabilities: { billing: 0.9, technical: 0.1 },
confidence: 0.8,
})
expect(response.answers.urgency).toEqual({
type: "score",
score: 1.2,
probabilities: { "0": 0, "1": 0.8, "2": 0.2 },
confidence: 0.6,
})
expect(response.answers.refund).toEqual({ type: "boolean", probability: 0.97 })
expect(response.usage?.totalTokens).toBe(36)
expect(response.providerMetadata).toEqual({
typesafe: {
confidence: { department: 0.8, urgency: 0.6 },
legend: { urgency: { "0": "Can wait", "1": "Needs attention", "2": "Blocking" } },
},
})
+2
View File
@@ -26,8 +26,10 @@ const request = Evaluation.request({
const result = EvaluationClient.evaluate(request)
type Result = Success<typeof result>
type Choice = Assert<Equal<Result["answers"]["topic"]["choice"], "billing" | "support">>
type Confidence = Assert<Equal<Result["answers"]["topic"]["confidence"], number | undefined>>
type ClientRequirements = Assert<Equal<Requirements<typeof result>, Service>>
void (true satisfies Choice)
void (true satisfies Confidence)
void (true satisfies ClientRequirements)
Effect.gen(function* () {
+5
View File
@@ -197,6 +197,11 @@ describe("public exports", () => {
expect(Google.configure({ apiKey: "fixture" }).transcription("gemini-3.5-transcribe").route.kind).toBe("stream")
expect(Deepgram.configure({ apiKey: "fixture" }).transcription("nova-3").route.kind).toBe("inline")
expect(AssemblyAI.configure({ apiKey: "fixture" }).transcription("universal-3-5-pro").route.kind).toBe("queued")
expect(ElevenLabs.configure({ apiKey: "fixture" }).transcription("scribe_v2").route.id).toBe(
"elevenlabs-transcription",
)
expect(ElevenLabs.configure({ apiKey: "fixture" }).transcription("scribe_v2").route.kind).toBe("inline")
expect(ElevenLabs.provider.transcription).toBe(ElevenLabs.transcription)
})
test("protocol barrels expose supported low-level routes", () => {
@@ -0,0 +1,32 @@
{
"version": 1,
"metadata": {
"tags": [
"prefix:elevenlabs-transcription",
"provider:elevenlabs",
"protocol:elevenlabs-transcription"
],
"name": "elevenlabs-transcription/groups-diarized-words-into-speaker-turns",
"recordedAt": "2026-09-27T09:35:28.265Z"
},
"interactions": [
{
"transport": "http",
"request": {
"method": "POST",
"url": "https://api.elevenlabs.io/v1/speech-to-text",
"headers": {
"content-type": "multipart/form-data; boundary=----WebKitFormBoundary356bdc14864a477dbacbfcf60d1ecceb"
},
"body": "--BOUNDARY\r\nContent-Disposition: form-data; name=\"file\"; filename=\"audio.mp3\"\r\nContent-Type: audio/mpeg\r\n\r\n[audio]\r\n--BOUNDARY\r\nContent-Disposition: form-data; name=\"model_id\"\r\n\r\nscribe_v2\r\n--BOUNDARY\r\nContent-Disposition: form-data; name=\"diarize\"\r\n\r\ntrue\r\n--BOUNDARY--\r\n"
},
"response": {
"status": 200,
"headers": {
"content-type": "application/json"
},
"body": "{\"language_code\":\"eng\",\"language_probability\":0.9495430588722229,\"text\":\"Did the release ship? Yes, it shipped this morning\",\"words\":[{\"text\":\"Did\",\"start\":0.34,\"end\":0.44,\"type\":\"word\",\"speaker_id\":\"speaker_0\",\"logprob\":-1.7881377516459906e-6},{\"text\":\" \",\"start\":0.44,\"end\":0.48,\"type\":\"spacing\",\"speaker_id\":\"speaker_0\",\"logprob\":-1.1920928244535389e-7},{\"text\":\"the\",\"start\":0.48,\"end\":0.56,\"type\":\"word\",\"speaker_id\":\"speaker_0\",\"logprob\":-1.1920928244535389e-7},{\"text\":\" \",\"start\":0.56,\"end\":0.6,\"type\":\"spacing\",\"speaker_id\":\"speaker_0\",\"logprob\":-7.152531907195225e-6},{\"text\":\"release\",\"start\":0.6,\"end\":0.92,\"type\":\"word\",\"speaker_id\":\"speaker_0\",\"logprob\":-7.152531907195225e-6},{\"text\":\" \",\"start\":0.92,\"end\":0.94,\"type\":\"spacing\",\"speaker_id\":\"speaker_0\",\"logprob\":-8.344646857949556e-7},{\"text\":\"ship?\",\"start\":0.94,\"end\":1.26,\"type\":\"word\",\"speaker_id\":\"speaker_0\",\"logprob\":-7.414704032271402e-6},{\"text\":\" \",\"start\":1.26,\"end\":1.26,\"type\":\"spacing\",\"speaker_id\":\"speaker_0\",\"logprob\":-0.0009363081189803779},{\"text\":\"Yes,\",\"start\":1.68,\"end\":2.02,\"type\":\"word\",\"speaker_id\":\"speaker_1\",\"logprob\":-0.003542040009030245},{\"text\":\" \",\"start\":2.02,\"end\":2.48,\"type\":\"spacing\",\"speaker_id\":\"speaker_1\",\"logprob\":-3.933898824470816e-6},{\"text\":\"it\",\"start\":2.48,\"end\":2.62,\"type\":\"word\",\"speaker_id\":\"speaker_1\",\"logprob\":-3.933898824470816e-6},{\"text\":\" \",\"start\":2.62,\"end\":2.64,\"type\":\"spacing\",\"speaker_id\":\"speaker_1\",\"logprob\":-0.000013589766240329482},{\"text\":\"shipped\",\"start\":2.66,\"end\":2.9,\"type\":\"word\",\"speaker_id\":\"speaker_1\",\"logprob\":-0.000013589766240329482},{\"text\":\" \",\"start\":2.9,\"end\":2.94,\"type\":\"spacing\",\"speaker_id\":\"speaker_1\",\"logprob\":0.0},{\"text\":\"this\",\"start\":2.94,\"end\":3.12,\"type\":\"word\",\"speaker_id\":\"speaker_1\",\"logprob\":0.0},{\"text\":\" \",\"start\":3.12,\"end\":3.18,\"type\":\"spacing\",\"speaker_id\":\"speaker_1\",\"logprob\":-3.576278118089249e-7},{\"text\":\"morning\",\"start\":3.18,\"end\":3.5,\"type\":\"word\",\"speaker_id\":\"speaker_1\",\"logprob\":-3.576278118089249e-7}],\"transcription_id\":\"cs3I2282TH8hjw12brNg\",\"audio_duration_secs\":3.5526875}"
}
}
]
}
@@ -0,0 +1,50 @@
{
"version": 1,
"metadata": {
"tags": [
"prefix:elevenlabs-transcription",
"provider:elevenlabs",
"protocol:elevenlabs-transcription"
],
"name": "elevenlabs-transcription/transcribes-audio-with-word-timestamps",
"recordedAt": "2026-09-27T09:35:27.686Z"
},
"interactions": [
{
"transport": "http",
"request": {
"method": "POST",
"url": "https://api.elevenlabs.io/v1/speech-to-text",
"headers": {
"content-type": "multipart/form-data; boundary=----WebKitFormBoundarye2be7b31e94441bbbeb35a9c890a9d74"
},
"body": "--BOUNDARY\r\nContent-Disposition: form-data; name=\"file\"; filename=\"audio.mp3\"\r\nContent-Type: audio/mpeg\r\n\r\n[audio]\r\n--BOUNDARY\r\nContent-Disposition: form-data; name=\"model_id\"\r\n\r\nscribe_v2\r\n--BOUNDARY--\r\n"
},
"response": {
"status": 200,
"headers": {
"content-type": "application/json"
},
"body": "{\"language_code\":\"eng\",\"language_probability\":0.6618340611457825,\"text\":\"Hello from OpenCode\",\"words\":[{\"text\":\"Hello\",\"start\":0.4,\"end\":0.66,\"type\":\"word\",\"logprob\":-0.000014781842764932662},{\"text\":\" \",\"start\":0.66,\"end\":0.74,\"type\":\"spacing\",\"logprob\":-3.814689989667386e-6},{\"text\":\"from\",\"start\":0.74,\"end\":0.84,\"type\":\"word\",\"logprob\":-3.814689989667386e-6},{\"text\":\" \",\"start\":0.84,\"end\":0.9,\"type\":\"spacing\",\"logprob\":-0.018268775194883347},{\"text\":\"OpenCode\",\"start\":0.9,\"end\":1.44,\"type\":\"word\",\"logprob\":-0.1251817401498556}],\"transcription_id\":\"D4VfnANM2ArCHTujIb9q\",\"audio_duration_secs\":1.54125}"
}
},
{
"transport": "http",
"request": {
"method": "POST",
"url": "https://api.elevenlabs.io/v1/speech-to-text",
"headers": {
"content-type": "multipart/form-data; boundary=----WebKitFormBoundaryfb80d0e44d9e44d299416ed546a04056"
},
"body": "--BOUNDARY\r\nContent-Disposition: form-data; name=\"file\"; filename=\"audio.mp3\"\r\nContent-Type: audio/mpeg\r\n\r\n[audio]\r\n--BOUNDARY\r\nContent-Disposition: form-data; name=\"model_id\"\r\n\r\nscribe_v2\r\n--BOUNDARY--\r\n"
},
"response": {
"status": 200,
"headers": {
"content-type": "application/json"
},
"body": "{\"language_code\":\"eng\",\"language_probability\":0.6618340611457825,\"text\":\"Hello from OpenCode\",\"words\":[{\"text\":\"Hello\",\"start\":0.4,\"end\":0.66,\"type\":\"word\",\"logprob\":-0.000023007127310847864},{\"text\":\" \",\"start\":0.66,\"end\":0.74,\"type\":\"spacing\",\"logprob\":-2.3841830625315197e-6},{\"text\":\"from\",\"start\":0.74,\"end\":0.84,\"type\":\"word\",\"logprob\":-2.3841830625315197e-6},{\"text\":\" \",\"start\":0.84,\"end\":0.9,\"type\":\"spacing\",\"logprob\":-0.008306833915412426},{\"text\":\"OpenCode\",\"start\":0.9,\"end\":1.44,\"type\":\"word\",\"logprob\":-0.1075385226868093}],\"transcription_id\":\"SkYplzfq1DW8Ae3bWnoy\",\"audio_duration_secs\":1.54125}"
}
}
]
}
@@ -0,0 +1,54 @@
{
"version": 1,
"metadata": {
"model": "gemini-3.8-flash",
"tags": [
"prefix:openai-compatible-chat",
"provider:google",
"protocol:openai-chat",
"tool",
"tool-loop",
"continuation"
],
"name": "gemini-parallel-tool-signatures",
"recordedAt": "2026-09-28T03:12:05.083Z"
},
"interactions": [
{
"transport": "http",
"request": {
"method": "POST",
"url": "https://generativelanguage.googleapis.com/v1beta/openai/chat/completions",
"headers": {
"content-type": "application/json"
},
"body": "{\"model\":\"gemini-3.8-flash\",\"messages\":[{\"role\":\"system\",\"content\":\"Call get_weather for every requested city in parallel, then answer in one short sentence.\"},{\"role\":\"user\",\"content\":\"What is the weather in Paris and in Tokyo?\"}],\"tools\":[{\"type\":\"function\",\"function\":{\"name\":\"get_weather\",\"description\":\"Get current weather for a city.\",\"parameters\":{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},\"required\":[\"city\"],\"additionalProperties\":false},\"strict\":false}}],\"stream\":true,\"stream_options\":{\"include_usage\":true}}"
},
"response": {
"status": 200,
"headers": {
"content-type": "text/event-stream"
},
"body": "data: {\"choices\":[{\"delta\":{\"role\":\"assistant\",\"tool_calls\":[{\"extra_content\":{\"google\":{\"thought_signature\":\"ErYCCrMCAWkUfRNGh+/Zbc8YUSzk1yfWAfROcjA4HyF1x69jz6167w8zd4n6kZQQ5FDeBZ5HZMEbEkQ4ENOpzsQL8roCR6wONkhXpiduWrTD6XwbP8KGNkf6D1tX/JlBh7G5Cl+0rdjiSOl/mdY1lcjbkfyCRFs5T8odNWMG7WD3rCrXJDFQ/5QfOl+tqVTceKGz2yyXBhOhvsQDU33ulR9tHQJo/Fmx3HNyDdwvmyKUXm+kgqHsYZnkdv6y6xZwy9zBXGUO4QwxelMw3Rrc24Mp2rNbEDZS3YEeP72Jn/hIP5a8XWOx6+zop/4/CjYxxSfN/tJRY5d48NAKRNFzUN1Az8SvAs/N5X2I3/5ZMaDnQWW/BVdS9fm2KFYJ2baqtBiRRnJxMq4ELArkKjGZzHpd97XG8YjV6g==\"}},\"function\":{\"arguments\":\"{\\\"city\\\":\\\"Paris\\\"}\",\"name\":\"get_weather\"},\"id\":\"call_723181\",\"type\":\"function\"}]},\"index\":0}],\"created\":1790565123,\"id\":\"A9u5aoufBp3rz7IPke_3oAo\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":16,\"prompt_tokens\":74,\"total_tokens\":132}}\n\ndata: {\"choices\":[{\"delta\":{\"role\":\"assistant\",\"tool_calls\":[{\"function\":{\"arguments\":\"{\\\"city\\\":\\\"Tokyo\\\"}\",\"name\":\"get_weather\"},\"id\":\"call_723184\",\"type\":\"function\"}]},\"index\":0}],\"created\":1790565123,\"id\":\"A9u5aoufBp3rz7IPke_3oAo\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":32,\"prompt_tokens\":74,\"total_tokens\":148}}\n\ndata: {\"choices\":[{\"delta\":{\"role\":\"assistant\"},\"finish_reason\":\"stop\",\"index\":0}],\"created\":1790565124,\"id\":\"A9u5aoufBp3rz7IPke_3oAo\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":32,\"prompt_tokens\":74,\"total_tokens\":148}}\n\ndata: [DONE]\n\n"
}
},
{
"transport": "http",
"request": {
"method": "POST",
"url": "https://generativelanguage.googleapis.com/v1beta/openai/chat/completions",
"headers": {
"content-type": "application/json"
},
"body": "{\"model\":\"gemini-3.8-flash\",\"messages\":[{\"role\":\"system\",\"content\":\"Call get_weather for every requested city in parallel, then answer in one short sentence.\"},{\"role\":\"user\",\"content\":\"What is the weather in Paris and in Tokyo?\"},{\"role\":\"assistant\",\"content\":null,\"tool_calls\":[{\"id\":\"call_723181\",\"type\":\"function\",\"function\":{\"name\":\"get_weather\",\"arguments\":\"{\\\"city\\\":\\\"Paris\\\"}\"},\"extra_content\":{\"google\":{\"thought_signature\":\"ErYCCrMCAWkUfRNGh+/Zbc8YUSzk1yfWAfROcjA4HyF1x69jz6167w8zd4n6kZQQ5FDeBZ5HZMEbEkQ4ENOpzsQL8roCR6wONkhXpiduWrTD6XwbP8KGNkf6D1tX/JlBh7G5Cl+0rdjiSOl/mdY1lcjbkfyCRFs5T8odNWMG7WD3rCrXJDFQ/5QfOl+tqVTceKGz2yyXBhOhvsQDU33ulR9tHQJo/Fmx3HNyDdwvmyKUXm+kgqHsYZnkdv6y6xZwy9zBXGUO4QwxelMw3Rrc24Mp2rNbEDZS3YEeP72Jn/hIP5a8XWOx6+zop/4/CjYxxSfN/tJRY5d48NAKRNFzUN1Az8SvAs/N5X2I3/5ZMaDnQWW/BVdS9fm2KFYJ2baqtBiRRnJxMq4ELArkKjGZzHpd97XG8YjV6g==\"}}},{\"id\":\"call_723184\",\"type\":\"function\",\"function\":{\"name\":\"get_weather\",\"arguments\":\"{\\\"city\\\":\\\"Tokyo\\\"}\"}}]},{\"role\":\"tool\",\"tool_call_id\":\"call_723181\",\"content\":\"{\\\"temperature\\\":22,\\\"condition\\\":\\\"sunny\\\"}\"},{\"role\":\"tool\",\"tool_call_id\":\"call_723184\",\"content\":\"{\\\"temperature\\\":0,\\\"condition\\\":\\\"unknown\\\"}\"}],\"tools\":[{\"type\":\"function\",\"function\":{\"name\":\"get_weather\",\"description\":\"Get current weather for a city.\",\"parameters\":{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},\"required\":[\"city\"],\"additionalProperties\":false},\"strict\":false}}],\"stream\":true,\"stream_options\":{\"include_usage\":true}}"
},
"response": {
"status": 200,
"headers": {
"content-type": "text/event-stream"
},
"body": "data: {\"choices\":[{\"delta\":{\"content\":\"It is currently sunny and 22°C in Paris,\",\"role\":\"assistant\"},\"index\":0}],\"created\":1790565125,\"id\":\"BNu5auWeB-2fz7IPxY6l4QY\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":13,\"prompt_tokens\":150,\"total_tokens\":226}}\n\ndata: {\"choices\":[{\"delta\":{\"content\":\" while Tokyo is 0°C.\",\"role\":\"assistant\"},\"index\":0}],\"created\":1790565125,\"id\":\"BNu5auWeB-2fz7IPxY6l4QY\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":21,\"prompt_tokens\":150,\"total_tokens\":234}}\n\ndata: {\"choices\":[{\"delta\":{\"extra_content\":{\"google\":{\"thought_signature\":\"EpcDCpQDAWkUfRM325TAUtfFiOeQIEWn/TsCU9oi2js4VPHFeeLfAu+2k8PJm0fN/OaF0y4ovau7S9QIAsuOPI2w2aIyQ2kMGj1XvUyRTvz30DOZgtq1km6W6YGzZyyTCNSeBcpwtJtziHZVVWq9xEI/HHB8Ta1Ot215xnFyDL7iUGEwgGu45/mInpk+SOCYBy9biDddpDxcDi14BGoleArY9XEFAzYLxXssl7HMWjpfee5095im7gD125Nripq1Jf3nGY/2TxqjgQAdJpQybwct63p74O1szGHQxrkBt7AwphDgbOWtLpUP/QJBRdl8qhrozqRe611NQ6V5lMSwpO7OhQ/IDRtWMwOyrrKblZfmMnnPl2/9xDfZRsYnfmWq+7PeAptJl1cDRlMBKhj5iRn31xvN43EiuvWwWPsSndiWvxrMvVBd89TR4u0+z0pYCYZcsFkYKFlPA7pZsdnh6CON24AA5WUkO4dQNoqYdik6aO5hE9jLHG5FHTdr0W69qgPNozbxnO0ptRcOtkGpB9xYUVoTxcSXXZE=\"}},\"role\":\"assistant\"},\"finish_reason\":\"stop\",\"index\":0}],\"created\":1790565125,\"id\":\"BNu5auWeB-2fz7IPxY6l4QY\",\"model\":\"gemini-3.8-flash\",\"object\":\"chat.completion.chunk\",\"usage\":{\"completion_tokens\":21,\"prompt_tokens\":192,\"total_tokens\":276}}\n\ndata: [DONE]\n\n"
}
}
]
}
+85 -1
View File
@@ -22,7 +22,8 @@ const chatBody = sseEvents(
/**
* Executor layer that answers chat completions with SSE text, image generations with one base64 PNG, Runway video
* tasks with a queued submission that succeeds on the second poll, speech with raw audio or SSE audio deltas, OpenAI
* transcription with JSON or SSE text deltas, and AssemblyAI transcripts that complete on the first poll.
* transcription with JSON or SSE text deltas, AssemblyAI transcripts that complete on the first poll, and `slow.test`
* chat completions that send one text delta and never finish.
*/
const executor = (seen: Array<string>) =>
RequestExecutor.layer.pipe(
@@ -55,6 +56,18 @@ const executor = (seen: Array<string>) =>
output: "https://replicate.test/a.webp",
urls: { get: "https://replicate.test/p_1", cancel: "https://replicate.test/p_1/cancel" },
})
if (web.url.startsWith("https://slow.test"))
return input.respond(
new ReadableStream({
start: (controller) =>
controller.enqueue(
new TextEncoder().encode(
`data: ${JSON.stringify({ choices: [{ delta: { content: "Hello" } }] })}\n\n`,
),
),
}),
{ headers: { "content-type": "text/event-stream" } },
)
if (web.url.endsWith("/chat/completions"))
return input.respond(chatBody, { headers: { "content-type": "text/event-stream" } })
if (web.url.endsWith("/audio/speech"))
@@ -304,6 +317,77 @@ describe("AI promise client", () => {
await ai.dispose()
})
test("aborted calls reject and aborted streams throw with the signal's reason", async () => {
const ai = AI.make({ layer: executor([]) })
const slow = OpenAI.configure({ apiKey: "test", baseURL: "https://slow.test/v1" }).chat("gpt-4o-mini")
const aborted = new AbortController()
aborted.abort()
const reason = new Error("mine")
const rejected = await ai.run(Effect.never, { signal: aborted.signal }).catch((error: unknown) => error)
expect(rejected).toBe(aborted.signal.reason)
expect(rejected).toMatchObject({ name: "AbortError" })
const inFlight = new AbortController()
setTimeout(() => inFlight.abort(reason), 10)
expect(
await ai.llm
.generate({ model: slow, prompt: "Hello" }, { signal: inFlight.signal })
.catch((error: unknown) => error),
).toBe(reason)
const preAborted = await Array.fromAsync(
ai.speech.stream({ model: openai.speech("gpt-4o-mini-tts"), text: "Hello" }, { signal: aborted.signal }),
).catch((error: unknown) => error)
expect(preAborted).toBe(aborted.signal.reason)
expect(preAborted).toMatchObject({ name: "AbortError" })
const midStream = new AbortController()
const deltas: Array<string> = []
const midStreamFailure = await Array.fromAsync(
ai.llm.stream({ model: slow, prompt: "Hello" }, { signal: midStream.signal }),
(event) => {
if (!LLMEvent.is.textDelta(event)) return
deltas.push(event.text)
midStream.abort()
},
).catch((error: unknown) => error)
expect(deltas).toEqual(["Hello"])
expect(midStreamFailure).toBe(midStream.signal.reason)
expect(midStreamFailure).toMatchObject({ name: "AbortError" })
const model = Runway.configure({ apiKey: "test", baseURL: "https://runway.test/v1" }).video("gen4.5")
const generation = await ai.video.start({ model, prompt: "A kite" })
const polling = new AbortController()
const events: Array<string> = []
const eventsFailure = await Array.fromAsync(
generation.events({ poll: { interval: 60_000 }, signal: polling.signal }),
(event) => {
events.push(event.type)
polling.abort(reason)
},
).catch((error: unknown) => error)
expect(events).toEqual(["generation-progress"])
expect(eventsFailure).toBe(reason)
await ai.dispose()
})
test("breaking out of an abortable stream cleans up without throwing", async () => {
const ai = AI.make({ layer: executor([]) })
const slow = OpenAI.configure({ apiKey: "test", baseURL: "https://slow.test/v1" }).chat("gpt-4o-mini")
const controller = new AbortController()
const deltas: Array<string> = []
for await (const event of ai.llm.stream({ model: slow, prompt: "Hello" }, { signal: controller.signal })) {
if (!LLMEvent.is.textDelta(event)) continue
deltas.push(event.text)
break
}
controller.abort()
expect(deltas).toEqual(["Hello"])
await ai.dispose()
})
test("the default client is created lazily and can be disposed", async () => {
expect(typeof AI.ai.llm.generate).toBe("function")
expect(typeof AI.ai.image.generate).toBe("function")
@@ -0,0 +1,50 @@
import { describe, expect } from "bun:test"
import { Effect, Stream } from "effect"
import { Transcription } from "../../src/index.js"
import { ElevenLabs } from "../../src/providers.js"
import { recordedTests } from "../recorded-test.js"
import { TRANSCRIPT, audio, audioRecording, dialog } from "./transcription-recording.js"
const model = ElevenLabs.configure({ apiKey: process.env.ELEVENLABS_API_KEY ?? "fixture" }).transcription("scribe_v2")
const recorded = recordedTests({
prefix: "elevenlabs-transcription",
provider: "elevenlabs",
protocol: "elevenlabs-transcription",
requires: ["ELEVENLABS_API_KEY"],
options: audioRecording,
})
describe("ElevenLabs Transcription recorded", () => {
recorded.effect("transcribes audio with word timestamps", () =>
Effect.gen(function* () {
const request = Transcription.request({ model, audio: yield* audio, timestamps: "word" })
const response = yield* Transcription.generate(request)
expect(response.text).toMatch(TRANSCRIPT)
expect(response.words?.map((word) => word.text)).toEqual(["Hello", "from", "OpenCode"])
expect(response.words?.every((word) => word.speaker === undefined && (word.confidence ?? 0) > 0)).toBe(true)
expect(response.segments).toBeUndefined()
expect(response.language).toBe("eng")
expect(response.durationSeconds).toBeGreaterThan(0)
expect(response.usage).toEqual({ type: "seconds", seconds: response.durationSeconds })
expect(response.providerMetadata?.elevenlabs?.transcriptionId).toEqual(expect.any(String))
const events = Array.from(yield* Stream.runCollect(Transcription.stream(request)))
expect(events.map((event) => event.type)).toEqual(["finish"])
}),
)
recorded.effect("groups diarized words into speaker turns", () =>
Effect.gen(function* () {
const response = yield* Transcription.generate({ model, audio: yield* dialog, diarize: true })
expect(response.segments?.map((segment) => segment.speaker)).toEqual(["speaker_0", "speaker_1"])
expect(response.segments?.[0].text).toMatch(/^Did the release ship\?$/)
expect(response.segments?.[1].text).toMatch(/^Yes, it shipped this morning\.?$/)
expect(response.segments?.map((segment) => segment.text).join(" ")).toBe(response.text)
expect(response.words?.some((word) => word.text.trim() === "")).toBe(false)
expect(new Set(response.words?.map((word) => word.speaker))).toEqual(new Set(["speaker_0", "speaker_1"]))
}),
)
})
@@ -92,5 +92,6 @@ const assertEvaluation = <Options extends EvaluationOptions>(
expect(response.answers.refund.probability).toBeGreaterThan(0.5)
expect(response.usage?.inputTokens).toBeGreaterThan(0)
expect(response.usage?.outputTokens).toBeGreaterThan(0)
expect(response.providerMetadata?.[metadataKey]?.confidence).toBeDefined()
expect(response.answers.department.confidence).toBeGreaterThan(0)
expect(response.answers.urgency.confidence).toBeGreaterThan(0)
})
@@ -0,0 +1,68 @@
import { describe, expect } from "bun:test"
import { Effect } from "effect"
import { LLM, LLMEvent, LLMRequest, Message, ToolRuntime, toDefinitions } from "../../src/index.js"
import * as OpenAICompatible from "../../src/providers/openai-compatible.js"
import { LLMClient } from "../../src/route.js"
import { compileRequest } from "../../src/route/client.js"
import { recordedTests } from "../recorded-test.js"
import { weatherRuntimeTool, weatherToolName } from "../recorded-scenarios.js"
const model = OpenAICompatible.configure({
provider: "google",
baseURL: "https://generativelanguage.googleapis.com/v1beta/openai",
apiKey: process.env.GOOGLE_GENERATIVE_AI_API_KEY ?? "fixture",
}).model("gemini-3.8-flash")
const recorded = recordedTests({
prefix: "openai-compatible-chat",
provider: "google",
protocol: "openai-chat",
requires: ["GOOGLE_GENERATIVE_AI_API_KEY"],
tags: ["tool", "tool-loop", "continuation"],
metadata: { model: model.id },
})
describe("Gemini OpenAI-compatible Chat recorded", () => {
recorded.effect.with(
"replays thought signatures through a parallel tool loop",
{ cassette: "openai-compatible-chat/gemini-parallel-tool-signatures" },
() =>
Effect.gen(function* () {
const tools = { [weatherToolName]: weatherRuntimeTool }
const request = LLM.request({
model,
system: "Call get_weather for every requested city in parallel, then answer in one short sentence.",
prompt: "What is the weather in Paris and in Tokyo?",
tools: toDefinitions(tools),
cache: "none",
})
const first = yield* LLMClient.generate(request)
const calls = first.events.filter(LLMEvent.is.toolCall)
expect(calls.map((call) => call.input)).toEqual([{ city: "Paris" }, { city: "Tokyo" }])
const extraContent = calls[0]?.providerMetadata?.google?.extraContent
expect(extraContent).toEqual({ google: { thought_signature: expect.any(String) } })
const results = yield* Effect.forEach(calls, (call) => ToolRuntime.dispatch(tools, call))
const continuation = LLMRequest.update(request, {
messages: [
...request.messages,
first.message,
...calls.map((call, index) =>
Message.tool({ id: call.id, name: call.name, result: results[index]!.result }),
),
],
})
const prepared = yield* compileRequest(continuation)
const assistant = prepared.body.messages.find((message) => message.role === "assistant")
expect(assistant?.role === "assistant" ? assistant.tool_calls?.[0]?.extra_content : undefined).toEqual(
extraContent,
)
const second = yield* LLMClient.generate(continuation)
expect(second.events.filter(LLMEvent.is.toolCall)).toHaveLength(0)
expect(second.text).toMatch(/Paris/)
expect(second.text).toMatch(/Tokyo/)
}),
60_000,
)
})
@@ -472,6 +472,45 @@ describe("OpenAI Chat route", () => {
}),
)
it.effect("replays Gemini thought signatures as tool call extra content", () =>
Effect.gen(function* () {
const prepared = yield* compileRequest(
LLM.request({
model,
messages: [
Message.user("Weather in Paris and Tokyo?"),
Message.assistant([
ToolCallPart.make({
id: "call_1",
name: "lookup",
input: { city: "Paris" },
providerMetadata: { openai: { extraContent: { google: { thought_signature: "sig_1" } } } },
}),
ToolCallPart.make({ id: "call_2", name: "lookup", input: { city: "Tokyo" } }),
]),
Message.tool({ id: "call_1", name: "lookup", result: "Sunny" }),
Message.tool({ id: "call_2", name: "lookup", result: "Rainy" }),
],
}),
)
const assistant = prepared.body.messages[1]
expect(assistant?.role === "assistant" ? assistant.tool_calls : undefined).toEqual([
{
id: "call_1",
type: "function",
function: { name: "lookup", arguments: encodeJson({ city: "Paris" }) },
extra_content: { google: { thought_signature: "sig_1" } },
},
{
id: "call_2",
type: "function",
function: { name: "lookup", arguments: encodeJson({ city: "Tokyo" }) },
},
])
}),
)
it.effect("limits OpenAI and Azure Chat tool call IDs to 40 characters", () =>
Effect.gen(function* () {
const id = `call_${"a".repeat(48)}`
@@ -1805,6 +1844,78 @@ describe("OpenAI Chat route", () => {
}),
)
it.effect("preserves Gemini thought signatures on streamed parallel tool calls", () =>
Effect.gen(function* () {
// Gemini's OpenAI-compatible endpoint omits `index`, streams each call whole,
// and signs only the first call of a parallel batch.
const body = sseEvents(
deltaChunk({
role: "assistant",
tool_calls: [
{
extra_content: { google: { thought_signature: "sig_1" } },
id: "call_1",
type: "function",
function: { name: "lookup", arguments: '{"city":"Paris"}' },
},
],
}),
deltaChunk({
role: "assistant",
tool_calls: [{ id: "call_2", type: "function", function: { name: "lookup", arguments: '{"city":"Tokyo"}' } }],
}),
deltaChunk({}, "stop"),
)
const response = yield* LLMClient.generate(
LLMRequest.update(request, {
tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}),
).pipe(Effect.provide(fixedResponse(body)))
expect(response.events.filter(LLMEvent.is.toolCall)).toEqual([
{
type: "tool-call",
id: "call_1",
name: "lookup",
input: { city: "Paris" },
providerExecuted: undefined,
providerMetadata: { openai: { extraContent: { google: { thought_signature: "sig_1" } } } },
},
{
type: "tool-call",
id: "call_2",
name: "lookup",
input: { city: "Tokyo" },
providerExecuted: undefined,
providerMetadata: undefined,
},
])
}),
)
it.effect("keeps extra content that arrives before the tool identity", () =>
Effect.gen(function* () {
const body = sseEvents(
deltaChunk({
tool_calls: [
{ index: 0, extra_content: { google: { thought_signature: "sig_1" } }, function: { arguments: "{" } },
],
}),
deltaChunk({ tool_calls: [{ index: 0, id: "call_1", function: { name: "lookup", arguments: "}" } }] }),
deltaChunk({}, "tool_calls"),
)
const response = yield* LLMClient.generate(
LLMRequest.update(request, {
tools: [ToolDefinition.make({ name: "lookup", description: "Lookup data", inputSchema: { type: "object" } })],
}),
).pipe(Effect.provide(fixedResponse(body)))
expect(response.events.filter(LLMEvent.is.toolCall).map((event) => event.providerMetadata)).toEqual([
{ openai: { extraContent: { google: { thought_signature: "sig_1" } } } },
])
}),
)
it.effect("does not finalize streamed tool calls when content is filtered", () =>
Effect.gen(function* () {
const body = sseEvents(
+113 -1
View File
@@ -3,7 +3,7 @@ import { Effect, Fiber, Layer, Stream } from "effect"
import * as TestClock from "effect/testing/TestClock"
import { HttpClientRequest } from "effect/unstable/http"
import { Media, Transcription, TranscriptionClient, type TranscriptionEvent } from "../src/index.js"
import { AssemblyAI, Deepgram, Google, OpenAI } from "../src/providers.js"
import { AssemblyAI, Deepgram, ElevenLabs, Google, OpenAI } from "../src/providers.js"
import { it } from "./lib/effect.js"
import { dynamicResponse, json, observe, type Call } from "./lib/http.js"
import { sseEvents } from "./lib/sse.js"
@@ -35,6 +35,9 @@ const formFields = (call: Call) =>
const assemblyai = AssemblyAI.configure({ apiKey: "aai-key", baseURL: "https://assemblyai.test" }).transcription(
"universal-3-5-pro",
)
const elevenlabs = ElevenLabs.configure({ apiKey: "test", baseURL: "https://elevenlabs.test" }).transcription(
"scribe_v2",
)
describe("Transcription", () => {
it.effect("rejects what a route cannot honor before sending anything", () =>
@@ -459,6 +462,115 @@ describe("Transcription", () => {
}),
)
it.effect("rejects ElevenLabs prompts, webhooks, per-channel transcripts, and untimed diarization", () =>
Effect.gen(function* () {
const errors = yield* Effect.all(
[
Transcription.generate({ model: elevenlabs, audio, prompt: "OpenCode" }),
Transcription.generate({ model: elevenlabs, audio, providerOptions: { webhook: true } }),
Transcription.generate({ model: elevenlabs, audio, http: { body: { use_multi_channel: true } } }),
Transcription.generate({
model: elevenlabs,
audio,
diarize: true,
providerOptions: { timestamps_granularity: "none" },
}),
Transcription.generate({
model: elevenlabs,
audio: Media.ref("file_1", { provider: "elevenlabs", mediaType: "audio/mpeg" }),
}),
].map((effect) => Effect.flip(effect)),
)
expect(errors.map((error) => [error.reason._tag, "operation" in error.reason && error.reason.operation])).toEqual(
[
["UnsupportedOperation", "media.prompt"],
["UnsupportedOperation", "transcription.webhook"],
["UnsupportedOperation", "transcription.multichannel"],
["UnsupportedOperation", "media.timestamps"],
["InvalidRequest", false],
],
)
}).pipe(Effect.provide(layer(() => Effect.die("an unsupported request reached the network")))),
)
it.effect("sends ElevenLabs URL audio as source_url and groups diarized words into speaker turns", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const token = (text: string, type: string, start: number, end: number, speaker_id?: string) => ({
text,
type,
start,
end,
speaker_id,
logprob: 0,
})
const response = yield* Transcription.generate({
model: elevenlabs,
audio: Media.url("https://a.test/call.mp3"),
language: "en",
speakers: 2,
providerOptions: { keyterms: ["OpenCode", "Scribe"], tag_audio_events: true, diarize: false },
}).pipe(
Effect.provide(
layer((input) =>
observe(calls, input).pipe(
Effect.as(
json(input, {
language_code: "ENG",
text: "Ready? (laughs) Yes. Go",
words: [
token("Ready?", "word", 0, 0.5, "speaker_0"),
token(" ", "spacing", 0.5, 0.6, "speaker_0"),
token("(laughs)", "audio_event", 0.6, 1, "speaker_0"),
token(" ", "spacing", 1, 1.1, "speaker_0"),
token("Yes.", "word", 1.2, 1.5, "speaker_1"),
token(" ", "spacing", 1.5, 1.6, "speaker_1"),
token("Go", "word", 1.6, 1.9, "speaker_0"),
],
transcription_id: "tr_1",
audio_duration_secs: 2,
}),
),
),
),
),
)
// `observe` re-encodes the FormData with a new boundary, so read the boundary from the sent body.
const boundary = /^--(\S+)/.exec(calls[0].body)?.[1]
const form = yield* Effect.promise(() =>
new Response(calls[0].body, {
headers: { "content-type": `multipart/form-data; boundary=${boundary}` },
}).formData(),
)
expect(calls[0].url).toBe("https://elevenlabs.test/v1/speech-to-text")
expect(calls[0].headers.get("xi-api-key")).toBe("test")
expect(Array.from(form.entries())).toEqual([
["model_id", "scribe_v2"],
["source_url", "https://a.test/call.mp3"],
["language_code", "en"],
["diarize", "true"],
["num_speakers", "2"],
["keyterms", "OpenCode"],
["keyterms", "Scribe"],
["tag_audio_events", "true"],
])
expect(response.segments).toEqual([
{ text: "Ready?", startSeconds: 0, endSeconds: 0.5, speaker: "speaker_0" },
{ text: "Yes.", startSeconds: 1.2, endSeconds: 1.5, speaker: "speaker_1" },
{ text: "Go", startSeconds: 1.6, endSeconds: 1.9, speaker: "speaker_0" },
])
expect(response.words?.map((word) => [word.text, word.speaker, word.confidence])).toEqual([
["Ready?", "speaker_0", 1],
["Yes.", "speaker_1", 1],
["Go", "speaker_0", 1],
])
expect(response.language).toBe("eng")
expect(response.usage).toEqual({ type: "seconds", seconds: 2 })
expect(response.providerMetadata).toEqual({ elevenlabs: { transcriptionId: "tr_1" } })
}),
)
it.effect("rejects reading an AssemblyAI result before the transcript finishes", () =>
Effect.gen(function* () {
const generation = yield* Transcription.resume(assemblyai, { transcriptID: "tr_1" })
+235 -26
View File
@@ -1,9 +1,10 @@
import { describe, expect } from "bun:test"
import { Effect, Layer, Stream } from "effect"
import { Effect, Fiber, Layer, Stream } from "effect"
import * as TestClock from "effect/testing/TestClock"
import { Media, Video, VideoClient, type GenerationEvent, type VideoEvent } from "../src/index.js"
import { Fal, Google, Runway, XAI } from "../src/providers.js"
import { it } from "./lib/effect.js"
import { dynamicResponse, json, observe, settle, type Call } from "./lib/http.js"
import { dynamicResponse, json, observe, settle, type Call, type HandlerInput } from "./lib/http.js"
const layer = (handler: Parameters<typeof dynamicResponse>[0]) =>
VideoClient.layer.pipe(Layer.provideMerge(dynamicResponse(handler)))
@@ -162,27 +163,39 @@ describe("Video / Google Veo", () => {
),
)
it.effect("surfaces an operation error as a failed generation with the provider body", () =>
Effect.gen(function* () {
const failure = {
name: operation,
done: true,
error: { code: 3, message: "Prompt violates policy", status: "INVALID_ARGUMENT" },
}
const error = yield* Video.generate({ model, prompt: "nope" }).pipe(
Effect.flip,
Effect.provide(
layer((input) =>
Effect.succeed(input.request.method === "POST" ? json(input, { name: operation }) : json(input, failure)),
),
),
)
expect(error.reason._tag).toBe("ProviderInternal")
expect(error.message).toBe("Google Veo operation failed: Prompt violates policy")
expect(error.reason.body).toBe(JSON.stringify(failure))
expect(error.reason.http?.status).toBe(200)
}),
)
for (const terminal of [
{ error: { code: 3, message: "Prompt violates policy", status: "INVALID_ARGUMENT" }, tag: "InvalidRequest" },
{ error: { code: 9, message: "Unsupported resolution", status: "FAILED_PRECONDITION" }, tag: "InvalidRequest" },
{ error: { code: 11, message: "Duration out of range", status: "OUT_OF_RANGE" }, tag: "InvalidRequest" },
{ error: { code: 7, message: "Permission denied", status: "PERMISSION_DENIED" }, tag: "Authentication" },
{ error: { code: 16, message: "Invalid credentials", status: "UNAUTHENTICATED" }, tag: "Authentication" },
{ error: { code: 8, message: "Quota exceeded", status: "RESOURCE_EXHAUSTED" }, tag: "RateLimit" },
{ error: { code: 13, message: "Internal error", status: "INTERNAL" }, tag: "ProviderInternal" },
{ error: { code: 14, message: "Service unavailable", status: "UNAVAILABLE" }, tag: "ProviderInternal" },
{ error: { message: "Something broke" }, tag: "ProviderInternal" },
]) {
it.effect(
`surfaces ${terminal.error.status ?? "an uncoded"} operation error as ${terminal.tag} with the provider body`,
() =>
Effect.gen(function* () {
const failure = { name: operation, done: true, error: terminal.error }
const error = yield* Video.generate({ model, prompt: "nope" }).pipe(
Effect.flip,
Effect.provide(
layer((input) =>
Effect.succeed(
input.request.method === "POST" ? json(input, { name: operation }) : json(input, failure),
),
),
),
)
expect(error.reason._tag).toBe(terminal.tag)
expect(error.message).toBe(`Google Veo operation failed: ${terminal.error.message}`)
expect(error.reason.body).toBe(JSON.stringify(failure))
expect(error.reason.http?.status).toBe(200)
}),
)
}
it.effect("reports fully filtered output as a content policy failure", () =>
Effect.gen(function* () {
@@ -332,12 +345,37 @@ describe("Video / xAI", () => {
for (const terminal of [
{
body: { status: "failed", error: { code: "invalid_argument", message: "Prompt cannot be empty." } },
tag: "ProviderInternal",
tag: "InvalidRequest",
message: "xAI Video generation failed (invalid_argument): Prompt cannot be empty.",
},
{
body: { status: "failed", error: { code: "failed_precondition", message: "Extension is not supported." } },
tag: "InvalidRequest",
message: "xAI Video generation failed (failed_precondition): Extension is not supported.",
},
{
body: { status: "failed", error: { code: "permission_denied", message: "Team lacks access." } },
tag: "Authentication",
message: "xAI Video generation failed (permission_denied): Team lacks access.",
},
{
body: { status: "failed", error: { code: "service_unavailable", message: "Overloaded." } },
tag: "ProviderInternal",
message: "xAI Video generation failed (service_unavailable): Overloaded.",
},
{
body: { status: "failed", error: { code: "internal_error", message: "Generation failed." } },
tag: "ProviderInternal",
message: "xAI Video generation failed (internal_error): Generation failed.",
},
{
body: { status: "failed", error: { code: "constructor", message: "Future code." } },
tag: "ProviderInternal",
message: "xAI Video generation failed (constructor): Future code.",
},
{ body: { status: "expired" }, tag: "InvalidRequest", message: "xAI Video request req_1 expired" },
]) {
it.effect(`surfaces ${terminal.body.status} generations with the provider body`, () =>
it.effect(`surfaces ${terminal.body.error?.code ?? terminal.body.status} generations with the provider body`, () =>
Effect.gen(function* () {
const error = yield* Video.generate({ model, prompt: "x" }).pipe(Effect.flip)
expect(error.reason._tag).toBe(terminal.tag)
@@ -558,7 +596,10 @@ describe("Video / fal", () => {
]) {
it.effect(`fails await for ${failure.name} with the response_url body and HTTP context`, () =>
Effect.gen(function* () {
const error = yield* Video.generate({ model, prompt: "x" }).pipe(Effect.flip)
// A transient 500 on the result fetch is retried first; the body and HTTP context survive the final failure.
const fiber = yield* Effect.forkChild(Video.generate({ model, prompt: "x" }).pipe(Effect.flip))
yield* TestClock.adjust("5 minutes")
const error = yield* Fiber.join(fiber)
expect(error.reason._tag).toBe(failure.tag)
expect(error.reason.body).toBe(JSON.stringify(failure.result.body))
expect(error.reason.http).toMatchObject({ url: urls.response, status: failure.result.status })
@@ -765,6 +806,11 @@ describe("Video / Runway", () => {
tag: "ProviderInternal",
message: "Runway task failed (INTERNAL.BAD_OUTPUT.CODE01): Something broke",
},
{
body: { status: "FAILED", failure: "Unsupported dimensions", failureCode: "ASSET.INVALID" },
tag: "InvalidRequest",
message: "Runway task failed (ASSET.INVALID): Unsupported dimensions",
},
{ body: { status: "CANCELLED" }, tag: "InvalidRequest", message: "Runway task task_1 was cancelled" },
]) {
it.effect(`surfaces ${terminal.body.failureCode ?? terminal.body.status} with the task body`, () =>
@@ -908,6 +954,169 @@ describe("Video / Runway", () => {
)
})
// ---------------------------------------------------------------------------
// Transient read failures
// ---------------------------------------------------------------------------
describe("Video / transient read failures", () => {
const model = Runway.configure({ apiKey: "test", baseURL: "https://runway.test/v1" }).video("gen4.5")
const succeeded = { id: "task_1", status: "SUCCEEDED", output: ["https://runway.test/out.mp4"] }
const failure = (input: HandlerInput, status: number, headers?: Record<string, string>) =>
json(input, { error: `HTTP ${status}` }, { status, headers })
const methods = (calls: ReadonlyArray<Call>) => calls.map((call) => call.method)
it.effect("retries a 503 status poll and a 503 result read, then returns the result", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const response = yield* settle(
Video.generate({ model, prompt: "x" }, { poll: { interval: "1 second" } }),
5,
).pipe(
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call, nth } = yield* observe(calls, input)
if (call.method === "POST") return json(input, { id: "task_1" })
// 1: status fails, 2: status succeeds, 3: result fails, 4: result succeeds.
if (nth === 1 || nth === 3) return failure(input, 503)
return json(input, succeeded)
}),
),
),
)
expect(response.video.source).toMatchObject({ type: "url", url: "https://runway.test/out.mp4" })
expect(methods(calls)).toEqual(["POST", "GET", "GET", "GET", "GET"])
}),
)
it.effect("waits for a 429 retry-after before polling again", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const fiber = yield* Effect.forkChild(
Video.generate({ model, prompt: "x" }, { poll: { interval: "1 second" } }).pipe(
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call, nth } = yield* observe(calls, input)
if (call.method === "POST") return json(input, { id: "task_1" })
if (nth === 1) return failure(input, 429, { "retry-after": "10" })
return json(input, succeeded)
}),
),
),
),
)
yield* TestClock.adjust("9 seconds")
expect(methods(calls)).toEqual(["POST", "GET"])
yield* TestClock.adjust("1 second")
yield* Fiber.join(fiber)
expect(methods(calls)).toEqual(["POST", "GET", "GET", "GET"])
}),
)
it.effect("fails a 400 status poll without retrying", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const error = yield* Video.generate({ model, prompt: "x" }).pipe(
Effect.flip,
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call } = yield* observe(calls, input)
return call.method === "POST" ? json(input, { id: "task_1" }) : failure(input, 400)
}),
),
),
)
expect(error.reason._tag).toBe("InvalidRequest")
expect(methods(calls)).toEqual(["POST", "GET"])
}),
)
it.effect("stops retrying at poll.timeout with a Timeout reason", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const error = yield* settle(
Video.generate({ model, prompt: "x" }, { poll: { interval: "1 second", timeout: "5 seconds" } }).pipe(
Effect.flip,
),
6,
).pipe(
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call } = yield* observe(calls, input)
return call.method === "POST" ? json(input, { id: "task_1" }) : failure(input, 503)
}),
),
),
)
expect(error.reason._tag).toBe("Timeout")
expect(calls.filter((call) => call.method === "GET").length).toBeGreaterThan(1)
}),
)
it.effect("bounds a streamed result read's retries by poll.timeout", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const error = yield* settle(
Video.stream({ model, prompt: "x" }, { poll: { interval: "1 second", timeout: "5 seconds" } }).pipe(
Stream.runCollect,
Effect.flip,
),
6,
).pipe(
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call, nth } = yield* observe(calls, input)
if (call.method === "POST") return json(input, { id: "task_1" })
return nth === 1 ? json(input, succeeded) : failure(input, 503)
}),
),
),
)
expect(error.reason._tag).toBe("Timeout")
expect(calls.filter((call) => call.method === "GET").length).toBeGreaterThan(2)
}),
)
it.effect("never retries a failed submit", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const error = yield* Video.generate({ model, prompt: "x" }).pipe(
Effect.flip,
Effect.provide(layer((input) => observe(calls, input).pipe(Effect.map(() => failure(input, 503))))),
)
expect(error.reason._tag).toBe("ProviderInternal")
expect(methods(calls)).toEqual(["POST"])
}),
)
it.effect("never retries a failed cancel", () =>
Effect.gen(function* () {
const calls: Array<Call> = []
const error = yield* Effect.gen(function* () {
const generation = yield* Video.start({ model, prompt: "x" })
return yield* generation.cancel().pipe(Effect.flip)
}).pipe(
Effect.provide(
layer((input) =>
Effect.gen(function* () {
const { call } = yield* observe(calls, input)
if (call.method === "POST") return json(input, { id: "task_1" })
if (call.method === "DELETE") return failure(input, 503)
return json(input, { id: "task_1", status: "RUNNING" })
}),
),
),
)
expect(error.reason._tag).toBe("ProviderInternal")
expect(methods(calls)).toEqual(["POST", "GET", "DELETE"])
}),
)
})
// ---------------------------------------------------------------------------
// Shared queued behavior
// ---------------------------------------------------------------------------
-6
View File
@@ -66,12 +66,6 @@
"node": "./src/shell/parser-wasm.node.ts",
"default": "./src/shell/parser-wasm.bun.ts"
},
"#process-lock-ffi": {
"workerd": "./src/util/process-lock-ffi.workerd.ts",
"bun": "./src/util/process-lock-ffi.bun.ts",
"node": "./src/util/process-lock-ffi.node.ts",
"default": "./src/util/process-lock-ffi.bun.ts"
},
"#v1-migration": {
"types": "./src/database/v1-migration.bun.ts",
"bun": "./src/database/v1-migration.bun.ts",
-1
View File
@@ -26,7 +26,6 @@ const result = await Bun.build({
"#fff",
"#photon-wasm",
"#shell-parser-wasm",
"#process-lock-ffi",
"#v1-migration",
],
splitting: true,
-6
View File
@@ -30,12 +30,6 @@ export interface ExternalDirectoryAuthorization {
readonly save: string
}
export const externalDirectoryPermission = (input: ExternalDirectoryAuthorization) => ({
action: input.action,
resources: [input.resource],
save: [input.save],
})
export interface Target {
readonly absolute: AbsolutePath
/** Location-relative for internal paths, absolute for external paths. */
-6
View File
@@ -1,6 +0,0 @@
export * as File from "./file.js"
import { FileDiff } from "@opencode/schema/file-diff"
export const Diff = FileDiff.Info
export type Diff = typeof Diff.Type
+19 -9
View File
@@ -1,6 +1,7 @@
export * as Generate from "./generate.js"
import { LLM, LLMClient, AIError } from "@opencode/ai"
import { SessionID } from "@opencode/schema/session-id"
import { Context, Effect, Layer, Schema } from "effect"
import { makeLocationNode } from "@opencode/util/effect/app-node"
import { llmClient } from "./effect/app-node-platform.js"
@@ -60,15 +61,24 @@ export const layer = Layer.effect(
? `Model unavailable: ${input.model.providerID}/${input.model.id}`
: "No model specified and no supported model is available",
})
const response = yield* llm.generate(LLM.request({ model: resolved.model, prompt: input.prompt })).pipe(
Effect.mapError(
(error: AIError) =>
new UnavailableError({
message: error.message,
service: resolved.ref.providerID,
}),
),
)
const response = yield* llm
.generate(
LLM.request({
model: resolved.model,
prompt: input.prompt,
// Gateways require session attribution even for a stateless call; no Session is stored.
http: { headers: { "x-opencode-session": SessionID.create() } },
}),
)
.pipe(
Effect.mapError(
(error: AIError) =>
new UnavailableError({
message: error.message,
service: resolved.ref.providerID,
}),
),
)
return response.text
})
+3 -3
View File
@@ -7,7 +7,7 @@ import { AbsolutePath, RelativePath } from "./schema.js"
import { FSUtil } from "@opencode/util/fs-util"
import { AppProcess } from "@opencode/util/process"
import { makeGlobalNode } from "@opencode/util/effect/app-node"
import { File } from "./file.js"
import { FileDiff } from "@opencode/schema/file-diff"
import { KeyedMutex } from "./effect/keyed-mutex.js"
import { VcsPatch } from "./vcs/patch.js"
import { gitExecutable } from "./util/git-executable.js"
@@ -152,7 +152,7 @@ export interface Interface {
to: TreeID
context?: number
paths?: readonly RelativePath[]
}) => Effect.Effect<readonly File.Diff[], OperationError>
}) => Effect.Effect<readonly FileDiff.Info[], OperationError>
readonly restore: (input: {
repository: Repository
files: ReadonlyMap<RelativePath, TreeID>
@@ -571,7 +571,7 @@ const layer = Layer.effect(
additions: stat?.additions ?? 0,
deletions: stat?.deletions ?? 0,
patch: stat?.binary ? "" : (patches.get(entry.file) ?? VcsPatch.emptyPatch(entry.file)),
} satisfies File.Diff
} satisfies FileDiff.Info
})
})
+1 -1
View File
@@ -152,7 +152,7 @@ export function layer(ref: Location.Ref, options: Options = {}): Layer.Layer<Ser
const replacements: LayerNode.Replacements = [
...(options.discovery === false ? vanillaReplacements : []),
...(options.replacements ?? []),
Location.node.replace(Location.boundNode(ref, { discovery: options.discovery })),
Location.node.replace(Location.boundNode(ref)),
InstancePlugins.node.replace(InstancePlugins.bound(options.plugins ?? [])),
]
-3
View File
@@ -1,3 +0,0 @@
/** @deprecated Use FileAccess for path resolution and authorization. */
export { FileAccess as LocationMutation } from "./file-access.js"
export * from "./file-access.js"
-1
View File
@@ -8,7 +8,6 @@ import { LocationServiceMap } from "./location-service-map.js"
export { LocationServiceMap } from "./location-service-map.js"
export type LocationServices = Instance.Services
export type LocationError = Instance.Error
export function buildLocationServiceMap(
replacements: LayerNode.Replacements = [],
+4 -4
View File
@@ -16,12 +16,12 @@ export class Service extends Context.Service<Service, Interface>()("@opencode/Lo
export const node = LayerNode.unbound(Service, tags.values.location)
const layer = (ref: Ref, options?: { readonly discovery?: boolean }) =>
const layer = (ref: Ref) =>
Layer.effect(
Service,
Effect.gen(function* () {
const project = yield* Project.Service
const resolved = yield* project.resolve(ref.directory, options)
const resolved = yield* project.resolve(ref.directory)
return Service.of({
directory: ref.directory,
workspaceID: ref.workspaceID,
@@ -31,9 +31,9 @@ const layer = (ref: Ref, options?: { readonly discovery?: boolean }) =>
}),
)
export const boundNode = (ref: Ref, options?: { readonly discovery?: boolean }) =>
export const boundNode = (ref: Ref) =>
makeLocationNode({
service: Service,
layer: layer(ref, options),
layer: layer(ref),
deps: [Project.node],
})
-2
View File
@@ -45,8 +45,6 @@ export const ResourceTemplate = Mcp.ResourceTemplate
export type ResourceTemplate = Mcp.ResourceTemplate
export const ResourceCatalog = Mcp.ResourceCatalog
export type ResourceCatalog = Mcp.ResourceCatalog
export const ResourceContentPart = Mcp.ResourceContentPart
export type ResourceContentPart = Mcp.ResourceContentPart
export const ResourceContent = Mcp.ResourceContent
export type ResourceContent = Mcp.ResourceContent
File diff suppressed because one or more lines are too long
@@ -11,6 +11,7 @@ import { Model } from "../../model.js"
import { Agent } from "../../agent.js"
import { define } from "@opencode/plugin/effect/plugin"
import { Provider } from "../../provider.js"
import { SessionAffinity } from "../../session/affinity.js"
import type { PluginInternal } from "../internal.js"
const clientID = "Ov23li8tweQw6odWQebz"
@@ -272,7 +273,7 @@ export const GithubCopilotPlugin = define({
.pipe(Effect.orElseSucceed(() => undefined))
const interaction = interactionType(evt.kind, session?.parentID !== undefined)
evt.headers["X-Interaction-Type"] = interaction
evt.headers["X-Interaction-Id"] = evt.sessionID
evt.headers["X-Interaction-Id"] = session ? SessionAffinity.of(session) : evt.sessionID
if (interaction !== "conversation-agent") evt.headers["x-initiator"] = "agent"
}),
{ providerID: Provider.ID.githubCopilot },
+7 -2
View File
@@ -9,6 +9,7 @@ import { Bus } from "../../bus.js"
import { Integration } from "../../integration.js"
import { OauthCallbackPage } from "../../oauth/page.js"
import { Provider } from "../../provider.js"
import { SessionAffinity } from "../../session/affinity.js"
import type { PluginInternal } from "../internal.js"
const clientID = "app_EMoamEEZ73f0CkXaXp7hrann"
@@ -299,12 +300,16 @@ export const OpenAIPlugin = define({
yield* ctx.session.hook(
"model.request",
(evt) =>
Effect.sync(() => {
Effect.gen(function* () {
if (!chatgpt) return
if (evt.baseURL && URL.canParse(evt.baseURL) && new URL(evt.baseURL).origin === "https://api.openai.com")
evt.baseURL = codexBaseURL
const session = yield* ctx.session
.get({ sessionID: evt.sessionID })
.pipe(Effect.orElseSucceed(() => undefined))
evt.headers.originator = "opencode"
evt.headers["session-id"] = evt.sessionID
// ChatGPT routes its prompt cache on this header, so children share the parent's.
evt.headers["session-id"] = session ? SessionAffinity.of(session) : evt.sessionID
}),
{ providerID: Provider.ID.openai },
)
+2 -5
View File
@@ -65,7 +65,7 @@ export interface Interface {
/** Records Project activity for recency ordering, at most once per minute per Project. */
readonly activate: (projectID: ID) => Effect.Effect<void>
/** Resolves and persists the owning Project. */
readonly resolve: (input: AbsolutePath, options?: { readonly discovery?: boolean }) => Effect.Effect<Resolved>
readonly resolve: (input: AbsolutePath) => Effect.Effect<Resolved>
}
export class Service extends Context.Service<Service, Interface>()("@opencode/Project") {}
@@ -334,10 +334,7 @@ const layer = Layer.effect(
}
})
const resolve = Effect.fn("Project.resolve")(function* (
input: AbsolutePath,
_options?: { readonly discovery?: boolean },
) {
const resolve = Effect.fn("Project.resolve")(function* (input: AbsolutePath) {
const directory = AbsolutePath.make(yield* fs.resolve(input))
const native = yield* fs.up({ targets: [".git", ".hg"], start: directory, mode: "first" }).pipe(
Effect.map((matches) => matches[0]),
+20 -18
View File
@@ -6,7 +6,6 @@ import { Context, Effect, Layer, Schema, Types } from "effect"
import { Pty } from "@opencode/schema/pty"
import { Bus } from "./bus.js"
import { Location } from "./location.js"
import { PtyID } from "./pty/schema.js"
import { ShellSelect } from "./shell/select.js"
import { lazy } from "./util/lazy.js"
@@ -35,6 +34,9 @@ type Active = {
listeners: Disp[]
}
export const ID = Pty.ID
export type ID = Pty.ID
export const Info = Pty.Info
export type Info = Types.DeepMutable<typeof Info.Type>
@@ -69,21 +71,21 @@ export type Attachment = {
}
export class NotFoundError extends Schema.TaggedError<NotFoundError>()("Pty.NotFoundError", {
ptyID: PtyID,
ptyID: ID,
}) {}
export class ExitedError extends Schema.TaggedError<ExitedError>()("Pty.ExitedError", {
ptyID: PtyID,
ptyID: ID,
}) {}
export interface Interface {
readonly list: () => Effect.Effect<Info[]>
readonly get: (id: PtyID) => Effect.Effect<Info, NotFoundError>
readonly get: (id: ID) => Effect.Effect<Info, NotFoundError>
readonly create: (input: CreateInput) => Effect.Effect<Info>
readonly update: (id: PtyID, input: UpdateInput) => Effect.Effect<Info, NotFoundError>
readonly remove: (id: PtyID) => Effect.Effect<void, NotFoundError>
readonly write: (id: PtyID, data: string) => Effect.Effect<void, NotFoundError>
readonly attach: (id: PtyID, input: AttachInput) => Effect.Effect<Attachment, NotFoundError | ExitedError>
readonly update: (id: ID, input: UpdateInput) => Effect.Effect<Info, NotFoundError>
readonly remove: (id: ID) => Effect.Effect<void, NotFoundError>
readonly write: (id: ID, data: string) => Effect.Effect<void, NotFoundError>
readonly attach: (id: ID, input: AttachInput) => Effect.Effect<Attachment, NotFoundError | ExitedError>
}
export class Service extends Context.Service<Service, Interface>()("@opencode/Pty") {}
@@ -96,8 +98,8 @@ const layer = Layer.effect(
const shell = yield* ShellSelect.Service
const context = yield* Effect.context()
const runFork = Effect.runForkWith(context)
const sessions = new Map<PtyID, Active>()
const exitOrder: PtyID[] = []
const sessions = new Map<ID, Active>()
const exitOrder: ID[] = []
function notifyEnd(session: Active, event: { exitCode?: number }) {
for (const subscriber of session.subscribers.values()) {
@@ -131,13 +133,13 @@ const layer = Layer.effect(
}),
)
const requireSession = Effect.fn("Pty.requireSession")(function* (id: PtyID) {
const requireSession = Effect.fn("Pty.requireSession")(function* (id: ID) {
const session = sessions.get(id)
if (!session) return yield* new NotFoundError({ ptyID: id })
return session
})
const removeSession = Effect.fnUntraced(function* (id: PtyID) {
const removeSession = Effect.fnUntraced(function* (id: ID) {
const session = sessions.get(id)
if (!session) return
sessions.delete(id)
@@ -148,7 +150,7 @@ const layer = Layer.effect(
yield* bus.publish(Pty.Event.Deleted, { id: session.info.id })
})
const remove = Effect.fn("Pty.remove")(function* (id: PtyID) {
const remove = Effect.fn("Pty.remove")(function* (id: ID) {
yield* requireSession(id)
yield* removeSession(id)
})
@@ -157,12 +159,12 @@ const layer = Layer.effect(
return Array.from(sessions.values()).map((session) => session.info)
})
const get = Effect.fn("Pty.get")(function* (id: PtyID) {
const get = Effect.fn("Pty.get")(function* (id: ID) {
return (yield* requireSession(id)).info
})
const create = Effect.fn("Pty.create")(function* (input: CreateInput) {
const id = PtyID.ascending()
const id = ID.ascending()
const command = input.command || (yield* shell.resolve({ priority: "config" }))
const args = ShellSelect.login(command) ? [...(input.args ?? []), "-l"] : [...(input.args ?? [])]
const cwd = input.cwd || location.directory
@@ -242,7 +244,7 @@ const layer = Layer.effect(
return info
})
const update = Effect.fn("Pty.update")(function* (id: PtyID, input: UpdateInput) {
const update = Effect.fn("Pty.update")(function* (id: ID, input: UpdateInput) {
const session = yield* requireSession(id)
if (input.title) session.info.title = input.title
if (input.size && session.info.status === "running") session.process.resize(input.size.cols, input.size.rows)
@@ -250,12 +252,12 @@ const layer = Layer.effect(
return session.info
})
const write = Effect.fn("Pty.write")(function* (id: PtyID, data: string) {
const write = Effect.fn("Pty.write")(function* (id: ID, data: string) {
const session = yield* requireSession(id)
if (session.info.status === "running") session.process.write(data)
})
const attach = Effect.fn("Pty.attach")(function* (id: PtyID, input: AttachInput) {
const attach = Effect.fn("Pty.attach")(function* (id: ID, input: AttachInput) {
const session = yield* requireSession(id)
if (session.info.status !== "running") return yield* new ExitedError({ ptyID: id })
yield* Effect.logInfo("client attached to session", { id, directory: location.directory })
-1
View File
@@ -1 +0,0 @@
export { ID as PtyID } from "@opencode/schema/pty"
+2 -2
View File
@@ -2,7 +2,7 @@ export * as PtyTicket from "./ticket.js"
import type { Workspace } from "@opencode/schema/workspace"
import { PtyTicket } from "@opencode/schema/pty-ticket"
import { PtyID } from "./schema.js"
import type { Pty } from "@opencode/schema/pty"
import { Cache, Context, Duration, Effect, Layer } from "effect"
import { makeGlobalNode } from "@opencode/util/effect/app-node"
@@ -12,7 +12,7 @@ const CAPACITY = 10_000
export const ConnectToken = PtyTicket.ConnectToken
export type Scope = {
readonly ptyID: PtyID
readonly ptyID: Pty.ID
readonly directory?: string
readonly workspaceID?: Workspace.ID
}
+2 -1
View File
@@ -145,8 +145,9 @@ function parts(input: string) {
.filter(Boolean)
}
// cachePath makes each `:`-separated host part a directory.
function safeHost(input: string) {
return Boolean(input) && !input.startsWith("-") && !/[\s/\\]/.test(input)
return Boolean(input) && !input.startsWith("-") && input.split(":").every(safeSegment)
}
function safeSegment(input: string) {
+8
View File
@@ -0,0 +1,8 @@
export * as SessionAffinity from "./affinity.js"
import type { SessionSchema } from "./schema.js"
// TODO: Should the `model.request` hook expose affinity so plugins stop deriving it from `ctx.session.get`?
/** The Session ID that groups model requests for provider cache routing: children share the parent's, forks the fork source's. */
export const of = (session: Pick<SessionSchema.Info, "id" | "parentID" | "fork">) =>
session.parentID ?? session.fork?.sessionID ?? session.id
-9
View File
@@ -50,15 +50,6 @@ export class StepFailedError extends Schema.TaggedError<StepFailedError>()("Sess
}
}
export class UserInterruptedError extends Schema.TaggedError<UserInterruptedError>()(
"Session.UserInterruptedError",
{},
) {
override get message() {
return "Session interrupted by user"
}
}
export class PromptConflictError extends Schema.TaggedError<PromptConflictError>()("Session.PromptConflictError", {
sessionID: SessionSchema.ID,
messageID: SessionMessage.ID,
+1 -4
View File
@@ -12,7 +12,6 @@ import { SessionRunner } from "./runner/index.js"
import { SessionSchema } from "./schema.js"
import { SessionStore } from "./store.js"
import { toSessionError } from "./to-session-error.js"
import { UserInterruptedError } from "./error.js"
import { SessionInbox } from "./inbox.js"
export interface Interface {
@@ -51,9 +50,7 @@ type InterruptReason = "user" | "shutdown" | "inactivity"
export function terminal(exit: Exit.Exit<void, SessionRunner.RunError>, reason?: InterruptReason) {
if (Exit.isSuccess(exit)) return { type: "succeeded" as const }
if (Cause.hasInterrupts(exit.cause)) return { type: "interrupted" as const, reason: reason ?? "shutdown" }
const failure = Cause.squash(exit.cause)
if (failure instanceof UserInterruptedError) return { type: "interrupted" as const, reason: "user" as const }
return { type: "failed" as const, error: toSessionError(failure) }
return { type: "failed" as const, error: toSessionError(Cause.squash(exit.cause)) }
}
/** Process-local execution: drains run in this process using the selected instance. */
+5 -4
View File
@@ -31,6 +31,7 @@ import { Permission } from "../permission.js"
import { PluginHooks } from "../plugin/hooks.js"
import { QuestionTool } from "../tool/plugin/question.js"
import { Tool } from "../tool.js"
import { SessionAffinity } from "./affinity.js"
import { SessionModelTransport } from "./model-transport.js"
import { SessionProviderContext } from "./provider-context.js"
import { SessionRunnerModel } from "./runner/model.js"
@@ -268,17 +269,17 @@ export const layer = Layer.effect(
const entries = Object.entries(shaped.options)
const generation = Object.fromEntries(entries.filter(([k]) => GENERATION_KEYS.has(k))) as GenerationOptionsFields
const providerOptions = Object.fromEntries(entries.filter(([k]) => !GENERATION_KEYS.has(k)))
const affinity = session.parentID ?? session.fork?.sessionID ?? session.id
const affinity = SessionAffinity.of(session)
const base = LLM.request({
model: model.model,
http: {
headers: {
"x-session-affinity": session.id,
"X-Session-Id": session.id,
"x-session-affinity": affinity,
"X-Session-Id": affinity,
...(session.parentID ? { "x-parent-session-id": session.parentID } : {}),
"User-Agent": App.useragent(app),
"x-opencode-project": session.projectID,
"x-opencode-session": session.id,
"x-opencode-session": affinity,
"x-opencode-client": app.name,
},
},
+1 -2
View File
@@ -4,7 +4,7 @@ import type { AIError } from "@opencode/ai"
import { Context, Data, Effect } from "effect"
import { SessionSchema } from "../schema.js"
import type { Promotable } from "../inbox.js"
import type { AgentNotFoundError, MessageDecodeError, StepFailedError, UserInterruptedError } from "../error.js"
import type { AgentNotFoundError, MessageDecodeError, StepFailedError } from "../error.js"
import { SessionRunnerModel } from "./model.js"
import type { Instructions } from "../../instructions/index.js"
@@ -14,7 +14,6 @@ export type RunError =
| MessageDecodeError
| AgentNotFoundError
| StepFailedError
| UserInterruptedError
| Instructions.InitializationBlocked
export type Continuation = { readonly step: number }
-13
View File
@@ -29,19 +29,6 @@ export class ModelUnavailableError extends Schema.TaggedError<ModelUnavailableEr
return `Model unavailable: ${this.providerID}/${this.modelID}`
}
}
export const VariantUnavailableError = ModelResolver.VariantUnavailableError
export type VariantUnavailableError = ModelResolver.VariantUnavailableError
export const UnsupportedPackageError = ModelResolver.UnsupportedPackageError
export type UnsupportedPackageError = ModelResolver.UnsupportedPackageError
export const ModelConfigurationError = ModelResolver.ModelConfigurationError
export type ModelConfigurationError = ModelResolver.ModelConfigurationError
export const ModelInitializationError = ModelResolver.ModelInitializationError
export type ModelInitializationError = ModelResolver.ModelInitializationError
export const UnresolvedProviderVariablesError = ModelResolver.UnresolvedProviderVariablesError
export type UnresolvedProviderVariablesError = ModelResolver.UnresolvedProviderVariablesError
export const UnsupportedCompactionError = ModelResolver.UnsupportedCompactionError
export type UnsupportedCompactionError = ModelResolver.UnsupportedCompactionError
export type Error = ModelNotSelectedError | ModelUnavailableError | ModelResolver.Error
export type Resolved = ModelResolver.Resolved
+3 -57
View File
@@ -1,6 +1,6 @@
export * as SessionRunnerRetry from "./retry.js"
import { AIError, isContextOverflowFailure } from "@opencode/ai"
import { AIError, isRetryable } from "@opencode/ai"
import { Agent } from "@opencode/schema/agent"
import { Model } from "@opencode/schema/model"
import { SessionError } from "@opencode/schema/session-error"
@@ -10,7 +10,8 @@ import type { PluginHooks } from "../../plugin/hooks.js"
import { SessionEvent } from "../event.js"
import { SessionMessage } from "../message.js"
import { SessionSchema } from "../schema.js"
import { toSessionError } from "../to-session-error.js"
export { isRetryable }
interface Input {
readonly cause: AIError
@@ -27,43 +28,6 @@ export interface Decision {
readonly delay: number
}
export function isRetryable(error: AIError) {
const override = error.reason.http?.headers["x-should-retry"]
if (override === "true") return true
if (override === "false") return false
switch (error.reason._tag) {
case "RateLimit":
case "ProviderInternal":
return true
// A WebSocket acknowledgment marks delivery accepted before model output may exist.
// Read failures can still recover; the Step chooses retry versus continuation from durable output.
case "Transport":
return (
error.reason.delivery !== "rejected" &&
(error.reason.delivery !== "accepted" || error.reason.operation === "read")
)
case "InvalidProviderOutput":
return error.reason.classification === "incomplete-stream"
// Unrecognized failures retry: classification records affirmative
// deterministic evidence, and transient failures are exactly the ones
// that arrive in shapes no classifier anticipates.
case "UnknownProvider":
return true
case "Authentication":
case "QuotaExceeded":
case "ContentPolicy":
case "InvalidRequest":
case "UnsupportedOperation":
case "NoRoute":
case "Timeout":
return false
default: {
const exhaustive: never = error.reason
return exhaustive
}
}
}
/** Bound provider-requested delays so a hostile or buggy retry-after cannot stall a session for hours. */
const RETRY_AFTER_MAX = Duration.toMillis("15 minutes")
@@ -119,24 +83,6 @@ export const policy = (sessionID: SessionSchema.ID) =>
})
})
/**
* Retries one auxiliary request's transient failures under a shared `policy` allowance, letting the
* session retry hook adjust each decision. Context overflow is never transient: callers recover it.
*/
export const transient =
(decide: Effect.Success<ReturnType<typeof policy>>, input: Pick<Input, "agent" | "model" | "hook">) =>
<A, R>(effect: Effect.Effect<A, AIError, R>) =>
Effect.retry(effect, {
while: (cause) =>
Effect.gen(function* () {
if (isContextOverflowFailure(cause)) return false
const decision = yield* decide({ ...input, cause, error: toSessionError(cause), retry: isRetryable(cause) })
if (!decision.retry) return false
yield* Effect.sleep(decision.delay)
return true
}),
})
export const make = (bus: Bus.Interface, sessionID: SessionSchema.ID) =>
Effect.gen(function* () {
const decide = yield* policy(sessionID)
@@ -3,7 +3,8 @@ import { Tool } from "@opencode/schema/tool"
import { SessionError } from "@opencode/schema/session-error"
import { Permission } from "../permission.js"
import { Integration } from "../integration.js"
import { AgentNotFoundError, StepFailedError, UserInterruptedError } from "./error.js"
import { AgentNotFoundError, StepFailedError } from "./error.js"
import { ModelResolver } from "../model-resolver.js"
import { SessionRunnerModel } from "./runner/model.js"
export function toSessionError(cause: unknown): SessionError.Error {
@@ -48,18 +49,17 @@ export function toSessionError(cause: unknown): SessionError.Error {
return unwrapped.message === "" ? { ...unwrapped, type: "tool.execution", message: cause.message } : unwrapped
}
if (cause instanceof StepFailedError) return cause.error
if (cause instanceof SessionRunnerModel.UnsupportedCompactionError)
if (cause instanceof ModelResolver.UnsupportedCompactionError)
return { type: "provider.unsupported-operation", message: cause.message }
if (cause instanceof AgentNotFoundError) return { type: "unknown", message: cause.message }
if (cause instanceof UserInterruptedError) return { type: "aborted", message: cause.message }
if (
cause instanceof SessionRunnerModel.ModelNotSelectedError ||
cause instanceof SessionRunnerModel.ModelUnavailableError ||
cause instanceof SessionRunnerModel.VariantUnavailableError ||
cause instanceof SessionRunnerModel.UnsupportedPackageError ||
cause instanceof SessionRunnerModel.ModelConfigurationError ||
cause instanceof SessionRunnerModel.ModelInitializationError ||
cause instanceof SessionRunnerModel.UnresolvedProviderVariablesError
cause instanceof ModelResolver.VariantUnavailableError ||
cause instanceof ModelResolver.UnsupportedPackageError ||
cause instanceof ModelResolver.ModelConfigurationError ||
cause instanceof ModelResolver.ModelInitializationError ||
cause instanceof ModelResolver.UnresolvedProviderVariablesError
)
return { type: "provider.no-route", message: cause.message }
if (cause instanceof Integration.AuthorizationError) return { type: "provider.auth", message: cause.message }
+2 -2
View File
@@ -3,7 +3,7 @@ export * as Snapshot from "./snapshot.js"
import { makeLocationNode } from "@opencode/util/effect/app-node"
import path from "path"
import { Context, Effect, Fiber, Layer, Schema, Scope } from "effect"
import { File } from "./file.js"
import { FileDiff } from "@opencode/schema/file-diff"
import { FSUtil } from "@opencode/util/fs-util"
import { Git } from "./git.js"
import { Global } from "@opencode/util/global"
@@ -58,7 +58,7 @@ export interface Interface extends State.Transformable<Editor> {
* Generate structured per-file diffs between two captured trees. `context`
* controls unchanged lines around each unified diff hunk.
*/
readonly diff: (input: DiffInput) => Effect.Effect<readonly File.Diff[], Error>
readonly diff: (input: DiffInput) => Effect.Effect<readonly FileDiff.Info[], Error>
/**
* Restore selected project-relative paths from their associated trees. A path
-10
View File
@@ -1,10 +0,0 @@
export function findLast<T>(
items: readonly T[],
predicate: (item: T, index: number, items: readonly T[]) => boolean,
): T | undefined {
for (let i = items.length - 1; i >= 0; i -= 1) {
const item = items[i]
if (predicate(item, i, items)) return item
}
return undefined
}
@@ -1,50 +0,0 @@
import { dlopen, read, type Pointer } from "bun:ffi"
import { existsSync } from "node:fs"
export type LockResult =
| { readonly acquired: true }
| { readonly acquired: false; readonly held: true }
| { readonly acquired: false; readonly held: false; readonly code: number }
const LOCK_EX = 2
const LOCK_NB = 4
const DARWIN_EWOULDBLOCK = 35
const LINUX_EWOULDBLOCK = 11
export function lockDarwin(fd: number): LockResult {
const library = dlopen("/usr/lib/libSystem.B.dylib", {
flock: { args: ["i32", "i32"], returns: "i32" },
__error: { args: [], returns: "ptr" },
})
try {
const result = library.symbols.flock(fd, LOCK_EX | LOCK_NB)
const code = result === 0 ? 0 : errorCode(library.symbols.__error())
if (result === 0) return { acquired: true }
if (code === DARWIN_EWOULDBLOCK) return { acquired: false, held: true }
return { acquired: false, held: false, code }
} finally {
library.close()
}
}
export function lockLinux(fd: number): LockResult {
const musl = `/lib/libc.musl-${process.arch === "arm64" ? "aarch64" : "x86_64"}.so.1`
const library = dlopen(existsSync(musl) ? musl : "libc.so.6", {
flock: { args: ["i32", "i32"], returns: "i32" },
__errno_location: { args: [], returns: "ptr" },
})
try {
const result = library.symbols.flock(fd, LOCK_EX | LOCK_NB)
const code = result === 0 ? 0 : errorCode(library.symbols.__errno_location())
if (result === 0) return { acquired: true }
if (code === LINUX_EWOULDBLOCK) return { acquired: false, held: true }
return { acquired: false, held: false, code }
} finally {
library.close()
}
}
function errorCode(pointer: Pointer | bigint | null) {
if (pointer === null) throw new Error("Failed to read process lock error code")
return read.i32(pointer, 0)
}
@@ -1,43 +0,0 @@
import { dlopen, getInt32 } from "node:ffi"
export type LockResult =
| { readonly acquired: true }
| { readonly acquired: false; readonly held: true }
| { readonly acquired: false; readonly held: false; readonly code: number }
const LOCK_EX = 2
const LOCK_NB = 4
const DARWIN_EWOULDBLOCK = 35
const LINUX_EWOULDBLOCK = 11
export function lockDarwin(fd: number): LockResult {
const library = dlopen("/usr/lib/libSystem.B.dylib", {
flock: { arguments: ["int32", "int32"], return: "int32" },
__error: { arguments: [], return: "pointer" },
})
try {
const result = library.functions.flock(fd, LOCK_EX | LOCK_NB)
const code = result === 0 ? 0 : getInt32(library.functions.__error(), 0)
if (result === 0) return { acquired: true }
if (code === DARWIN_EWOULDBLOCK) return { acquired: false, held: true }
return { acquired: false, held: false, code }
} finally {
library.lib.close()
}
}
export function lockLinux(fd: number): LockResult {
const library = dlopen("libc.so.6", {
flock: { arguments: ["int32", "int32"], return: "int32" },
__errno_location: { arguments: [], return: "pointer" },
})
try {
const result = library.functions.flock(fd, LOCK_EX | LOCK_NB)
const code = result === 0 ? 0 : getInt32(library.functions.__errno_location(), 0)
if (result === 0) return { acquired: true }
if (code === LINUX_EWOULDBLOCK) return { acquired: false, held: true }
return { acquired: false, held: false, code }
} finally {
library.lib.close()
}
}
@@ -1,14 +0,0 @@
export type LockResult =
| { readonly acquired: true }
| { readonly acquired: false; readonly held: true }
| { readonly acquired: false; readonly held: false; readonly code: number }
// workerd has no FFI and no cross-process file locking; a Durable Object is
// already single-threaded per instance, so nothing on this runtime should
// reach these.
const unavailable = (_fd: number): LockResult => {
throw new Error("Process locks are unavailable on the workerd runtime")
}
export const lockDarwin = unavailable
export const lockLinux = unavailable
-131
View File
@@ -1,131 +0,0 @@
import { lockDarwin, lockLinux, type LockResult } from "#process-lock-ffi"
import { closeSync, mkdirSync, openSync } from "node:fs"
import { connect, createServer, type Server, type Socket } from "node:net"
import path from "node:path"
import { Effect, Schema } from "effect"
import { Hash } from "@opencode/util/hash"
export namespace ProcessLock {
export class HeldError extends Schema.TaggedError<HeldError>()("ProcessLockHeldError", {
file: Schema.String,
}) {
override get message() {
return `Process lock is already held: ${this.file}`
}
}
export class SystemError extends Schema.TaggedError<SystemError>()("ProcessLockSystemError", {
file: Schema.String,
operation: Schema.Literals(["open", "acquire"]),
code: Schema.String,
}) {
override get message() {
return `Process lock ${this.operation} failed for ${this.file}: ${this.code}`
}
}
export type LockError = HeldError | SystemError
const acquirePosix = Effect.fnUntraced(function* (file: string) {
const fd = yield* Effect.try({
try: () => {
mkdirSync(path.dirname(file), { recursive: true })
return openSync(file, "a+", 0o600)
},
catch: (cause) =>
new SystemError({
file,
operation: "open",
code: cause instanceof Error ? cause.message : String(cause),
}),
})
const result = yield* Effect.try({
try: () => lock(fd),
catch: (cause) =>
new SystemError({
file,
operation: "acquire",
code: cause instanceof Error ? cause.message : String(cause),
}),
}).pipe(
Effect.tapError(() =>
Effect.sync(() => {
closeSync(fd)
}),
),
)
if (result.acquired) return fd
closeSync(fd)
return yield* result.held
? new HeldError({ file })
: new SystemError({ file, operation: "acquire", code: String(result.code) })
})
export const acquire = Effect.fn("ProcessLock.acquire")(function* (file: string) {
if (process.platform === "win32") {
yield* Effect.acquireRelease(acquireWindows(file), closeWindows)
return
}
yield* Effect.acquireRelease(acquirePosix(file), (fd) =>
Effect.sync(() => {
closeSync(fd)
}),
)
})
}
function lock(fd: number): LockResult {
if (process.platform === "darwin") return lockDarwin(fd)
if (process.platform === "linux") return lockLinux(fd)
throw new Error(`Unsupported process lock platform: ${process.platform}`)
}
function acquireWindows(file: string) {
return Effect.callback<Server, ProcessLock.LockError>((resume) => {
const server = createServer()
let probe: Socket | undefined
const pipe = `\\\\.\\pipe\\opencode-process-lock-${Hash.sha256(path.resolve(file).toLowerCase())}`
const onError = (cause: NodeJS.ErrnoException) => {
server.off("listening", onListening)
probe = connect(pipe)
const onProbeError = () => {
probe?.off("connect", onConnect)
resume(
Effect.fail(
new ProcessLock.SystemError({
file,
operation: "acquire",
code: cause.code ?? cause.message,
}),
),
)
}
const onConnect = () => {
probe?.off("error", onProbeError)
probe?.destroy()
resume(Effect.fail(new ProcessLock.HeldError({ file })))
}
probe.once("connect", onConnect)
probe.once("error", onProbeError)
}
const onListening = () => {
server.off("error", onError)
resume(Effect.succeed(server))
}
server.once("error", onError)
server.once("listening", onListening)
server.on("connection", (socket) => socket.destroy())
server.listen(pipe)
return Effect.sync(() => {
probe?.destroy()
server.close()
})
})
}
function closeWindows(server: Server) {
return Effect.callback<void>((resume) => {
if (!server.listening) return resume(Effect.void)
server.close((error) => resume(error ? Effect.die(error) : Effect.void))
})
}
@@ -1,17 +0,0 @@
import { ProcessLock } from "@opencode/core/util/process-lock"
import { Effect, Schema } from "effect"
import fs from "node:fs/promises"
const input = Schema.decodeUnknownSync(
Schema.fromJsonString(Schema.Struct({ file: Schema.String, ready: Schema.String })),
)(process.argv[2])
await Effect.runPromise(
Effect.scoped(
Effect.gen(function* () {
yield* ProcessLock.acquire(input.file)
yield* Effect.promise(() => fs.writeFile(input.ready, String(process.pid)))
return yield* Effect.never
}),
),
)
+56 -1
View File
@@ -1,5 +1,6 @@
import { expect } from "bun:test"
import { LanguageModel } from "@opencode/ai"
import { LanguageModel, LLMClient } from "@opencode/ai"
import { RequestExecutor } from "@opencode/ai/route"
import { OpenAIChat } from "@opencode/ai/protocols"
import { TestLLM } from "@opencode/ai/testing"
import { AISDK } from "@opencode/core/aisdk"
@@ -10,6 +11,7 @@ import { ID, Info, Model, Ref } from "@opencode/core/model"
import { Provider } from "@opencode/core/provider"
import { Npm } from "@opencode/util/npm"
import { Effect, Layer } from "effect"
import { HttpClient, HttpClientResponse } from "effect/unstable/http"
import { testEffect } from "./lib/effect"
const selected = Info.make({
@@ -89,3 +91,56 @@ resolverIt.effect("resolves dynamic models with their catalog metadata", () =>
})
}),
)
testEffect(Layer.empty).effect("attributes each stateless completion without creating a stored session", () =>
Effect.gen(function* () {
const sessions: string[] = []
const http = Layer.succeed(
HttpClient.HttpClient,
HttpClient.make((request) =>
Effect.sync(() => {
const session = request.headers["x-opencode-session"]
if (!session)
return HttpClientResponse.fromWeb(
request,
Response.json(
{
error: { type: "MissingSessionID", message: "Session ID is required" },
},
{ status: 400 },
),
)
sessions.push(session)
return HttpClientResponse.fromWeb(
request,
new Response(
`data: ${JSON.stringify({
id: "completion",
object: "chat.completion.chunk",
created: 1,
model: "gemini",
choices: [{ index: 0, delta: { content: "OK" }, finish_reason: "stop" }],
})}\n\ndata: [DONE]\n\n`,
{ headers: { "content-type": "text/event-stream" } },
),
)
}),
),
)
const native = LLMClient.layer.pipe(Layer.provide(RequestExecutor.layer.pipe(Layer.provide(http))))
yield* Effect.gen(function* () {
const generate = yield* Generate.Service
for (let index = 0; index < 2; index++) {
expect(
yield* generate.text({
prompt: "Return exactly OK",
model: Ref.make({ providerID: selected.providerID, id: selected.id }),
}),
).toBe("OK")
}
}).pipe(Effect.provide(Generate.layer.pipe(Layer.provide(Layer.merge(resolver, native)))))
expect(sessions).toHaveLength(2)
expect(sessions[0]).toStartWith("ses_")
expect(sessions[1]).not.toBe(sessions[0])
}),
)
@@ -183,10 +183,11 @@ describe("GithubCopilotPlugin", () => {
it.effect("classifies child-session steps as subagent interactions", () =>
Effect.gen(function* () {
yield* addPlugin()
const event = yield* modelRequest((yield* sessions()).child, "primary")
const ids = yield* sessions()
const event = yield* modelRequest(ids.child, "primary")
expect(event.headers).toEqual({
"X-Interaction-Type": "conversation-subagent",
"X-Interaction-Id": event.sessionID,
"X-Interaction-Id": ids.parent,
"x-initiator": "agent",
})
}),
@@ -207,10 +208,11 @@ describe("GithubCopilotPlugin", () => {
it.effect("classifies compaction requests by kind rather than agent", () =>
Effect.gen(function* () {
yield* addPlugin()
const event = yield* modelRequest((yield* sessions()).child, "compaction", "build")
const ids = yield* sessions()
const event = yield* modelRequest(ids.child, "compaction", "build")
expect(event.headers).toEqual({
"X-Interaction-Type": "conversation-compaction",
"X-Interaction-Id": event.sessionID,
"X-Interaction-Id": ids.parent,
"x-initiator": "agent",
})
}),
@@ -1,6 +1,6 @@
import { Money } from "@opencode/schema/money"
import { Agent } from "@opencode/schema/agent"
import { Session } from "@opencode/schema/session"
import { Session } from "@opencode/core/session"
import { OpenAIResponses } from "@opencode/ai/protocols/openai-responses"
import { describe, expect } from "bun:test"
import { ConfigProvider, DateTime, Effect } from "effect"
@@ -41,10 +41,14 @@ function required<T>(value: T | undefined): T {
return value
}
const request = Effect.fn(function* (providerID: Provider.ID, baseURL: string) {
const request = Effect.fn(function* (
providerID: Provider.ID,
baseURL: string,
sessionID = Session.ID.make("ses_test"),
) {
const hooks = yield* PluginHooks.Service
const event = yield* hooks.trigger("session", "model.request", {
sessionID: Session.ID.make("ses_test"),
sessionID,
agent: Agent.ID.make("build"),
model: Model.Ref.make({ providerID, id: Model.ID.make("gpt-5.5") }),
kind: "primary",
@@ -156,6 +160,12 @@ describe("OpenAIPlugin", () => {
expect(custom.headers).not.toHaveProperty("originator")
expect(proxy.baseURL).toBe("https://proxy.example/v1?region=us")
expect(proxy.headers).toMatchObject({ originator: "opencode", "session-id": "ses_test" })
const sessions = yield* Session.Service
const location = yield* Location.Service
const parent = yield* sessions.create({ location: { directory: location.directory } })
const child = yield* sessions.create({ parentID: parent.id })
const childRequest = yield* request(Provider.ID.openai, "https://api.openai.com/v1", child.id)
expect(childRequest.headers).toMatchObject({ "session-id": parent.id })
const eligible = required(yield* models.get(Provider.ID.openai, Model.ID.make("gpt-5.5")))
expect(eligible.package).toBe("@opencode/ai/providers/openai")
expect(eligible.headers).toMatchObject({ originator: "opencode", "chatgpt-account-id": "acct_123" })
+4 -5
View File
@@ -5,13 +5,12 @@ import { LayerNode } from "@opencode/util/effect/layer-node"
import { Bus } from "@opencode/core/bus"
import { Location } from "@opencode/core/location"
import { Pty } from "@opencode/core/pty"
import { PtyID } from "@opencode/core/pty/schema"
import { AbsolutePath } from "@opencode/core/schema"
import { ShellSelect } from "@opencode/core/shell/select"
import { location } from "../fixture/location"
import { testEffect } from "../lib/effect"
type PtyEvent = { type: "created" | "exited" | "deleted"; id: PtyID }
type PtyEvent = { type: "created" | "exited" | "deleted"; id: Pty.ID }
const locationLayer = Layer.succeed(
Location.Service,
@@ -46,7 +45,7 @@ const createPty = Effect.fn("PtySessionTest.createPty")(function* (command: stri
)
})
const waitForEvents = (events: Queue.Queue<PtyEvent>, id: PtyID, count: number) =>
const waitForEvents = (events: Queue.Queue<PtyEvent>, id: Pty.ID, count: number) =>
Effect.gen(function* () {
const picked: Array<PtyEvent["type"]> = []
while (picked.length < count) {
@@ -61,7 +60,7 @@ const waitForEvents = (events: Queue.Queue<PtyEvent>, id: PtyID, count: number)
}),
)
const attachCollecting = Effect.fn("PtySessionTest.attachCollecting")(function* (id: PtyID, cursor?: number) {
const attachCollecting = Effect.fn("PtySessionTest.attachCollecting")(function* (id: Pty.ID, cursor?: number) {
const pty = yield* Pty.Service
const output = yield* Queue.unbounded<string>()
const ended = yield* Deferred.make<{ exitCode?: number }>()
@@ -90,7 +89,7 @@ describe("pty", () => {
it.live("returns typed not found errors for missing sessions", () =>
Effect.gen(function* () {
const pty = yield* Pty.Service
const id = PtyID.make("pty_missing")
const id = Pty.ID.make("pty_missing")
for (const result of [
yield* pty.get(id).pipe(Effect.asVoid, Effect.exit),
+5 -5
View File
@@ -1,7 +1,7 @@
import { describe, expect } from "bun:test"
import { Effect, Layer } from "effect"
import { LayerNode } from "@opencode/util/effect/layer-node"
import { PtyID } from "@opencode/core/pty/schema"
import { Pty } from "@opencode/core/pty"
import { PtyTicket } from "@opencode/core/pty/ticket"
import { Workspace } from "@opencode/core/workspace"
import { testEffect } from "../lib/effect"
@@ -17,7 +17,7 @@ describe("PTY websocket tickets", () => {
it.live("consumes tickets once", () =>
Effect.gen(function* () {
const tickets = yield* PtyTicket.Service
const scope = { ptyID: PtyID.ascending(), directory: "/tmp/a" }
const scope = { ptyID: Pty.ID.ascending(), directory: "/tmp/a" }
const issued = yield* tickets.issue(scope)
expect(yield* tickets.consume({ ...scope, ticket: issued.ticket })).toBe(true)
@@ -28,7 +28,7 @@ describe("PTY websocket tickets", () => {
it.live("rejects tickets scoped to a different request", () =>
Effect.gen(function* () {
const tickets = yield* PtyTicket.Service
const ptyID = PtyID.ascending()
const ptyID = Pty.ID.ascending()
const issued = yield* tickets.issue({ ptyID, directory: "/tmp/a" })
expect(yield* tickets.consume({ ptyID, directory: "/tmp/b", ticket: issued.ticket })).toBe(false)
@@ -39,7 +39,7 @@ describe("PTY websocket tickets", () => {
itExpiring.live("rejects tickets after the TTL elapses", () =>
Effect.gen(function* () {
const tickets = yield* PtyTicket.Service
const ptyID = PtyID.ascending()
const ptyID = Pty.ID.ascending()
const issued = yield* tickets.issue({ ptyID })
yield* Effect.promise(() => new Promise((resolve) => setTimeout(resolve, 25)))
@@ -51,7 +51,7 @@ describe("PTY websocket tickets", () => {
it.live("rejects tickets scoped to a different workspace", () =>
Effect.gen(function* () {
const tickets = yield* PtyTicket.Service
const ptyID = PtyID.ascending()
const ptyID = Pty.ID.ascending()
const workspaceID = Workspace.ID.ascending()
const issued = yield* tickets.issue({ ptyID, workspaceID })
+18
View File
@@ -41,6 +41,12 @@ describe("Repository", () => {
})
})
test("caches a host with a port under a directory per host part", () => {
expect(Repository.cachePath("/cache", Repository.parseRemote("ssh://git@example.com:2222/owner/repo"))).toBe(
path.join("/cache", "example.com", "2222", "owner", "repo"),
)
})
test("keeps local file repositories distinct from remote repositories", () => {
const localPath = path.resolve("repo.git")
const reference = Repository.parse(pathToFileURL(localPath).href)
@@ -62,6 +68,18 @@ describe("Repository", () => {
expect(() => Repository.validateBranch("bad branch")).toThrow(Repository.InvalidBranchError)
})
test.each([
"..:repo",
"git@..:repo",
"../owner/repo",
"ssh://../repo",
"ssh://..:22/repo",
"https://%2e%2e/owner/repo",
".:repo",
])("rejects %s because its host contains a relative path segment", (input) => {
expect(() => Repository.parseRemote(input)).toThrow(Repository.InvalidReferenceError)
})
test("compares cache identity independent of input spelling", () => {
const shorthand = Repository.parseRemote("owner/repo")
@@ -445,12 +445,12 @@ it.effect("manual compaction summarizes short context instead of no-op", () =>
expect(requests).toHaveLength(1)
expect(requests[0]?.promptCacheKey).toBe(parentID)
expect(requests[0]?.http?.headers).toEqual({
"x-session-affinity": sessionID,
"X-Session-Id": sessionID,
"x-session-affinity": parentID,
"X-Session-Id": parentID,
"x-parent-session-id": parentID,
"User-Agent": App.useragent(App.make()),
"x-opencode-project": Project.ID.global,
"x-opencode-session": sessionID,
"x-opencode-session": parentID,
"x-opencode-client": "opencode",
})
expect(requests[0]?.generation).toEqual(GenerationOptions.make({ maxTokens: 20_000 }))
@@ -15,7 +15,6 @@ import { AbsolutePath } from "@opencode/core/schema"
import { Session } from "@opencode/core/session"
import { SessionExecution } from "@opencode/core/session/execution"
import { SessionRestart } from "@opencode/core/session/execution/restart"
import { UserInterruptedError } from "@opencode/core/session/error"
import { SessionEvent } from "@opencode/core/session/event"
import { SessionInbox } from "@opencode/core/session/inbox"
import { SessionMessage } from "@opencode/core/session/message"
@@ -50,10 +49,6 @@ describe("SessionExecution lifecycle", () => {
const interrupted = Effect.runSyncExit(Effect.interrupt)
expect(SessionExecution.terminal(interrupted)).toEqual({ type: "interrupted", reason: "shutdown" })
expect(SessionExecution.terminal(interrupted, "user")).toEqual({ type: "interrupted", reason: "user" })
expect(SessionExecution.terminal(Exit.fail(new UserInterruptedError()))).toEqual({
type: "interrupted",
reason: "user",
})
})
it.effect("the sweep only lists claimed top-level Sessions", () =>
+6 -1
View File
@@ -4693,7 +4693,12 @@ describe("SessionRunnerLLM", () => {
.pipe(Effect.orDie)
yield* s.runPrompt("Run child request")
expect(s.requests[0]?.http?.headers?.["x-parent-session-id"]).toBe(parentID)
expect(s.requests[0]?.http?.headers).toMatchObject({
"x-session-affinity": parentID,
"X-Session-Id": parentID,
"x-parent-session-id": parentID,
"x-opencode-session": parentID,
})
expect(s.requests[0]?.promptCacheKey).toBe(parentID)
})
@@ -1,71 +0,0 @@
import { expect } from "bun:test"
import { ProcessLock } from "@opencode/core/util/process-lock"
import { Effect } from "effect"
import fs from "node:fs/promises"
import os from "node:os"
import path from "node:path"
import { it } from "../lib/effect"
const worker = path.join(import.meta.dir, "../fixture/process-lock-worker.ts")
it.live(
"releases ownership when the scope closes",
Effect.gen(function* () {
const root = yield* temp("opencode-process-lock-")
const file = path.join(root, "service.lock")
yield* Effect.scoped(ProcessLock.acquire(file))
yield* Effect.scoped(ProcessLock.acquire(file))
}),
)
it.live(
"releases ownership when the process dies",
Effect.gen(function* () {
const root = yield* temp("opencode-process-lock-death-")
const file = path.join(root, "service.lock")
const ready = path.join(root, "ready")
const child = yield* Effect.acquireRelease(
Effect.sync(() =>
Bun.spawn([process.execPath, worker, JSON.stringify({ file, ready })], {
stdout: "ignore",
stderr: "pipe",
}),
),
(child) =>
Effect.promise(async () => {
kill(child)
await child.exited
}),
)
yield* Effect.promise(async () => {
for (let attempt = 0; attempt < 100 && !(await Bun.file(ready).exists()); attempt++) await Bun.sleep(20)
})
expect(yield* Effect.promise(() => Bun.file(ready).exists())).toBe(true)
const error = yield* Effect.scoped(ProcessLock.acquire(file)).pipe(Effect.flip)
expect(error._tag).toBe("ProcessLockHeldError")
if (process.platform !== "win32") {
process.kill(child.pid, "SIGSTOP")
const paused = yield* Effect.scoped(ProcessLock.acquire(file)).pipe(Effect.flip)
expect(paused._tag).toBe("ProcessLockHeldError")
process.kill(child.pid, "SIGCONT")
}
kill(child)
yield* Effect.promise(() => child.exited)
yield* Effect.scoped(ProcessLock.acquire(file))
}),
)
function temp(prefix: string) {
return Effect.acquireRelease(
Effect.promise(() => fs.mkdtemp(path.join(os.tmpdir(), prefix))),
(root) => Effect.promise(() => fs.rm(root, { recursive: true, force: true })),
)
}
function kill(child: Bun.Subprocess) {
if (process.platform === "win32") return child.kill()
return child.kill("SIGKILL")
}
+2 -2
View File
@@ -23,8 +23,8 @@ import type { ServerOptions } from "./options"
*
* - Database runs on the injected `DurableObjectStorage` SQLite.
* - Watcher and fff are disabled through their existing option flags; pty, fff,
* shell-parser, photon, and process-lock native modules resolve to inert
* stubs under the `workerd` bundle condition.
* shell-parser, and photon native modules resolve to inert stubs under the
* `workerd` bundle condition.
* - Bare locations use a typed no-execution-plane process spawner; FileSystem,
* FileSystemSearch, and Pty fail with a clear defect until a remote sandbox
* backs them; Snapshot and Vcs degrade to no-op results.
@@ -130,8 +130,8 @@ export function Answer(props: {
<text attributes={TextAttributes.BOLD} fg={theme.text.base} flexShrink={0}>
/btw
</text>
<text fg={theme.text.muted} wrapMode="word" flexGrow={1}>
{props.question}
<text fg={theme.text.muted} wrapMode="none" flexGrow={1} truncate>
{props.question.replace(/\s+/g, " ")}
</text>
<text fg={theme.text.muted} flexShrink={0} onMouseUp={() => dialog.clear()}>
esc
+1 -1
View File
@@ -2808,7 +2808,7 @@ function Shell(props: ToolProps) {
command={stringValue(props.input.command)}
workdir={stringValue(props.input.workdir)}
status={props.part.state.status}
background={Boolean(stringValue(props.metadata.shellID)) && props.part.state.status !== "running"}
background={props.part.state.status === "completed" && props.metadata.status === "running"}
output={stringValue(props.metadata.shellID) ? undefined : props.output}
/>
)
+4 -2
View File
@@ -61,7 +61,9 @@ if ((await editor.exited) !== 0) {
const document = await Bun.file(review).text()
const notes = document.match(/<!-- changelog:start -->\s*([\s\S]*?)\s*<!-- changelog:end -->/)
if (!notes) throw new Error("Release review is missing its changelog markers")
await Bun.write(changelog, `${notes[1].trim()}\n`)
const releaseNotes = notes[1].trim()
if (!releaseNotes) throw new Error("Release review has no changelog")
await Bun.write(changelog, `${releaseNotes}\n`)
const answer = prompt(`Trigger the ${version} release? [y/N]`)
if (answer?.trim().toLowerCase() !== "y" && answer?.trim().toLowerCase() !== "yes") {
@@ -69,7 +71,7 @@ if (answer?.trim().toLowerCase() !== "y" && answer?.trim().toLowerCase() !== "ye
process.exit(0)
}
await $`gh workflow run publish.yml --ref v2 ${input}`
await $`gh workflow run publish.yml --ref v2 ${input} -f release_notes=${releaseNotes}`
console.log(`Triggered the ${version} release`)
async function generateReview(base: string) {
@@ -5,6 +5,7 @@ import Card from "./Card.astro"
import CardGroup from "./CardGroup.astro"
import CodeBlock from "./CodeBlock.astro"
import CodeTabs from "./CodeTabs.astro"
import PlanTabs from "./PlanTabs.astro"
import DocsLayout from "../layouts/DocsLayout.astro"
interface Props {
@@ -22,5 +23,5 @@ const rendered = await render(entry)
headings={rendered.headings}
showTableOfContents={entry.data.tableOfContents !== false}
>
<rendered.Content components={{ Callout, Card, CardGroup, CodeBlock, CodeTabs }} />
<rendered.Content components={{ Callout, Card, CardGroup, CodeBlock, CodeTabs, PlanTabs }} />
</DocsLayout>
@@ -0,0 +1,103 @@
---
interface Props {
id: string
label: string
syncKey: string
}
const plans = [
{ id: "go", label: "Go" },
{ id: "go-plus", label: "Go Plus" },
] as const
---
<div class="docs-plan-tabs" data-plan-tabs data-sync-key={Astro.props.syncKey}>
<div class="docs-plan-tabs-list" role="tablist" aria-label={Astro.props.label}>
{
plans.map((plan, index) => (
<button
type="button"
role="tab"
id={`${Astro.props.id}-tab-${plan.id}`}
aria-controls={`${Astro.props.id}-panel-${plan.id}`}
aria-selected={index === 0 ? "true" : "false"}
tabindex={index === 0 ? 0 : -1}
data-plan-tab={plan.id}
>
{plan.label}
</button>
))
}
</div>
{
plans.map((plan, index) => (
<div
role="tabpanel"
id={`${Astro.props.id}-panel-${plan.id}`}
aria-labelledby={`${Astro.props.id}-tab-${plan.id}`}
hidden={index !== 0}
data-plan-panel={plan.id}
>
<slot name={plan.id} />
</div>
))
}
</div>
<script>
const selectPlan = (tabs: HTMLElement, plan: string) => {
tabs.querySelectorAll<HTMLButtonElement>("[data-plan-tab]").forEach((button) => {
const selected = button.dataset.planTab === plan
button.setAttribute("aria-selected", String(selected))
button.tabIndex = selected ? 0 : -1
})
tabs.querySelectorAll<HTMLElement>("[data-plan-panel]").forEach((panel) => {
panel.hidden = panel.dataset.planPanel !== plan
})
}
const selectSyncedPlan = (source: HTMLElement, plan: string) => {
const syncKey = source.dataset.syncKey
if (!syncKey) return
document.querySelectorAll<HTMLElement>("[data-plan-tabs]").forEach((tabs) => {
if (tabs.dataset.syncKey === syncKey) selectPlan(tabs, plan)
})
localStorage.setItem(`docs-plan-tabs:${syncKey}`, plan)
}
document.querySelectorAll<HTMLElement>("[data-plan-tabs]").forEach((tabs) => {
const syncKey = tabs.dataset.syncKey
const plan = syncKey ? localStorage.getItem(`docs-plan-tabs:${syncKey}`) : undefined
if (plan) selectPlan(tabs, plan)
})
document.addEventListener("click", (event) => {
if (!(event.target instanceof Element)) return
const button = event.target.closest<HTMLButtonElement>("[data-plan-tab]")
const tabs = button?.closest<HTMLElement>("[data-plan-tabs]")
if (!button || !tabs || !button.dataset.planTab) return
selectSyncedPlan(tabs, button.dataset.planTab)
})
document.addEventListener("keydown", (event) => {
if (!(event.target instanceof HTMLButtonElement) || !event.target.matches("[data-plan-tab]")) return
const tabs = event.target.closest<HTMLElement>("[data-plan-tabs]")
if (!tabs) return
const buttons = [...tabs.querySelectorAll<HTMLButtonElement>("[data-plan-tab]")]
const selected = buttons.indexOf(event.target)
const next =
event.key === "Home"
? buttons[0]
: event.key === "End"
? buttons.at(-1)
: event.key === "ArrowRight"
? buttons[(selected + 1) % buttons.length]
: event.key === "ArrowLeft"
? buttons[(selected - 1 + buttons.length) % buttons.length]
: undefined
if (!next?.dataset.planTab) return
event.preventDefault()
selectSyncedPlan(tabs, next.dataset.planTab)
next.focus()
})
</script>
+184 -115
View File
@@ -1,17 +1,22 @@
---
title: "Go"
description: "Low cost subscription for open coding models."
description: "Reliable access to open coding models with two usage tiers."
---
OpenCode Go is a low cost **$10/month subscription** that gives you reliable access to popular open coding models.
OpenCode Go gives you reliable access to popular open coding models, with two monthly plans:
| Plan | Price | Included usage |
| ---- | ----- | -------------- |
| **Go** | **$10/month** | Lower-cost access to the models below |
| **Go Plus** | **$40/month** | Higher usage limits across the models below |
Go works like any other provider in OpenCode. You subscribe to OpenCode Go and get your API key. It's **completely optional** and you don't need it to use OpenCode.
It is designed primarily for international users and provides stable global access.
The service is designed primarily for international users and provides stable global access.
## How it works
1. Sign in to the [OpenCode console](https://opencode.ai/console), subscribe to Go, add your billing details, and copy your API key.
1. Sign in to the [OpenCode Console](https://opencode.ai/console), subscribe to Go or Go Plus, add your billing details, and copy your API key.
2. Run `/connect` in the TUI, select **OpenCode Go**, and paste your API key.
```text
@@ -24,7 +29,7 @@ It is designed primarily for international users and provides stable global acce
/models
```
<Callout>Only one member per workspace can subscribe to OpenCode Go.</Callout>
<Callout>Only one member per workspace can subscribe to OpenCode Go or Go Plus.</Callout>
The current list of models includes:
@@ -33,13 +38,13 @@ The current list of models includes:
- **GLM-5.3-Flash**
- **GLM-5.3**
- **GLM-5.2**
- **GLM-5.1**
- **GPT 6 Luna**
- **GPT 5.6 Luna**
- **Kimi K3**
- **Kimi K2.7 Code**
- **Kimi K2.6**
- **LongCat-2.0**
- **LongCat 2.5 Preview Free** (limited time)
- **MiMo-V2.6-Flash**
- **MiMo-V2.6-Pro**
- **MiMo-V2.5**
@@ -50,9 +55,7 @@ The current list of models includes:
- **Muse Spark 1.2 Contributor** ([limited regions](https://ai.developer.meta.com/legal/geographic-use-policy))
- **Qwen3.8 Max**
- **Qwen3.8 Flash**
- **Qwen3.7 Max**
- **Qwen3.7 Plus**
- **Qwen3.6 Plus**
- **DeepSeek V4.1 Flash**
- **DeepSeek V4 Pro**
- **DeepSeek V4 Flash**
@@ -60,7 +63,6 @@ The current list of models includes:
- **Hy4 preview**
- **Hy3**
- **Space Bunny Free** (limited time)
- **LongCat 2.5 Preview Free** (limited time)
The list of models may change as we test and add new ones.
@@ -109,127 +111,202 @@ investigated. The linked reports track fixes and workarounds.
## Usage limits
Usage limits are defined as monthly dollar amounts. The table below shows the
monthly limit and token costs for each model.
monthly limit for each plan and the token costs for each model. Token pricing is
the same for Go and Go Plus.
Each model has the following usage limits: 5-hour — 20% of the monthly limit;
weekly — 50%; and monthly — 100%.
For example, if a model has a $60 monthly limit, you can spend up to:
- **5-hour limit** — $12 of usage
- **Weekly limit** — $30 of usage
- **Monthly limit** — $60 of usage
Each model's monthly limit below determines how its usage counts toward those allowances.
Token prices are per 1M tokens.
<div class="docs-table-scroll" role="region" aria-label="Go model pricing" tabIndex={0}>
<PlanTabs id="go-pricing" label="Go plan" syncKey="go-plan">
<div slot="go">
| Model | Input | Output | Cached Read | Cached Write | Monthly limit |
| --------------------------------------- | ------ | ------ | ----------- | ------------ | ---------------------------------------------------- |
| GLM-5.3-Flash | $0.15 | $0.50 | $0.03 | - | **$60** |
| GLM-5.3 | $1.40 | $4.40 | $0.26 | - | **$15** |
| GLM-5.2 | $1.40 | $4.40 | $0.26 | - | **$60** |
| GLM-5.1 | $1.40 | $4.40 | $0.26 | - | **$60** |
| Kimi K3 | $3.00 | $15.00 | $0.30 | - | **$15** |
| Kimi K2.7 Code | $0.95 | $4.00 | $0.19 | - | **$60** |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | - | **$60** |
| LongCat-2.0 | $0.30 | $1.20 | $0.006 | - | **$60** |
| MiMo-V2.6-Flash | $0.14 | $0.28 | $0.0028 | - | **$60** |
| MiMo-V2.6-Pro | $0.435 | $0.87 | $0.003625 | - | **$15** |
| MiMo-V2.5 | $0.14 | $0.28 | $0.0028 | - | **$60** |
| MiMo-V2.5-Pro | $0.435 | $0.87 | $0.003625 | - | **$15** |
| MiniMax M3 | $0.30 | $1.20 | $0.06 | - | **$60** |
| MiniMax M2.7 | $0.30 | $1.20 | $0.06 | $0.375 | **$60** |
| MiniMax M2.5 | $0.30 | $1.20 | $0.06 | $0.375 | **$60** |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | $0.002 | - | **$60** |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | $0.002 | - | **$60** |
| Qwen3.8 Max | $2.00 | $6.00 | $0.25 | $2.50 | **$15** |
| Qwen3.8 Flash | $0.15 | $0.47 | $0.016 | $0.20 | **$30** |
| Qwen3.7 Max | $2.50 | $7.50 | $0.50 | $3.125 | **$30** |
| Qwen3.7 Plus (≤ 256K tokens) | $0.40 | $1.60 | $0.04 | $0.50 | **$60** |
| Qwen3.7 Plus (> 256K tokens) | $1.20 | $4.80 | $0.12 | $1.50 | **$60** |
| Qwen3.6 Plus (≤ 256K tokens) | $0.50 | $3.00 | $0.05 | $0.625 | **$60** |
| Qwen3.6 Plus (> 256K tokens) | $2.00 | $6.00 | $0.20 | $2.50 | **$60** |
| DeepSeek V4.1 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$60** |
| DeepSeek V4.1 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$60** |
| DeepSeek V4 Pro (Off-Peak) | $0.66 | $1.98 | $0.022 | - | **$15** |
| DeepSeek V4 Pro (Peak) | $1.32 | $3.96 | $0.044 | - | **$15** |
| DeepSeek V4 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$30** |
| DeepSeek V4 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$30** |
| DeepSeek V4 Flash Vision Exp (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$15** |
| DeepSeek V4 Flash Vision Exp (Peak) | $0.30 | $1.20 | $0.006 | - | **$15** |
| Hy4 preview | $0.834 | $2.501 | $0.042 | - | **$30** |
| Hy3 | $0.14 | $0.58 | $0.035 | - | **$60** |
| Space Bunny Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| LongCat 2.5 Preview Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| Grok 4.7 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$15** |
| Grok 4.7 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$15** |
| Grok 4.6 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$15** |
| Grok 4.6 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$15** |
| GPT 6 Luna (≤ 272K tokens) | $0.10 | $0.50 | $0.01 | $0.125 | **$15** |
| GPT 6 Luna (> 272K tokens) | $0.20 | $0.75 | $0.02 | $0.25 | **$15** |
| GPT 5.6 Luna (≤ 272K tokens) | $0.20 | $1.20 | $0.02 | $0.25 | **$15** |
| GPT 5.6 Luna (> 272K tokens) | $0.40 | $1.80 | $0.04 | $0.50 | **$15** |
| Model | Input | Output | Cached Read | Cached Write | Monthly limit |
| --------------------------------------- | ------ | ------ | ----------- | ------------ | ---------------------------------------------- |
| GLM-5.3-Flash | $0.15 | $0.50 | $0.03 | - | **$60** |
| GLM-5.3 | $1.40 | $4.40 | $0.26 | - | **$15** |
| GLM-5.2 | $1.40 | $4.40 | $0.26 | - | **$60** |
| Kimi K3 | $3.00 | $15.00 | $0.30 | - | **$15** |
| Kimi K2.7 Code | $0.95 | $4.00 | $0.19 | - | **$60** |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | - | **$60** |
| LongCat-2.0 | $0.30 | $1.20 | $0.006 | - | **$60** |
| LongCat 2.5 Preview Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| MiMo-V2.6-Flash | $0.14 | $0.28 | $0.0028 | - | **$60** |
| MiMo-V2.6-Pro | $0.435 | $0.87 | $0.003625 | - | **$15** |
| MiMo-V2.5 | $0.14 | $0.28 | $0.0028 | - | **$60** |
| MiMo-V2.5-Pro | $0.435 | $0.87 | $0.003625 | - | **$15** |
| MiniMax M3 | $0.30 | $1.20 | $0.06 | - | **$60** |
| MiniMax M2.7 | $0.30 | $1.20 | $0.06 | $0.375 | **$60** |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | $0.002 | - | **$60** |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | $0.002 | - | **$60** |
| Qwen3.8 Max | $2.00 | $6.00 | $0.25 | $2.50 | **$15** |
| Qwen3.8 Flash | $0.15 | $0.47 | $0.016 | $0.20 | **$30** |
| Qwen3.7 Plus (≤ 256K tokens) | $0.40 | $1.60 | $0.04 | $0.50 | **$60** |
| Qwen3.7 Plus (> 256K tokens) | $1.20 | $4.80 | $0.12 | $1.50 | **$60** |
| DeepSeek V4.1 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$60** |
| DeepSeek V4.1 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$60** |
| DeepSeek V4 Pro (Off-Peak) | $0.66 | $1.98 | $0.022 | - | **$15** |
| DeepSeek V4 Pro (Peak) | $1.32 | $3.96 | $0.044 | - | **$15** |
| DeepSeek V4 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$30** |
| DeepSeek V4 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$30** |
| DeepSeek V4 Flash Vision Exp (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$15** |
| DeepSeek V4 Flash Vision Exp (Peak) | $0.30 | $1.20 | $0.006 | - | **$15** |
| Hy4 preview | $0.834 | $2.501 | $0.042 | - | **$30** |
| Hy3 | $0.14 | $0.58 | $0.035 | - | **$60** |
| Space Bunny Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| Grok 4.7 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$15** |
| Grok 4.7 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$15** |
| Grok 4.6 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$15** |
| Grok 4.6 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$15** |
| GPT 6 Luna (≤ 272K tokens) | $0.10 | $0.50 | $0.01 | $0.125 | **$15** |
| GPT 6 Luna (> 272K tokens) | $0.20 | $0.75 | $0.02 | $0.25 | **$15** |
| GPT 5.6 Luna (≤ 272K tokens) | $0.20 | $1.20 | $0.02 | $0.25 | **$15** |
| GPT 5.6 Luna (> 272K tokens) | $0.40 | $1.80 | $0.04 | $0.50 | **$15** |
</div>
</div>
<div slot="go-plus">
| Model | Input | Output | Cached Read | Cached Write | Monthly limit |
| --------------------------------------- | ------ | ------ | ----------- | ------------ | ---------------------------------------------- |
| GLM-5.3-Flash | $0.15 | $0.50 | $0.03 | - | **$180** |
| GLM-5.3 | $1.40 | $4.40 | $0.26 | - | **$120** |
| GLM-5.2 | $1.40 | $4.40 | $0.26 | - | **$180** |
| Kimi K3 | $3.00 | $15.00 | $0.30 | - | **$60** |
| Kimi K2.7 Code | $0.95 | $4.00 | $0.19 | - | **$180** |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | - | **$240** |
| LongCat-2.0 | $0.30 | $1.20 | $0.006 | - | **$240** |
| LongCat 2.5 Preview Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| MiMo-V2.6-Flash | $0.14 | $0.28 | $0.0028 | - | **$120** |
| MiMo-V2.6-Pro | $0.435 | $0.87 | $0.003625 | - | **$60** |
| MiMo-V2.5 | $0.14 | $0.28 | $0.0028 | - | **$120** |
| MiMo-V2.5-Pro | $0.435 | $0.87 | $0.003625 | - | **$60** |
| MiniMax M3 | $0.30 | $1.20 | $0.06 | - | **$180** |
| MiniMax M2.7 | $0.30 | $1.20 | $0.06 | $0.375 | **$240** |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | $0.002 | - | **$120** |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | $0.002 | - | **$120** |
| Qwen3.8 Max | $2.00 | $6.00 | $0.25 | $2.50 | **$60** |
| Qwen3.8 Flash | $0.15 | $0.47 | $0.016 | $0.20 | **$90** |
| Qwen3.7 Plus (≤ 256K tokens) | $0.40 | $1.60 | $0.04 | $0.50 | **$180** |
| Qwen3.7 Plus (> 256K tokens) | $1.20 | $4.80 | $0.12 | $1.50 | **$180** |
| DeepSeek V4.1 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$120** |
| DeepSeek V4.1 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$120** |
| DeepSeek V4 Pro (Off-Peak) | $0.66 | $1.98 | $0.022 | - | **$60** |
| DeepSeek V4 Pro (Peak) | $1.32 | $3.96 | $0.044 | - | **$60** |
| DeepSeek V4 Flash (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$120** |
| DeepSeek V4 Flash (Peak) | $0.30 | $1.20 | $0.006 | - | **$120** |
| DeepSeek V4 Flash Vision Exp (Off-Peak) | $0.15 | $0.60 | $0.003 | - | **$60** |
| DeepSeek V4 Flash Vision Exp (Peak) | $0.30 | $1.20 | $0.006 | - | **$60** |
| Hy4 preview | $0.834 | $2.501 | $0.042 | - | **$120** |
| Hy3 | $0.14 | $0.58 | $0.035 | - | **$240** |
| Space Bunny Free | Free | Free | Free | - | **Unlimited**<br /><small>limited time</small> |
| Grok 4.7 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$60** |
| Grok 4.7 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$60** |
| Grok 4.6 (≤ 200K tokens) | $2.00 | $6.00 | $0.50 | - | **$60** |
| Grok 4.6 (> 200K tokens) | $4.00 | $12.00 | $1.00 | - | **$60** |
| GPT 6 Luna (≤ 272K tokens) | $0.10 | $0.50 | $0.01 | $0.125 | **$60** |
| GPT 6 Luna (> 272K tokens) | $0.20 | $0.75 | $0.02 | $0.25 | **$60** |
| GPT 5.6 Luna (≤ 272K tokens) | $0.20 | $1.20 | $0.02 | $0.25 | **$60** |
| GPT 5.6 Luna (> 272K tokens) | $0.40 | $1.80 | $0.04 | $0.50 | **$60** |
</div>
</PlanTabs>
**Space Bunny Free:** Free for a limited time.
**LongCat 2.5 Preview Free:** Free for a limited time.
**DeepSeek V4.1 Flash / V4 Pro / V4 Flash Vision Exp:** Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; all other hours, including weekends, are Off-Peak. [Learn more](https://api-docs.deepseek.com/quick_start/pricing/).
**DeepSeek V4.1 Flash / V4 Pro / V4 Flash / V4 Flash Vision Exp:** Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday; all other hours, including weekends, are Off-Peak. [Learn more](https://api-docs.deepseek.com/quick_start/pricing/).
**DeepSeek V4 Flash Vision Exp:** Images are converted into tokens based on their dimensions and billed as input tokens alongside text tokens. [Learn more](https://api-docs.deepseek.com/quick_start/pricing/).
### Estimated requests
The table below provides an estimated request count based on typical Go usage patterns:
The tables below estimate request counts based on typical Go usage patterns. Go
Plus estimates scale with each model's higher usage limit.
<div class="docs-table-scroll" role="region" aria-label="Go estimated requests" tabIndex={0}>
<PlanTabs id="go-requests" label="Go plan" syncKey="go-plan">
<div slot="go">
| Model | requests per 5 hour | requests per week | requests per month |
| -------------------------------------------------------- | ------------------------- | -------------------------- | --------------------------- |
| GLM-5.3-Flash | 6,320 | 15,790 | 31,580 |
| GLM-5.3 | 220 | 540 | 1,080 |
| GLM-5.2 | 880 | 2,150 | 4,300 |
| GLM-5.1 | 880 | 2,150 | 4,300 |
| Kimi K3 | 110 | 250 | 490 |
| Kimi K2.7 Code | 1,350 | 3,380 | 6,750 |
| Kimi K2.6 | 1,150 | 2,880 | 5,750 |
| LongCat-2.0 | 11,400 | 28,600 | 57,200 |
| MiMo-V2.6-Flash | 30,100 | 75,200 | 150,400 |
| MiMo-V2.6-Pro | 3,250 | 8,150 | 16,300 |
| MiMo-V2.5 | 30,100 | 75,200 | 150,400 |
| MiMo-V2.5-Pro | 3,250 | 8,150 | 16,300 |
| MiniMax M3 | 3,200 | 8,000 | 16,000 |
| MiniMax M2.7 | 3,400 | 8,500 | 17,000 |
| Muse Spark 1.3 Contributor | 45,300 | 113,300 | 226,600 |
| Muse Spark 1.2 Contributor | 45,300 | 113,300 | 226,600 |
| Qwen3.8 Max | 160 | 400 | 810 |
| Qwen3.8 Flash | 5,400 | 13,500 | 27,000 |
| Qwen3.7 Max | 170 | 420 | 840 |
| Qwen3.7 Plus | 4,300 | 10,800 | 21,600 |
| Qwen3.6 Plus | 3,300 | 8,200 | 16,300 |
| DeepSeek V4.1 Flash | 26,000 | 65,000 | 130,000 |
| DeepSeek V4 Pro | 1,050 | 2,600 | 5,200 |
| DeepSeek V4 Flash | 13,000 | 32,500 | 65,000 |
| DeepSeek V4 Flash Vision Exp | 6,500 | 16,250 | 32,500 |
| Hy4 preview | 1,350 | 3,380 | 6,770 |
| Hy3 | 4,300 | 10,750 | 21,500 |
| Space Bunny Free | Unlimited | Unlimited | Unlimited |
| LongCat 2.5 Preview Free | Unlimited | Unlimited | Unlimited |
| Grok 4.7 | 169 | 423 | 845 |
| Grok 4.6 | 169 | 423 | 845 |
| GPT 6 Luna | 4,230 | 10,560 | 21,130 |
| GPT 5.6 Luna | 2,050 | 5,100 | 10,250 |
| Model | Requests per 5 hours | Requests per week | Requests per month |
| ---------------------------- | -------------------- | ----------------- | ------------------ |
| GLM-5.3-Flash | 6,320 | 15,790 | 31,580 |
| GLM-5.3 | 220 | 540 | 1,080 |
| GLM-5.2 | 880 | 2,150 | 4,300 |
| Kimi K3 | 110 | 250 | 490 |
| Kimi K2.7 Code | 1,350 | 3,380 | 6,750 |
| Kimi K2.6 | 1,150 | 2,880 | 5,750 |
| LongCat-2.0 | 11,400 | 28,600 | 57,200 |
| LongCat 2.5 Preview Free | Unlimited | Unlimited | Unlimited |
| MiMo-V2.6-Flash | 30,100 | 75,200 | 150,400 |
| MiMo-V2.6-Pro | 3,250 | 8,150 | 16,300 |
| MiMo-V2.5 | 30,100 | 75,200 | 150,400 |
| MiMo-V2.5-Pro | 3,250 | 8,150 | 16,300 |
| MiniMax M3 | 3,200 | 8,000 | 16,000 |
| MiniMax M2.7 | 3,400 | 8,500 | 17,000 |
| Muse Spark 1.3 Contributor | 45,300 | 113,300 | 226,600 |
| Muse Spark 1.2 Contributor | 45,300 | 113,300 | 226,600 |
| Qwen3.8 Max | 160 | 400 | 810 |
| Qwen3.8 Flash | 5,400 | 13,500 | 27,000 |
| Qwen3.7 Plus | 4,300 | 10,800 | 21,600 |
| DeepSeek V4.1 Flash | 26,000 | 65,000 | 130,000 |
| DeepSeek V4 Pro | 1,050 | 2,600 | 5,200 |
| DeepSeek V4 Flash | 13,000 | 32,500 | 65,000 |
| DeepSeek V4 Flash Vision Exp | 6,500 | 16,250 | 32,500 |
| Hy4 preview | 1,350 | 3,380 | 6,770 |
| Hy3 | 4,300 | 10,750 | 21,500 |
| Space Bunny Free | Unlimited | Unlimited | Unlimited |
| Grok 4.7 | 169 | 423 | 845 |
| Grok 4.6 | 169 | 423 | 845 |
| GPT 6 Luna | 4,230 | 10,560 | 21,130 |
| GPT 5.6 Luna | 2,050 | 5,100 | 10,250 |
</div>
</div>
<div slot="go-plus">
| Model | Requests per 5 hours | Requests per week | Requests per month |
| ---------------------------- | -------------------- | ----------------- | ------------------ |
| GLM-5.3-Flash | 18,960 | 47,370 | 94,740 |
| GLM-5.3 | 1,760 | 4,320 | 8,640 |
| GLM-5.2 | 2,640 | 6,450 | 12,900 |
| Kimi K3 | 440 | 1,000 | 1,960 |
| Kimi K2.7 Code | 4,050 | 10,140 | 20,250 |
| Kimi K2.6 | 4,600 | 11,520 | 23,000 |
| LongCat-2.0 | 45,600 | 114,400 | 228,800 |
| LongCat 2.5 Preview Free | Unlimited | Unlimited | Unlimited |
| MiMo-V2.6-Flash | 60,200 | 150,400 | 300,800 |
| MiMo-V2.6-Pro | 13,000 | 32,600 | 65,200 |
| MiMo-V2.5 | 60,200 | 150,400 | 300,800 |
| MiMo-V2.5-Pro | 13,000 | 32,600 | 65,200 |
| MiniMax M3 | 9,600 | 24,000 | 48,000 |
| MiniMax M2.7 | 13,600 | 34,000 | 68,000 |
| Muse Spark 1.3 Contributor | 90,600 | 226,600 | 453,200 |
| Muse Spark 1.2 Contributor | 90,600 | 226,600 | 453,200 |
| Qwen3.8 Max | 640 | 1,600 | 3,240 |
| Qwen3.8 Flash | 16,200 | 40,500 | 81,000 |
| Qwen3.7 Plus | 12,900 | 32,400 | 64,800 |
| DeepSeek V4.1 Flash | 52,000 | 130,000 | 260,000 |
| DeepSeek V4 Pro | 4,200 | 10,400 | 20,800 |
| DeepSeek V4 Flash | 52,000 | 130,000 | 260,000 |
| DeepSeek V4 Flash Vision Exp | 26,000 | 65,000 | 130,000 |
| Hy4 preview | 5,400 | 13,520 | 27,080 |
| Hy3 | 17,200 | 43,000 | 86,000 |
| Space Bunny Free | Unlimited | Unlimited | Unlimited |
| Grok 4.7 | 676 | 1,692 | 3,380 |
| Grok 4.6 | 676 | 1,692 | 3,380 |
| GPT 6 Luna | 16,920 | 42,240 | 84,520 |
| GPT 5.6 Luna | 8,200 | 20,400 | 41,000 |
</div>
</PlanTabs>
The estimates use the following token counts per request; actual usage varies.
- Grok 4.7/4.6 — 390 input, 32,500 cached, 120 output tokens per request
- GLM-5.3-Flash — 1,000 input, 55,000 cached, 200 output tokens per request
- GLM-5.3/5.2/5.1 — 700 input, 52,000 cached, 150 output tokens per request
- GLM-5.3/5.2 — 700 input, 52,000 cached, 150 output tokens per request
- GPT 6 Luna — 1,000 input, 50,000 cached, 220 output tokens per request
- GPT 5.6 Luna — 1,000 input, 50,000 cached, 220 output tokens per request
- Kimi K3 — 1,050 input, 76,500 cached, 300 output tokens per request
@@ -249,9 +326,7 @@ The estimates use the following token counts per request; actual usage varies.
- MiMo-V2.5-Pro — 790 input, 86,000 cached, 305 output tokens per request
- Qwen3.8 Max — 420 input, 66,000 cached, 200 output tokens per request
- Qwen3.8 Flash — 600 input, 58,000 cached, 200 output tokens per request
- Qwen3.7 Max — 420 input, 66,000 cached, 200 output tokens per request
- Qwen3.7 Plus — 500 input, 57,000 cached, 190 output tokens per request
- Qwen3.6 Plus — 500 input, 57,000 cached, 190 output tokens per request
- Hy4 preview — 830 input, 71,500 cached, 295 output tokens per request
- Hy3 — 830 input, 71,500 cached, 295 output tokens per request
@@ -261,6 +336,7 @@ You can track your current usage in the [console](https://opencode.ai/console).
Usage limits may change as we learn from early usage and feedback.
---
### Usage beyond limits
@@ -271,7 +347,7 @@ after you've reached your usage limits instead of blocking requests.
### Why some models have lower usage
With Go, you pay $10/month, and the included monthly usage varies by model.
With Go, the included monthly usage varies by model.
For most models, we make this work through bulk discounts and reserved GPU capacity. We then pass those savings on to you as higher monthly usage.
@@ -294,7 +370,6 @@ You can also access Go models through the following API endpoints.
| GLM-5.3-Flash | glm-5.3-flash | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| GLM-5.3 | glm-5.3 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| GLM-5.2 | glm-5.2 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| GLM-5.1 | glm-5.1 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| Kimi K3 | kimi-k3 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| Kimi K2.7 Code | kimi-k2.7-code | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| Kimi K2.6 | kimi-k2.6 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
@@ -309,14 +384,11 @@ You can also access Go models through the following API endpoints.
| MiMo-V2.5-Pro | mimo-v2.5-pro | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| MiniMax M3 | minimax-m3 | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| MiniMax M2.7 | minimax-m2.7 | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| MiniMax M2.5 | minimax-m2.5 | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Muse Spark 1.3 Contributor | muse-spark-1.3-contributor | `https://opencode.ai/zen/go/v1/responses` | `@ai-sdk/openai` |
| Muse Spark 1.2 Contributor | muse-spark-1.2-contributor | `https://opencode.ai/zen/go/v1/responses` | `@ai-sdk/openai` |
| Qwen3.8 Max | qwen3.8-max | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Qwen3.8 Flash | qwen3.8-flash | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Qwen3.7 Max | qwen3.7-max | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Qwen3.7 Plus | qwen3.7-plus | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Qwen3.6 Plus | qwen3.6-plus | `https://opencode.ai/zen/go/v1/messages` | `@ai-sdk/anthropic` |
| Hy4 preview | hy4-preview | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| Hy3 | hy3 | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
| Space Bunny Free | space-bunny-free | `https://opencode.ai/zen/go/v1/chat/completions` | `@ai-sdk/openai-compatible` |
@@ -353,7 +425,6 @@ curl https://opencode.ai/zen/go/v1/models
| GLM-5.3-Flash | Not used | 0 days |
| GLM-5.3 | Not used | 0 days |
| GLM-5.2 | Not used | 0 days |
| GLM-5.1 | Not used | 0 days |
| Kimi K3 | Not used | 0 days |
| Kimi K2.7 Code | Not used | 0 days |
| Kimi K2.6 | Not used | 0 days |
@@ -364,9 +435,7 @@ curl https://opencode.ai/zen/go/v1/models
| MiMo-V2.5 | Not used | 0 days |
| Qwen3.8 Max | Not used | 0 days |
| Qwen3.8 Flash | Not used | 0 days |
| Qwen3.7 Max | Not used | 0 days |
| Qwen3.7 Plus | Not used | 0 days |
| Qwen3.6 Plus | Not used | 0 days |
| MiniMax M3 | Not used | 0 days |
| MiniMax M2.7 | Not used | 0 days |
| Muse Spark 1.3 Contributor | Yes | Not ZDR |
@@ -384,7 +453,7 @@ curl https://opencode.ai/zen/go/v1/models
- **GPT 6 Luna / GPT 5.6 Luna:** Abuse monitoring logs are generated for all API feature usage and retained for up to 30 days. [Learn more](https://developers.openai.com/api/docs/guides/your-data#data-retention-controls-for-abuse-monitoring).
- **Muse Spark 1.3 Contributor:** Heavily discounted token pricing in exchange for permission to use your prompts and completions to train future Meta models. Availability is limited to regions permitted by Meta's [Geographic Use Policy](https://ai.developer.meta.com/legal/geographic-use-policy). [Learn more](https://dev.meta.ai/docs/pricing-rate-limits#contributor-tier).
- **Muse Spark 1.2 Contributor:** Heavily discounted token pricing in exchange for permission to use your prompts and completions to train future Meta models. Availability is limited to regions permitted by Meta's [Geographic Use Policy](https://ai.developer.meta.com/legal/geographic-use-policy). [Learn more](https://dev.meta.ai/docs/pricing-rate-limits#contributor-tier).
- **DeepSeek:** ZDR agreement is renewed monthly. The current agreement is valid through September 30, 2026.
- **DeepSeek:** ZDR agreement is renewed monthly. The current agreement is valid through October 31, 2026.
## Background
@@ -400,7 +469,7 @@ To fix this, we did a couple of things:
2. We worked with a few providers to make sure these were being served correctly.
3. We benchmarked the combination of the model/provider and came up with a list that we feel good recommending.
OpenCode Go gives you access to these models for **$10/month**.
Both Go and Go Plus provide access to these models; choose the plan that fits how much you use them.
## Goals
+42
View File
@@ -912,6 +912,48 @@ main {
overflow-x: auto;
}
.docs-plan-tabs {
margin-bottom: 1.5rem;
}
.docs-plan-tabs-list {
display: flex;
gap: 0;
margin-bottom: 1rem;
border-bottom: 1px solid var(--border);
}
.docs-plan-tabs-list button {
margin-bottom: -1px;
padding: 0.5rem 1rem;
border: 0;
border-bottom: 2px solid transparent;
background: transparent;
color: var(--muted);
font: inherit;
font-weight: 600;
cursor: pointer;
}
.docs-plan-tabs-list button[aria-selected="true"] {
border-bottom-color: var(--foreground);
color: var(--foreground);
}
.docs-plan-tabs-list button:focus-visible {
outline: 2px solid var(--link);
outline-offset: 2px;
}
.docs-plan-tabs [role="tabpanel"] {
overflow-x: auto;
}
.docs-plan-tabs [role="tabpanel"] > :last-child,
.docs-plan-tabs [role="tabpanel"] > :last-child > :last-child {
margin-bottom: 0;
}
.prose table {
width: 100%;
margin-bottom: 1.5rem;