mirror of
https://github.com/anomalyco/opencode.git
synced 2026-09-28 11:37:37 +00:00
Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
80241e9718 | ||
|
|
4b2dda60d6 |
+10
-3
@@ -789,7 +789,7 @@ await ai.write(video.video, "./kite.mp4")
|
||||
|
||||
Speech (text-to-speech) is one request whose response is parsed incrementally, so every route supports both
|
||||
`Speech.generate` (the whole file) and `Speech.stream` (audio chunks as they arrive). Models come from `.speech(...)`
|
||||
selectors on the `OpenAI`, `Google` (Gemini TTS), `ElevenLabs`, `Cartesia`, and `Deepgram` facades. Common fields
|
||||
selectors on the `OpenAI`, `Google` (Gemini TTS), `ElevenLabs`, `Cartesia`, `Deepgram`, and `XAI` facades. Common fields
|
||||
(`voice`, `format`, `speed`, `language`, `instructions`, `timestamps`) lower natively or fail with a typed `AIError`
|
||||
before any network call; provider-native controls live under `providerOptions`, inferred from the selected model.
|
||||
|
||||
@@ -854,6 +854,10 @@ Provider notes:
|
||||
- **Deepgram** Aura's voice is the model id (`aura-2-thalia-en`), so `voice` and `language` fail typed. `format`
|
||||
and `providerOptions` lower to query parameters (`encoding`, `container`, `sample_rate`, `bit_rate`); `pcm` is
|
||||
`linear16` without a container. Auth is `Authorization: Token <DEEPGRAM_API_KEY>`.
|
||||
- **xAI** (`POST /v1/tts`) has no model field, so the `.speech(...)` id (for example `"grok-tts"`) only names the
|
||||
model. `voice` is the `voice_id` (default `eve`), `language` defaults to `auto`, and `format` is the codec (`mp3`,
|
||||
`wav`, `pcm`, `mulaw`, `alaw`); `providerOptions.sampleRate` and `bitRate` complete `output_format`. Both
|
||||
`generate` and `stream` read the raw audio body. `instructions` and `timestamps` are not supported.
|
||||
|
||||
The promise client mirrors the Effect API; `ai.speech.stream` is an `AsyncIterable`.
|
||||
|
||||
@@ -873,8 +877,8 @@ for await (const event of ai.speech.stream({ model, text: "Hello from OpenCode."
|
||||
Transcription (speech-to-text) is the one modality whose providers use every route kind: OpenAI and Gemini stream,
|
||||
Deepgram answers inline, and AssemblyAI is queued. `Transcription.generate` and `Transcription.stream` work on all of
|
||||
them; `Transcription.start` / `resume` return a `Generation` on queued routes and fail with `UnsupportedOperation`
|
||||
elsewhere. Models come from `.transcription(...)` selectors on the `OpenAI`, `Google`, `Deepgram`, and `AssemblyAI`
|
||||
facades. Common fields (`language`, `prompt`, `timestamps: "none" | "segment" | "word"`, `diarize`, `speakers`) lower
|
||||
elsewhere. Models come from `.transcription(...)` selectors on the `OpenAI`, `Google`, `Deepgram`, `AssemblyAI`, and
|
||||
`XAI` facades. Common fields (`language`, `prompt`, `timestamps: "none" | "segment" | "word"`, `diarize`, `speakers`) lower
|
||||
natively or fail with a typed `AIError` before any network call; a route may return more than asked.
|
||||
|
||||
```ts
|
||||
@@ -922,6 +926,9 @@ Provider notes:
|
||||
- **Gemini** needs a transcribe model (`gemini-3.5-transcribe`); `prompt` and `speakers` fail typed.
|
||||
- **Deepgram** detects the language unless `language` is set; vocabulary goes in `providerOptions.keyterm`.
|
||||
- **AssemblyAI** uploads inline audio before submitting and is the only route that accepts `speakers`.
|
||||
- **xAI** (`grok-voice-transcribe-2.0`) answers inline and always returns words; `diarize` or `timestamps: "segment"`
|
||||
groups them into speaker-turn segments. Vocabulary goes in `providerOptions.keyterm`, and headerless PCM uploads
|
||||
send `audio_format` and `sample_rate` from `audio.info`. `prompt` and `speakers` fail typed.
|
||||
|
||||
The promise client mirrors the Effect API:
|
||||
|
||||
|
||||
@@ -199,7 +199,8 @@ hints (none of the four providers emit one). Later providers: Luma, Kling, MiniM
|
||||
#### Speech (TTS)
|
||||
|
||||
Shipped in phase 3 (`src/speech.ts`, `src/speech-client.ts`, protocols `openai-speech`, `google-speech`,
|
||||
`elevenlabs-speech`, `cartesia-speech`, `deepgram-speech`; new `ElevenLabs`, `Cartesia`, and `Deepgram` facades).
|
||||
`elevenlabs-speech`, `cartesia-speech`, `deepgram-speech`, `xai-speech`; new `ElevenLabs`, `Cartesia`, and `Deepgram`
|
||||
facades).
|
||||
|
||||
```ts
|
||||
const request = Speech.request({
|
||||
@@ -248,7 +249,7 @@ unsupported field: `UnsupportedOperation` with `operation: "media.format"`.
|
||||
|
||||
**Timestamps.** `timestamps: true` on the request asks for alignment. ElevenLabs selects the `with-timestamps`
|
||||
endpoints (character-level, NDJSON when streaming); Cartesia sets `add_timestamps` on `/tts/sse` (word-level; a
|
||||
`generate` with timestamps collects the SSE stream). OpenAI, Gemini, and Deepgram reject it.
|
||||
`generate` with timestamps collects the SSE stream). OpenAI, Gemini, Deepgram, and xAI reject it.
|
||||
|
||||
Common-field lowering per provider:
|
||||
|
||||
@@ -259,6 +260,7 @@ Common-field lowering per provider:
|
||||
| ElevenLabs | path voice id (required) | `voice_settings.speed` | `language_code` | unsupported | `with-timestamps` | `credits` from `character-cost` header |
|
||||
| Cartesia | `voice` (required) | `generation_config.speed` | `language` | unsupported | `add_timestamps` | none |
|
||||
| Deepgram | unsupported (voice is the model) | `speed` query | unsupported | unsupported | unsupported | `characters` from `dg-char-count` header |
|
||||
| xAI | `voice_id` (defaults to `eve`) | `speed` | `language` (`auto` when omitted) | unsupported | unsupported | none |
|
||||
|
||||
Deferred: `Speech.session(...)` — input-streaming TTS where text arrives incrementally over a WebSocket (ElevenLabs
|
||||
`stream-input`, Cartesia WebSocket contexts, Deepgram WebSocket speak) — is a separate scoped resource, not part of
|
||||
@@ -267,8 +269,8 @@ Deferred: `Speech.session(...)` — input-streaming TTS where text arrives incre
|
||||
#### Transcription (STT)
|
||||
|
||||
Shipped as the second half of phase 3 (`src/transcription.ts`, `src/transcription-client.ts`, protocols
|
||||
`openai-transcription`, `google-transcription`, `deepgram-transcription`, `assemblyai-transcription`; new `AssemblyAI`
|
||||
facade).
|
||||
`openai-transcription`, `google-transcription`, `deepgram-transcription`, `assemblyai-transcription`,
|
||||
`xai-transcription`; new `AssemblyAI` facade).
|
||||
|
||||
```ts
|
||||
const request = Transcription.request({
|
||||
@@ -327,6 +329,7 @@ Settled rules:
|
||||
| Gemini | stream (`generateContent` / `streamGenerateContent`) | `inlineData` or Gemini Files `fileData` | `audioTranscriptionConfig.wordTimestamp` | `audioTranscriptionConfig.diarization` | `prompt`, `speakers` | `tokens` |
|
||||
| Deepgram | inline | raw body, or JSON `{ url }` | words always; `segment` → `utterances` | `diarize_model=latest` + `utterances` | `prompt`, `speakers` | `seconds` (`metadata.duration`) |
|
||||
| AssemblyAI | queued (upload → submit → poll) | `/v2/upload` then `audio_url`, or a URL | words always; `segment` → `speaker_labels` | `speaker_labels` | — | `seconds` (`audio_duration`) |
|
||||
| xAI | inline (batch `/v1/stt`) | multipart `file` (last field), or `url` | words always; `segment` → `diarize` speaker turns | `diarize` | `prompt`, `speakers` | `seconds` (`duration`) |
|
||||
|
||||
Deferred: `Transcription.session(...)` — realtime STT over WebSocket (Deepgram live, AssemblyAI streaming, ElevenLabs
|
||||
realtime, OpenAI realtime transcription) — is the same future scoped `session` shape as input-streaming TTS and ships
|
||||
@@ -406,7 +409,7 @@ Existing facades gain per-modality selectors; the modality routes each facade pr
|
||||
|---|---|---|---|---|---|---|
|
||||
| `OpenAI` | responses (default), chat | Images API (stream) | Sora (deprecated 2026-09-24) | ✓ | ✓ | |
|
||||
| `Google` | Gemini | Gemini-native | Veo | Gemini TTS | `gemini-3.5-transcribe` | |
|
||||
| `XAI` | ✓ | ✓ | ✓ | | | |
|
||||
| `XAI` | ✓ | ✓ | ✓ | ✓ | ✓ (batch) | |
|
||||
| `ElevenLabs` | | | | ✓ | Scribe | soundEffect, music |
|
||||
| `Cartesia` | | | | ✓ | | |
|
||||
| `Deepgram` | | | | Aura | ✓ | |
|
||||
@@ -460,7 +463,7 @@ Foundation + Image ship together as the reference implementation, serially. Vide
|
||||
|
||||
1. **Foundation** — per-modality selectors, `Media`, `Generation`, `Poll`, `Usage` union, `MediaProtocol` kinds, `@opencode/ai/promise` with `llm` + `image`. Port the five existing image protocols onto it. Unify `MediaPart` and add the `media` LLM event (fixes Gemini image output being dropped).
|
||||
2. **Video** — ✅ Veo, xAI, fal, Runway shipped (`MediaProtocol.queued`, `Video.start/generate/resume/stream`, promise `ai.video`). Deferred: `Video.complete` (webhooks), Luma, Kling, MiniMax, Replicate.
|
||||
3. **Speech + Transcription** — ✅ Speech: OpenAI, Gemini TTS, ElevenLabs, Cartesia, Deepgram shipped (`MediaProtocol.stream`, `Speech.generate/stream`, promise `ai.speech`). ✅ Transcription: OpenAI, Gemini, Deepgram, AssemblyAI shipped across all three route kinds (`Transcription.generate/stream/start/resume`, promise `ai.transcription`). Pending: ElevenLabs Scribe. Deferred: `Speech.session` and `Transcription.session` (WebSocket streaming).
|
||||
3. **Speech + Transcription** — ✅ Speech: OpenAI, Gemini TTS, ElevenLabs, Cartesia, Deepgram, xAI shipped (`MediaProtocol.stream`, `Speech.generate/stream`, promise `ai.speech`). ✅ Transcription: OpenAI, Gemini, Deepgram, AssemblyAI, xAI shipped across all three route kinds (`Transcription.generate/stream/start/resume`, promise `ai.transcription`). Pending: ElevenLabs Scribe. Deferred: `Speech.session` and `Transcription.session` (WebSocket streaming).
|
||||
4. **Image queued routes and partials** — ✅ BFL, fal, Replicate, and Stability creative upscale queued; Stability generate inline; OpenAI `partial_images` streaming (`image-partial` restored). Imagen dropped: shut down on the Gemini API and discontinued on Vertex (2026-06-30). Deferred: Stability's synchronous edit and fast/conservative upscale endpoints.
|
||||
5. **Later** — ElevenLabs music/SFX, Lyria, `Speech.session` / `Transcription.session`, realtime.
|
||||
|
||||
|
||||
@@ -178,7 +178,7 @@ const fromRequest = Effect.fn("OpenAITranscription.fromRequest")(function* (requ
|
||||
{
|
||||
overlay: mergeJsonRecords(request.providerOptions, request.http?.body),
|
||||
reserved: RESERVED_FORM_FIELDS,
|
||||
repeatArrays: true,
|
||||
repeatArrays: "key[]",
|
||||
},
|
||||
)
|
||||
return MediaProtocol.multipart(form)
|
||||
|
||||
@@ -71,8 +71,8 @@ export const imageOutput = (
|
||||
}
|
||||
|
||||
/**
|
||||
* Append multipart text fields: strings as-is, other values as JSON, or arrays as repeated `key[]` parts with
|
||||
* `repeatArrays`. `overlay` keys in `reserved` are dropped so `http.body` cannot replace route-owned fields.
|
||||
* Append multipart text fields: strings as-is, other values as JSON, or arrays as repeated parts named `key[]` or
|
||||
* `key` with `repeatArrays`. `overlay` keys in `reserved` are dropped so `http.body` cannot replace route-owned fields.
|
||||
*/
|
||||
export const appendFields = (
|
||||
form: FormData,
|
||||
@@ -80,13 +80,13 @@ export const appendFields = (
|
||||
options: {
|
||||
readonly overlay?: Record<string, unknown>
|
||||
readonly reserved: ReadonlySet<string>
|
||||
readonly repeatArrays?: true
|
||||
readonly repeatArrays?: "key[]" | "key"
|
||||
},
|
||||
) => {
|
||||
const overlay = Object.entries(options.overlay ?? {}).filter(([key]) => !options.reserved.has(key))
|
||||
Object.entries(mergeJsonRecords(fields, Object.fromEntries(overlay)) ?? {}).forEach(([key, value]) => {
|
||||
if (Array.isArray(value) && options.repeatArrays)
|
||||
return value.forEach((item) => form.append(`${key}[]`, String(item)))
|
||||
if (Array.isArray(value) && options.repeatArrays !== undefined)
|
||||
return value.forEach((item) => form.append(options.repeatArrays === "key[]" ? `${key}[]` : key, String(item)))
|
||||
form.append(key, typeof value === "string" ? value : encodeJson(value))
|
||||
})
|
||||
}
|
||||
|
||||
@@ -0,0 +1,112 @@
|
||||
import { Effect } from "effect"
|
||||
import { MediaProtocol } from "../route/media-protocol.js"
|
||||
import { MediaRoute } from "../route/media.js"
|
||||
import { mergeJsonRecords } from "../schema/index.js"
|
||||
import { SpeechModel, type SpeechEvent, type SpeechRequestFor } from "../speech.js"
|
||||
import { SpeechStream } from "./utils/speech-stream.js"
|
||||
|
||||
const route = MediaProtocol.identity({ id: "xai-speech", name: "xAI Speech", provider: "xai" })
|
||||
export const DEFAULT_BASE_URL = "https://api.x.ai/v1"
|
||||
export const PATH = "/tts"
|
||||
const DEFAULT_SAMPLE_RATE = 24000
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 1. Public model input
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** `voice`, `format`, `speed`, and `language` are common request fields; other native body fields pass through. */
|
||||
export type XAISpeechOptions = {
|
||||
readonly sampleRate?: 8000 | 16000 | 22050 | 24000 | 44100 | 48000
|
||||
/** MP3 only. */
|
||||
readonly bitRate?: 32000 | 64000 | 96000 | 128000 | 192000
|
||||
readonly optimize_streaming_latency?: number
|
||||
readonly text_normalization?: boolean
|
||||
readonly replace?: Readonly<Record<string, string>>
|
||||
} & Record<string, unknown>
|
||||
|
||||
export type Request = SpeechRequestFor<XAISpeechOptions>
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 4. Parser state
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
type State = SpeechStream.Audio
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 5. Request body construction
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** `output_format.codec` values; headerless codecs map to the PCM encoding of their samples. */
|
||||
const CODECS = new Map<string, SpeechStream.PcmEncoding | undefined>([
|
||||
["mp3", undefined],
|
||||
["wav", undefined],
|
||||
["pcm", "pcm_s16le"],
|
||||
["mulaw", "pcm_mulaw"],
|
||||
["alaw", "pcm_alaw"],
|
||||
])
|
||||
|
||||
const outputFormat = Effect.fn("XAISpeech.outputFormat")(function* (request: Request) {
|
||||
const codec = request.format ?? "mp3"
|
||||
if (!CODECS.has(codec))
|
||||
return yield* route.unsupported(
|
||||
"media.format",
|
||||
`${route.name} supports the mp3, wav, pcm, mulaw, and alaw formats, not "${codec}"`,
|
||||
)
|
||||
return { codec, sample_rate: request.providerOptions?.sampleRate, bit_rate: request.providerOptions?.bitRate }
|
||||
})
|
||||
|
||||
// The TTS API has no model field, so the selected model id only names the model.
|
||||
const fromRequest = Effect.fn("XAISpeech.fromRequest")(function* (request: MediaProtocol.Addressed<Request>) {
|
||||
const { sampleRate: _sampleRate, bitRate: _bitRate, ...native } = request.providerOptions ?? {}
|
||||
return MediaProtocol.json(
|
||||
mergeJsonRecords(
|
||||
{
|
||||
text: request.text,
|
||||
voice_id: SpeechStream.voiceID(request.voice),
|
||||
// `language` is required; `auto` detects it from the text.
|
||||
language: request.language ?? "auto",
|
||||
output_format: yield* outputFormat(request),
|
||||
speed: request.speed,
|
||||
},
|
||||
native,
|
||||
request.http?.body,
|
||||
) ?? {},
|
||||
)
|
||||
})
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 6. Stream parsing
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const finish = Effect.fn("XAISpeech.finish")(function* (state: State, context: MediaProtocol.ResponseContext<Request>) {
|
||||
const format = yield* outputFormat(context.request)
|
||||
const sampleRate = format.sample_rate ?? DEFAULT_SAMPLE_RATE
|
||||
const encoding = CODECS.get(format.codec)
|
||||
return yield* SpeechStream.finish(
|
||||
route,
|
||||
state,
|
||||
encoding === undefined ? SpeechStream.container(format.codec, sampleRate) : SpeechStream.pcm(encoding, sampleRate),
|
||||
)
|
||||
})
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 7. Protocol and route
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** The response body is the raw audio in both modes, so `stream` forwards body chunks as they arrive. */
|
||||
export const protocol = MediaProtocol.stream<Request, SpeechEvent, Uint8Array, State>(route, {
|
||||
unsupported: ["instructions", "timestamps"],
|
||||
body: { from: fromRequest },
|
||||
frames: (bytes) => bytes,
|
||||
initial: () => ({ chunks: [] }),
|
||||
step: (state, frame) => Effect.succeed(SpeechStream.delta(state, frame)),
|
||||
finish,
|
||||
})
|
||||
|
||||
export const model = (input: MediaRoute.ModelInput) =>
|
||||
SpeechModel.fromRoute<XAISpeechOptions, Uint8Array, State>({ protocol, baseURL: DEFAULT_BASE_URL, path: PATH }, input)
|
||||
|
||||
export const XAISpeech = {
|
||||
protocol,
|
||||
model,
|
||||
} as const
|
||||
@@ -0,0 +1,179 @@
|
||||
import { Effect, Schema } from "effect"
|
||||
import type { HttpClientResponse } from "effect/unstable/http"
|
||||
import { MediaProtocol } from "../route/media-protocol.js"
|
||||
import { MediaRoute } from "../route/media.js"
|
||||
import { mergeJsonRecords, type OpenString } from "../schema/index.js"
|
||||
import {
|
||||
TranscriptionModel,
|
||||
TranscriptionResponse,
|
||||
type TranscriptionRequestFor,
|
||||
type TranscriptionSegment,
|
||||
type TranscriptionWord,
|
||||
} from "../transcription.js"
|
||||
import { mediaTypeExtension } from "../utils/media-type.js"
|
||||
import { ProviderShared } from "./shared.js"
|
||||
import { MediaInput } from "./utils/media-input.js"
|
||||
|
||||
const route = MediaProtocol.identity({ id: "xai-transcription", name: "xAI Transcription", provider: "xai" })
|
||||
export const DEFAULT_BASE_URL = "https://api.x.ai/v1"
|
||||
export const PATH = "/stt"
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 1. Public model input
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export type XAITranscriptionOptions = {
|
||||
/** Inverse text normalization ("one hundred dollars" → "$100"); requires `language`. */
|
||||
readonly format?: boolean
|
||||
readonly keyterm?: ReadonlyArray<string>
|
||||
readonly filler_words?: boolean
|
||||
/** Headerless audio only; derived from `audio.info.encoding` and `audio.info.sampleRate` when omitted. */
|
||||
readonly audio_format?: OpenString<"pcm" | "mulaw" | "alaw">
|
||||
readonly sample_rate?: 8000 | 16000 | 22050 | 24000 | 44100 | 48000
|
||||
readonly multichannel?: boolean
|
||||
readonly channels?: number
|
||||
readonly vad_threshold?: number
|
||||
} & Record<string, unknown>
|
||||
|
||||
export type Request = TranscriptionRequestFor<XAITranscriptionOptions>
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 2. Response schema
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const Word = Schema.Struct({
|
||||
text: Schema.String,
|
||||
start: Schema.Number,
|
||||
end: Schema.Number,
|
||||
confidence: Schema.optional(Schema.Number),
|
||||
speaker: Schema.optional(Schema.Number),
|
||||
})
|
||||
|
||||
const SttResponse = Schema.Struct({
|
||||
text: Schema.String,
|
||||
language: Schema.optional(Schema.String),
|
||||
duration: Schema.optional(Schema.Number),
|
||||
words: Schema.optional(Schema.Array(Word)),
|
||||
channels: Schema.optional(
|
||||
Schema.Array(
|
||||
Schema.Struct({
|
||||
index: Schema.Number,
|
||||
text: Schema.String,
|
||||
language: Schema.optional(Schema.String),
|
||||
words: Schema.optional(Schema.Array(Word)),
|
||||
}),
|
||||
),
|
||||
),
|
||||
})
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 5. Request body construction
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** xAI returns only words, so segments are diarized speaker turns, as with AssemblyAI utterances. */
|
||||
const wantsSegments = (request: Request) => request.diarize === true || request.timestamps === "segment"
|
||||
|
||||
const RAW_AUDIO_FORMATS: Readonly<Record<string, string>> = {
|
||||
pcm_s16le: "pcm",
|
||||
pcm_mulaw: "mulaw",
|
||||
pcm_alaw: "alaw",
|
||||
}
|
||||
|
||||
const RESERVED_FORM_FIELDS = new Set(["file", "url", "model", "language", "diarize"])
|
||||
|
||||
const fromRequest = Effect.fn("XAITranscription.fromRequest")(function* (request: Request) {
|
||||
const audioFormat = RAW_AUDIO_FORMATS[request.audio.info?.encoding ?? ""]
|
||||
const form = new FormData()
|
||||
MediaInput.appendFields(
|
||||
form,
|
||||
{
|
||||
model: request.model.id,
|
||||
language: request.language,
|
||||
diarize: wantsSegments(request) ? true : undefined,
|
||||
audio_format: audioFormat,
|
||||
sample_rate: audioFormat === undefined ? undefined : request.audio.info?.sampleRate,
|
||||
},
|
||||
{
|
||||
overlay: mergeJsonRecords(request.providerOptions, request.http?.body),
|
||||
reserved: RESERVED_FORM_FIELDS,
|
||||
repeatArrays: "key",
|
||||
},
|
||||
)
|
||||
// `file` must be the last field: options after it may be ignored for streamed uploads.
|
||||
const url = ProviderShared.mediaUrl(request.audio)
|
||||
if (url !== undefined) {
|
||||
form.append("url", url)
|
||||
return MediaProtocol.multipart(form)
|
||||
}
|
||||
const audio = yield* MediaInput.inlineBytes(route.id, request.audio)
|
||||
const extension = mediaTypeExtension(request.audio.mediaType)
|
||||
form.append(
|
||||
"file",
|
||||
MediaInput.blob(audio, request.audio.mediaType),
|
||||
extension === undefined ? "audio" : `audio.${extension}`,
|
||||
)
|
||||
return MediaProtocol.multipart(form)
|
||||
})
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 6. Response decoding
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const decodeStt = route.decodeJson(SttResponse)
|
||||
|
||||
const word = (value: typeof Word.Type): TranscriptionWord => ({
|
||||
text: value.text,
|
||||
startSeconds: value.start,
|
||||
endSeconds: value.end,
|
||||
speaker: value.speaker === undefined ? undefined : String(value.speaker),
|
||||
confidence: value.confidence,
|
||||
})
|
||||
|
||||
const speakerSegments = (words: ReadonlyArray<TranscriptionWord>) =>
|
||||
words.reduce<Array<TranscriptionSegment>>((turns, next) => {
|
||||
const last = turns.at(-1)
|
||||
if (last === undefined || last.speaker !== next.speaker)
|
||||
return [
|
||||
...turns,
|
||||
{ text: next.text, startSeconds: next.startSeconds, endSeconds: next.endSeconds, speaker: next.speaker },
|
||||
]
|
||||
turns[turns.length - 1] = { ...last, text: `${last.text} ${next.text}`, endSeconds: next.endSeconds }
|
||||
return turns
|
||||
}, [])
|
||||
|
||||
const decodeResponse = Effect.fn("XAITranscription.decodeResponse")(function* (
|
||||
response: HttpClientResponse.HttpClientResponse,
|
||||
context: MediaProtocol.DecodeContext<Request>,
|
||||
) {
|
||||
const output = yield* decodeStt(response)
|
||||
const transcript = output.value
|
||||
const words = transcript.words?.map(word)
|
||||
const duration = transcript.duration
|
||||
return new TranscriptionResponse({
|
||||
text: transcript.text,
|
||||
segments: words === undefined || !wantsSegments(context.request) ? undefined : speakerSegments(words),
|
||||
words,
|
||||
language: transcript.language?.toLowerCase(),
|
||||
durationSeconds: duration,
|
||||
usage: duration === undefined ? undefined : { type: "seconds", seconds: duration },
|
||||
providerMetadata: transcript.channels === undefined ? undefined : { xai: { channels: transcript.channels } },
|
||||
})
|
||||
})
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// 7. Protocol and route
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
export const protocol = MediaProtocol.inline<Request, TranscriptionResponse>(route, {
|
||||
unsupported: ["prompt", "speakers"],
|
||||
body: { from: fromRequest },
|
||||
response: { decode: decodeResponse },
|
||||
})
|
||||
|
||||
export const model = (input: MediaRoute.ModelInput) =>
|
||||
TranscriptionModel.fromRoute<XAITranscriptionOptions>({ protocol, baseURL: DEFAULT_BASE_URL, path: PATH }, input)
|
||||
|
||||
export const XAITranscription = {
|
||||
protocol,
|
||||
model,
|
||||
} as const
|
||||
@@ -7,6 +7,8 @@ import { OpenAIChat } from "../protocols/openai-chat.js"
|
||||
import { OpenResponsesChannel } from "../protocols/open-responses-channel.js"
|
||||
import { XAIResponses } from "../protocols/xai-responses.js"
|
||||
import { XAIImages } from "../protocols/xai-images.js"
|
||||
import { XAISpeech } from "../protocols/xai-speech.js"
|
||||
import { XAITranscription } from "../protocols/xai-transcription.js"
|
||||
import { XAIVideo } from "../protocols/xai-video.js"
|
||||
import type { OpenAIOptionsInput } from "./openai-options.js"
|
||||
import type { ProviderPackage } from "../provider-package.js"
|
||||
@@ -29,6 +31,8 @@ export type Settings = ProviderPackage.Settings &
|
||||
}
|
||||
|
||||
export type { XAIImageOptions } from "../protocols/xai-images.js"
|
||||
export type { XAISpeechOptions } from "../protocols/xai-speech.js"
|
||||
export type { XAITranscriptionOptions } from "../protocols/xai-transcription.js"
|
||||
export type { XAIVideoOptions } from "../protocols/xai-video.js"
|
||||
|
||||
const RESPONSES_WEBSOCKET_ROTATE_AFTER_MS = 24 * 60 * 1000
|
||||
@@ -98,6 +102,8 @@ export const configure = (input: LanguageModelOptions = {}) => {
|
||||
chat,
|
||||
image: (modelID: string | ModelID) => XAIImages.model({ ...media, id: modelID }),
|
||||
video: (modelID: string | ModelID) => XAIVideo.model({ ...media, id: modelID }),
|
||||
speech: (modelID: string | ModelID) => XAISpeech.model({ ...media, id: modelID }),
|
||||
transcription: (modelID: string | ModelID) => XAITranscription.model({ ...media, id: modelID }),
|
||||
configure,
|
||||
}
|
||||
}
|
||||
@@ -119,3 +125,5 @@ export const responses = provider.responses
|
||||
export const chat = provider.chat
|
||||
export const image = provider.image
|
||||
export const video = provider.video
|
||||
export const speech = provider.speech
|
||||
export const transcription = provider.transcription
|
||||
|
||||
@@ -88,15 +88,19 @@ const decodeProviderBody = Schema.decodeUnknownOption(
|
||||
Schema.fromJsonString(
|
||||
Schema.Struct({
|
||||
message: Schema.optionalKey(Schema.String),
|
||||
error: Schema.optionalKey(Schema.Struct({ message: Schema.optionalKey(Schema.String) })),
|
||||
// xAI sends `{ code, error }` with the readable reason as a plain string.
|
||||
error: Schema.optionalKey(
|
||||
Schema.Union([Schema.String, Schema.Struct({ message: Schema.optionalKey(Schema.String) })]),
|
||||
),
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
const providerMessage = (status: number, body: string | void) => {
|
||||
const decoded = body === undefined ? undefined : Option.getOrUndefined(decodeProviderBody(body))
|
||||
const error = typeof decoded?.error === "string" ? decoded.error : decoded?.error?.message
|
||||
return (
|
||||
[decoded?.error?.message, decoded?.message].find((message) => message?.trim()) ??
|
||||
[error, decoded?.message].find((message) => message?.trim()) ??
|
||||
`Provider request failed with HTTP ${status}`
|
||||
)
|
||||
}
|
||||
|
||||
@@ -299,6 +299,26 @@ describe("RequestExecutor", () => {
|
||||
),
|
||||
)
|
||||
|
||||
it.effect("reads provider messages sent as a plain error string", () =>
|
||||
Effect.gen(function* () {
|
||||
const executor = yield* RequestExecutor.Service
|
||||
const error = yield* executor.execute(request).pipe(Effect.flip)
|
||||
|
||||
expectAIError(error)
|
||||
expect(error.message).toBe("Your team has no credits for this endpoint")
|
||||
}).pipe(
|
||||
Effect.provide(
|
||||
fixedResponse(
|
||||
JSON.stringify({
|
||||
code: "The caller does not have permission to execute the specified operation",
|
||||
error: "Your team has no credits for this endpoint",
|
||||
}),
|
||||
{ status: 403 },
|
||||
),
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
it.effect("falls back when structured provider messages are empty", () =>
|
||||
Effect.gen(function* () {
|
||||
const executor = yield* RequestExecutor.Service
|
||||
|
||||
@@ -152,6 +152,10 @@ describe("public exports", () => {
|
||||
expect(XAI.configure({ apiKey: "fixture" }).responses("grok-4.3").route.id).toBe("openai-responses")
|
||||
expect(XAI.configure({ apiKey: "fixture" }).chat("grok-4.3").route.id).toBe("openai-compatible-chat")
|
||||
expect(XAI.configure({ apiKey: "fixture" }).video("grok-imagine-video-1.5").route.id).toBe("xai-video")
|
||||
expect(XAI.configure({ apiKey: "fixture" }).speech("grok-tts").route.id).toBe("xai-speech")
|
||||
expect(XAI.configure({ apiKey: "fixture" }).transcription("grok-voice-transcribe-2.0").route.kind).toBe("inline")
|
||||
expect(XAI.provider.speech).toBe(XAI.speech)
|
||||
expect(XAI.provider.transcription).toBe(XAI.transcription)
|
||||
expect(Fal.configure({ apiKey: "fixture" }).video("fal-ai/veo3.1").route.id).toBe("fal-video")
|
||||
expect(Runway.configure({ apiKey: "fixture" }).video("gen4.5").route.id).toBe("runway-video")
|
||||
expect(Runway.provider.video).toBe(Runway.video)
|
||||
|
||||
@@ -0,0 +1,37 @@
|
||||
import { describe, expect } from "bun:test"
|
||||
import { Effect } from "effect"
|
||||
import { Speech } from "../../src/index.js"
|
||||
import { XAI } from "../../src/providers.js"
|
||||
import { recordedTests } from "../recorded-test.js"
|
||||
import { TEXT, collectSpeech } from "./speech-recording.js"
|
||||
|
||||
const model = XAI.configure({ apiKey: process.env.XAI_API_KEY ?? "fixture" }).speech("grok-tts")
|
||||
|
||||
const recorded = recordedTests({
|
||||
prefix: "xai-speech",
|
||||
provider: "xai",
|
||||
protocol: "xai-speech",
|
||||
requires: ["XAI_API_KEY"],
|
||||
})
|
||||
|
||||
describe("xAI Speech recorded", () => {
|
||||
recorded.effect("generates speech", () =>
|
||||
Effect.gen(function* () {
|
||||
const response = yield* Speech.generate({ model, text: TEXT, voice: "eve", language: "en" })
|
||||
|
||||
expect(response.audio.mediaType).toBe("audio/mpeg")
|
||||
expect((yield* response.audio.bytes()).length).toBeGreaterThan(0)
|
||||
}),
|
||||
)
|
||||
|
||||
recorded.effect("streams speech", () =>
|
||||
Effect.gen(function* () {
|
||||
const { finish } = yield* collectSpeech(
|
||||
Speech.stream({ model, text: TEXT, voice: "eve", language: "en", format: "pcm" }),
|
||||
)
|
||||
|
||||
expect(finish.audio.mediaType).toBe("audio/pcm")
|
||||
expect(finish.audio.info).toMatchObject({ encoding: "pcm_s16le", sampleRate: 24000, channels: 1 })
|
||||
}),
|
||||
)
|
||||
})
|
||||
@@ -0,0 +1,146 @@
|
||||
import { describe, expect } from "bun:test"
|
||||
import { Effect, Layer, Stream } from "effect"
|
||||
import { Speech, SpeechClient, SpeechEvent } from "../../src/index.js"
|
||||
import { XAI } from "../../src/providers.js"
|
||||
import { it } from "../lib/effect.js"
|
||||
import { dynamicResponse, type Call, type Handler, observe } from "../lib/http.js"
|
||||
|
||||
const layer = (handler: Handler) => SpeechClient.layer.pipe(Layer.provideMerge(dynamicResponse(handler)))
|
||||
|
||||
const xai = XAI.configure({ apiKey: "test", baseURL: "https://api.xai.test/v1" })
|
||||
const model = xai.speech("grok-tts")
|
||||
|
||||
const respondAudio = (calls: Array<Call>, body: Uint8Array | ReadableStream<Uint8Array>, contentType: string) =>
|
||||
layer((input) =>
|
||||
observe(calls, input).pipe(Effect.map(() => input.respond(body, { headers: { "content-type": contentType } }))),
|
||||
)
|
||||
|
||||
describe("xAI Speech", () => {
|
||||
it.effect("lowers common fields into the TTS body and describes the requested container", () => {
|
||||
const calls: Array<Call> = []
|
||||
return Effect.gen(function* () {
|
||||
const wav = Uint8Array.from([0x52, 0x49, 0x46, 0x46, 0, 0, 0, 0, 0x57, 0x41, 0x56, 0x45])
|
||||
const response = yield* Speech.generate({
|
||||
model,
|
||||
text: "Hello from OpenCode.",
|
||||
voice: { id: "nlbqfwie" },
|
||||
speed: 1.2,
|
||||
language: "en",
|
||||
format: "wav",
|
||||
providerOptions: { sampleRate: 44100, text_normalization: true },
|
||||
http: { body: { replace: { OpenCode: "Open Code" } } },
|
||||
}).pipe(Effect.provide(respondAudio(calls, wav, "audio/wav")))
|
||||
|
||||
expect(calls.map((call) => [call.method, call.url, call.headers.get("authorization")])).toEqual([
|
||||
["POST", "https://api.xai.test/v1/tts", "Bearer test"],
|
||||
])
|
||||
expect(JSON.parse(calls[0].body)).toEqual({
|
||||
text: "Hello from OpenCode.",
|
||||
voice_id: "nlbqfwie",
|
||||
language: "en",
|
||||
output_format: { codec: "wav", sample_rate: 44100 },
|
||||
speed: 1.2,
|
||||
text_normalization: true,
|
||||
replace: { OpenCode: "Open Code" },
|
||||
})
|
||||
expect(response.audio.mediaType).toBe("audio/wav")
|
||||
expect(response.audio.info).toEqual({ format: "wav", sampleRate: 44100 })
|
||||
expect(yield* response.audio.bytes()).toEqual(wav)
|
||||
expect(response.usage).toBeUndefined()
|
||||
})
|
||||
})
|
||||
|
||||
it.effect("defaults to auto-detected language and 24 kHz MP3", () => {
|
||||
const calls: Array<Call> = []
|
||||
return Effect.gen(function* () {
|
||||
const response = yield* Speech.generate({ model, text: "Hi" }).pipe(
|
||||
Effect.provide(respondAudio(calls, Uint8Array.from([0x49, 0x44, 0x33, 4]), "audio/mpeg")),
|
||||
)
|
||||
|
||||
expect(JSON.parse(calls[0].body)).toEqual({ text: "Hi", language: "auto", output_format: { codec: "mp3" } })
|
||||
expect(response.audio.mediaType).toBe("audio/mpeg")
|
||||
expect(response.audio.info).toEqual({ format: "mp3", sampleRate: 24000 })
|
||||
})
|
||||
})
|
||||
|
||||
it.effect("streams raw body chunks as audio deltas and describes headerless PCM", () => {
|
||||
const calls: Array<Call> = []
|
||||
return Effect.gen(function* () {
|
||||
const body = new ReadableStream<Uint8Array>({
|
||||
start(controller) {
|
||||
controller.enqueue(Uint8Array.from([1, 2]))
|
||||
controller.enqueue(Uint8Array.from([3, 4, 5]))
|
||||
controller.close()
|
||||
},
|
||||
})
|
||||
const events = Array.from(
|
||||
yield* Stream.runCollect(Speech.stream({ model, text: "Hi", voice: "eve", format: "pcm" })).pipe(
|
||||
Effect.provide(respondAudio(calls, body, "audio/pcm")),
|
||||
),
|
||||
)
|
||||
|
||||
expect(JSON.parse(calls[0].body)).toEqual({
|
||||
text: "Hi",
|
||||
voice_id: "eve",
|
||||
language: "auto",
|
||||
output_format: { codec: "pcm" },
|
||||
})
|
||||
expect(events.map((event) => event.type)).toEqual(["audio-delta", "audio-delta", "finish"])
|
||||
expect(events.filter(SpeechEvent.is.audioDelta).map((event) => Array.from(event.chunk))).toEqual([
|
||||
[1, 2],
|
||||
[3, 4, 5],
|
||||
])
|
||||
const finish = events.find(SpeechEvent.is.finish)
|
||||
expect(finish?.audio.mediaType).toBe("audio/pcm")
|
||||
expect(finish?.audio.info).toEqual({ format: "pcm", encoding: "pcm_s16le", sampleRate: 24000, channels: 1 })
|
||||
expect(yield* finish!.audio.bytes()).toEqual(Uint8Array.from([1, 2, 3, 4, 5]))
|
||||
})
|
||||
})
|
||||
|
||||
it.effect("describes telephony codecs at the requested sample rate", () =>
|
||||
Effect.gen(function* () {
|
||||
const response = yield* Speech.generate({
|
||||
model,
|
||||
text: "Hi",
|
||||
format: "mulaw",
|
||||
providerOptions: { sampleRate: 8000 },
|
||||
}).pipe(Effect.provide(respondAudio([], Uint8Array.from([0xff, 0x7f]), "audio/basic")))
|
||||
|
||||
expect(response.audio.mediaType).toBe("audio/mulaw")
|
||||
expect(response.audio.info).toEqual({ format: "pcm", encoding: "pcm_mulaw", sampleRate: 8000, channels: 1 })
|
||||
}),
|
||||
)
|
||||
|
||||
it.effect("rejects what xAI cannot lower before sending anything", () =>
|
||||
Effect.gen(function* () {
|
||||
const errors = yield* Effect.all(
|
||||
[
|
||||
Speech.generate({ model, text: "Hi", instructions: "Warm." }),
|
||||
Speech.generate({ model, text: "Hi", timestamps: true }),
|
||||
Stream.runCollect(Speech.stream({ model, text: "Hi", format: "opus" })),
|
||||
].map((effect) => Effect.flip(effect)),
|
||||
)
|
||||
|
||||
expect(errors.map((error) => [error.reason._tag, "operation" in error.reason && error.reason.operation])).toEqual(
|
||||
[
|
||||
["UnsupportedOperation", "media.instructions"],
|
||||
["UnsupportedOperation", "media.timestamps"],
|
||||
["UnsupportedOperation", "media.format"],
|
||||
],
|
||||
)
|
||||
expect(errors[2].reason).toMatchObject({ provider: "xai", route: "xai-speech" })
|
||||
}).pipe(Effect.provide(layer(() => Effect.die("an unsupported request reached the network")))),
|
||||
)
|
||||
|
||||
it.effect("fails typed when the provider returns no audio", () =>
|
||||
Effect.gen(function* () {
|
||||
const error = yield* Speech.generate({ model, text: "Hi" }).pipe(
|
||||
Effect.provide(respondAudio([], new Uint8Array(), "audio/mpeg")),
|
||||
Effect.flip,
|
||||
)
|
||||
|
||||
expect(error.reason).toMatchObject({ _tag: "InvalidProviderOutput", route: "xai-speech" })
|
||||
expect(error.reason.http?.status).toBe(200)
|
||||
}),
|
||||
)
|
||||
})
|
||||
@@ -0,0 +1,39 @@
|
||||
import { describe, expect } from "bun:test"
|
||||
import { Effect } from "effect"
|
||||
import { Transcription } from "../../src/index.js"
|
||||
import { XAI } from "../../src/providers.js"
|
||||
import { recordedTests } from "../recorded-test.js"
|
||||
import { TRANSCRIPT, audio, audioRecording, dialog } from "./transcription-recording.js"
|
||||
|
||||
const model = XAI.configure({ apiKey: process.env.XAI_API_KEY ?? "fixture" }).transcription("grok-voice-transcribe-2.0")
|
||||
|
||||
const recorded = recordedTests({
|
||||
prefix: "xai-transcription",
|
||||
provider: "xai",
|
||||
protocol: "xai-transcription",
|
||||
requires: ["XAI_API_KEY"],
|
||||
options: audioRecording,
|
||||
})
|
||||
|
||||
describe("xAI Transcription recorded", () => {
|
||||
recorded.effect("transcribes with word timestamps", () =>
|
||||
Effect.gen(function* () {
|
||||
const response = yield* Transcription.generate({ model, audio: yield* audio, language: "en" })
|
||||
|
||||
expect(response.text).toMatch(TRANSCRIPT)
|
||||
expect(response.words?.length).toBeGreaterThan(2)
|
||||
expect(response.words?.every((word) => word.endSeconds >= word.startSeconds)).toBe(true)
|
||||
expect(response.language).toBe("en")
|
||||
expect(response.usage).toEqual({ type: "seconds", seconds: response.durationSeconds })
|
||||
}),
|
||||
)
|
||||
|
||||
recorded.effect("diarizes speakers into segments", () =>
|
||||
Effect.gen(function* () {
|
||||
const response = yield* Transcription.generate({ model, audio: yield* dialog, diarize: true })
|
||||
|
||||
expect(new Set(response.segments?.map((segment) => segment.speaker)).size).toBe(2)
|
||||
expect(response.words?.every((word) => word.speaker !== undefined)).toBe(true)
|
||||
}),
|
||||
)
|
||||
})
|
||||
@@ -0,0 +1,168 @@
|
||||
import { describe, expect } from "bun:test"
|
||||
import { Effect, Layer, Stream } from "effect"
|
||||
import { HttpClientRequest } from "effect/unstable/http"
|
||||
import { Media, Transcription, TranscriptionClient } from "../../src/index.js"
|
||||
import { XAI } from "../../src/providers.js"
|
||||
import { it } from "../lib/effect.js"
|
||||
import { dynamicResponse, json, type Handler } from "../lib/http.js"
|
||||
|
||||
const layer = (handler: Handler) => TranscriptionClient.layer.pipe(Layer.provideMerge(dynamicResponse(handler)))
|
||||
|
||||
const model = XAI.configure({ apiKey: "test", baseURL: "https://api.xai.test/v1" }).transcription(
|
||||
"grok-voice-transcribe-2.0",
|
||||
)
|
||||
const audio = Media.bytes(Uint8Array.from([0x49, 0x44, 0x33, 1, 2, 3]), "audio/mpeg")
|
||||
|
||||
interface Upload {
|
||||
readonly url: string
|
||||
readonly authorization: string | null
|
||||
readonly form: FormData
|
||||
}
|
||||
|
||||
/** Reply with `body` and keep every request's parsed multipart form. */
|
||||
const respondTranscript = (uploads: Array<Upload>, body: unknown) =>
|
||||
layer((input) =>
|
||||
Effect.gen(function* () {
|
||||
const web = yield* HttpClientRequest.toWeb(input.request).pipe(Effect.orDie)
|
||||
const form = yield* Effect.promise(() => web.formData())
|
||||
uploads.push({ url: web.url, authorization: web.headers.get("authorization"), form })
|
||||
return json(input, body)
|
||||
}),
|
||||
)
|
||||
|
||||
describe("xAI Transcription", () => {
|
||||
it.effect("uploads inline audio as multipart with file last and splits diarized words into speaker turns", () => {
|
||||
const uploads: Array<Upload> = []
|
||||
return Effect.gen(function* () {
|
||||
const request = Transcription.request({
|
||||
model,
|
||||
audio,
|
||||
language: "en",
|
||||
diarize: true,
|
||||
timestamps: "word",
|
||||
providerOptions: { format: true, keyterm: ["OpenCode", "Grok"], filler_words: false },
|
||||
http: { body: { vad_threshold: 0.3, model: "ignored" } },
|
||||
})
|
||||
const response = yield* Transcription.generate(request)
|
||||
const events = Array.from(yield* Stream.runCollect(Transcription.stream(request)))
|
||||
|
||||
const upload = uploads[0]
|
||||
expect([upload.url, upload.authorization]).toEqual(["https://api.xai.test/v1/stt", "Bearer test"])
|
||||
expect(Array.from(upload.form.keys())).toEqual([
|
||||
"model",
|
||||
"language",
|
||||
"diarize",
|
||||
"format",
|
||||
"keyterm",
|
||||
"keyterm",
|
||||
"filler_words",
|
||||
"vad_threshold",
|
||||
"file",
|
||||
])
|
||||
expect(upload.form.get("model")).toBe("grok-voice-transcribe-2.0")
|
||||
expect(upload.form.get("language")).toBe("en")
|
||||
expect(upload.form.get("diarize")).toBe("true")
|
||||
expect(upload.form.get("format")).toBe("true")
|
||||
expect(upload.form.getAll("keyterm")).toEqual(["OpenCode", "Grok"])
|
||||
expect(upload.form.get("filler_words")).toBe("false")
|
||||
expect(upload.form.get("vad_threshold")).toBe("0.3")
|
||||
const file = upload.form.get("file")
|
||||
if (!(file instanceof File)) throw new Error("Expected a file upload")
|
||||
expect([file.name, file.type]).toEqual(["audio.mp3", "audio/mpeg"])
|
||||
expect(new Uint8Array(yield* Effect.promise(() => file.arrayBuffer()))).toEqual(yield* audio.bytes())
|
||||
|
||||
expect(response.text).toBe("Did it ship? Yes.")
|
||||
expect(response.words).toEqual([
|
||||
{ text: "Did", startSeconds: 0.2, endSeconds: 0.4, speaker: "0", confidence: 0.9 },
|
||||
{ text: "it", startSeconds: 0.4, endSeconds: 0.5, speaker: "0", confidence: undefined },
|
||||
{ text: "ship?", startSeconds: 0.5, endSeconds: 0.9, speaker: "0", confidence: undefined },
|
||||
{ text: "Yes.", startSeconds: 1.2, endSeconds: 1.6, speaker: "1", confidence: 0.8 },
|
||||
])
|
||||
expect(response.segments).toEqual([
|
||||
{ text: "Did it ship?", startSeconds: 0.2, endSeconds: 0.9, speaker: "0" },
|
||||
{ text: "Yes.", startSeconds: 1.2, endSeconds: 1.6, speaker: "1" },
|
||||
])
|
||||
expect(response).toMatchObject({
|
||||
language: "en",
|
||||
durationSeconds: 1.75,
|
||||
usage: { type: "seconds", seconds: 1.75 },
|
||||
})
|
||||
expect(events.map((event) => event.type)).toEqual(["finish"])
|
||||
}).pipe(
|
||||
Effect.provide(
|
||||
respondTranscript(uploads, {
|
||||
text: "Did it ship? Yes.",
|
||||
language: "EN",
|
||||
duration: 1.75,
|
||||
words: [
|
||||
{ text: "Did", start: 0.2, end: 0.4, confidence: 0.9, speaker: 0 },
|
||||
{ text: "it", start: 0.4, end: 0.5, speaker: 0 },
|
||||
{ text: "ship?", start: 0.5, end: 0.9, speaker: 0 },
|
||||
{ text: "Yes.", start: 1.2, end: 1.6, confidence: 0.8, speaker: 1 },
|
||||
],
|
||||
}),
|
||||
),
|
||||
)
|
||||
})
|
||||
|
||||
it.effect("sends remote audio by URL and describes headerless PCM uploads", () => {
|
||||
const uploads: Array<Upload> = []
|
||||
return Effect.gen(function* () {
|
||||
const remote = yield* Transcription.generate({
|
||||
model,
|
||||
audio: Media.url("https://cdn.test/call.mp3", { mediaType: "audio/mpeg" }),
|
||||
timestamps: "segment",
|
||||
})
|
||||
const pcm = yield* Transcription.generate({
|
||||
model,
|
||||
audio: Media.bytes(Uint8Array.from([0, 1, 0, 1]), "audio/pcm", {
|
||||
info: { format: "pcm", encoding: "pcm_s16le", sampleRate: 16000, channels: 1 },
|
||||
}),
|
||||
})
|
||||
|
||||
expect(Array.from(uploads[0].form.entries())).toEqual([
|
||||
["model", "grok-voice-transcribe-2.0"],
|
||||
["diarize", "true"],
|
||||
["url", "https://cdn.test/call.mp3"],
|
||||
])
|
||||
expect(Array.from(uploads[1].form.keys())).toEqual(["model", "audio_format", "sample_rate", "file"])
|
||||
expect(uploads[1].form.get("audio_format")).toBe("pcm")
|
||||
expect(uploads[1].form.get("sample_rate")).toBe("16000")
|
||||
expect(remote.segments).toBeUndefined()
|
||||
expect(pcm).toMatchObject({ text: "Hi", words: undefined, segments: undefined })
|
||||
}).pipe(Effect.provide(respondTranscript(uploads, { text: "Hi", language: "en", duration: 0.5 })))
|
||||
})
|
||||
|
||||
it.effect("rejects what xAI cannot lower before sending anything", () =>
|
||||
Effect.gen(function* () {
|
||||
const errors = yield* Effect.all(
|
||||
[
|
||||
Transcription.generate({ model, audio, prompt: "OpenCode" }),
|
||||
Transcription.generate({ model, audio, speakers: 2 }),
|
||||
Transcription.start({ model, audio }),
|
||||
Transcription.generate({ model, audio: Media.ref("file_1", { provider: "xai", mediaType: "audio/mpeg" }) }),
|
||||
].map((effect) => Effect.flip(effect)),
|
||||
)
|
||||
|
||||
expect(errors.map((error) => [error.reason._tag, "operation" in error.reason && error.reason.operation])).toEqual(
|
||||
[
|
||||
["UnsupportedOperation", "media.prompt"],
|
||||
["UnsupportedOperation", "media.speakers"],
|
||||
["UnsupportedOperation", "transcription.start"],
|
||||
["InvalidRequest", false],
|
||||
],
|
||||
)
|
||||
expect(errors[0].reason).toMatchObject({ provider: "xai", route: "xai-transcription" })
|
||||
}).pipe(Effect.provide(layer(() => Effect.die("an unsupported request reached the network")))),
|
||||
)
|
||||
|
||||
it.effect("keeps the raw body when the transcript cannot be decoded", () => {
|
||||
const uploads: Array<Upload> = []
|
||||
return Effect.gen(function* () {
|
||||
const error = yield* Transcription.generate({ model, audio }).pipe(Effect.flip)
|
||||
|
||||
expect(error.reason).toMatchObject({ _tag: "InvalidProviderOutput", body: JSON.stringify({ words: [] }) })
|
||||
expect(error.reason.http?.status).toBe(200)
|
||||
}).pipe(Effect.provide(respondTranscript(uploads, { words: [] })))
|
||||
})
|
||||
})
|
||||
@@ -235,6 +235,12 @@ export type SessionForkInput = { readonly sessionID: Session.ID; readonly before
|
||||
export type SessionForkOutput = Session.Info
|
||||
export type SessionForkOperation<E = never> = (input: SessionForkInput) => Effect.Effect<SessionForkOutput, E>
|
||||
|
||||
export type SessionCompanionInput = { readonly sessionID: Session.ID }
|
||||
export type SessionCompanionOutput = Session.Info
|
||||
export type SessionCompanionOperation<E = never> = (
|
||||
input: SessionCompanionInput,
|
||||
) => Effect.Effect<SessionCompanionOutput, E>
|
||||
|
||||
export type SessionSwitchAgentInput = { readonly sessionID: Session.ID; readonly agent: Agent.ID }
|
||||
export type SessionSwitchAgentOutput = void
|
||||
export type SessionSwitchAgentOperation<E = never> = (
|
||||
@@ -445,6 +451,7 @@ export type SessionLogOutput =
|
||||
}
|
||||
readonly subpath?: RelativePath | undefined
|
||||
readonly parentID?: Session.ID | undefined
|
||||
readonly kind?: Session.Kind | undefined
|
||||
readonly slug: string
|
||||
readonly title?: string | undefined
|
||||
readonly agent?: Agent.ID | undefined
|
||||
@@ -1418,6 +1425,7 @@ export interface SessionApi<E = never> {
|
||||
readonly get: SessionGetOperation<E>
|
||||
readonly remove: SessionRemoveOperation<E>
|
||||
readonly fork: SessionForkOperation<E>
|
||||
readonly companion: SessionCompanionOperation<E>
|
||||
readonly switchAgent: SessionSwitchAgentOperation<E>
|
||||
readonly switchModel: SessionSwitchModelOperation<E>
|
||||
readonly update: SessionUpdateOperation<E>
|
||||
@@ -1513,6 +1521,30 @@ export interface GenerateApi<E = never> {
|
||||
readonly text: GenerateTextOperation<E>
|
||||
}
|
||||
|
||||
export type VoiceTranscribeInput = { readonly mediaType: string; readonly payload: globalThis.Uint8Array }
|
||||
export type VoiceTranscribeOutput = { readonly text: string }
|
||||
export type VoiceTranscribeOperation<E = never> = (
|
||||
input: VoiceTranscribeInput,
|
||||
) => Effect.Effect<VoiceTranscribeOutput, E>
|
||||
|
||||
export type VoiceSpeechInput = { readonly text: string }
|
||||
export type VoiceSpeechOutput =
|
||||
| {
|
||||
readonly type: "format"
|
||||
readonly format:
|
||||
| { readonly type: "mp3" }
|
||||
| { readonly type: "pcm"; readonly sampleRate: number; readonly channels: number }
|
||||
}
|
||||
| { readonly type: "audio"; readonly data: string }
|
||||
| { readonly type: "done" }
|
||||
| { readonly type: "error"; readonly message: string }
|
||||
export type VoiceSpeechOperation<E = never> = (input: VoiceSpeechInput) => Stream.Stream<VoiceSpeechOutput, E>
|
||||
|
||||
export interface VoiceApi<E = never> {
|
||||
readonly transcribe: VoiceTranscribeOperation<E>
|
||||
readonly speech: VoiceSpeechOperation<E>
|
||||
}
|
||||
|
||||
export type ProviderListInput = { readonly location?: { readonly directory?: string | undefined } | undefined }
|
||||
export type ProviderListOutput = { readonly location: Location.PublicRef; readonly data: ReadonlyArray<Provider.Info> }
|
||||
export type ProviderListOperation<E = never> = (input?: ProviderListInput) => Effect.Effect<ProviderListOutput, E>
|
||||
@@ -2342,6 +2374,7 @@ export interface AppApi<E = never> {
|
||||
readonly message: MessageApi<E>
|
||||
readonly model: ModelApi<E>
|
||||
readonly generate: GenerateApi<E>
|
||||
readonly voice: VoiceApi<E>
|
||||
readonly provider: ProviderApi<E>
|
||||
readonly integration: IntegrationApi<E>
|
||||
readonly mcp: McpApi<E>
|
||||
|
||||
@@ -39,6 +39,8 @@ import type {
|
||||
SessionRemoveOutput,
|
||||
SessionForkInput,
|
||||
SessionForkOutput,
|
||||
SessionCompanionInput,
|
||||
SessionCompanionOutput,
|
||||
SessionSwitchAgentInput,
|
||||
SessionSwitchAgentOutput,
|
||||
SessionSwitchModelInput,
|
||||
@@ -115,6 +117,10 @@ import type {
|
||||
ModelDefaultOutput,
|
||||
GenerateTextInput,
|
||||
GenerateTextOutput,
|
||||
VoiceTranscribeInput,
|
||||
VoiceTranscribeOutput,
|
||||
VoiceSpeechInput,
|
||||
VoiceSpeechOutput,
|
||||
ProviderListInput,
|
||||
ProviderListOutput,
|
||||
ProviderGetInput,
|
||||
@@ -450,6 +456,14 @@ const EndpointSessionFork = (raw: RawClient["server.session"]) => (input: Sessio
|
||||
),
|
||||
)
|
||||
|
||||
const EndpointSessionCompanion = (raw: RawClient["server.session"]) => (input: SessionCompanionInput) =>
|
||||
preserveEffect<SessionCompanionOutput>()(
|
||||
raw["session.companion"]({ params: { sessionID: input["sessionID"] } }).pipe(
|
||||
Effect.mapError(mapClientError),
|
||||
Effect.map((value) => value.data),
|
||||
),
|
||||
)
|
||||
|
||||
const EndpointSessionSwitchAgent = (raw: RawClient["server.session"]) => (input: SessionSwitchAgentInput) =>
|
||||
preserveEffect<SessionSwitchAgentOutput>()(
|
||||
raw["session.switchAgent"]({ params: { sessionID: input["sessionID"] }, payload: { agent: input["agent"] } }).pipe(
|
||||
@@ -762,6 +776,7 @@ const adaptGroupSession = (raw: RawClient["server.session"]) => ({
|
||||
get: EndpointSessionGet(raw),
|
||||
remove: EndpointSessionRemove(raw),
|
||||
fork: EndpointSessionFork(raw),
|
||||
companion: EndpointSessionCompanion(raw),
|
||||
switchAgent: EndpointSessionSwitchAgent(raw),
|
||||
switchModel: EndpointSessionSwitchModel(raw),
|
||||
update: EndpointSessionUpdate(raw),
|
||||
@@ -843,6 +858,33 @@ const EndpointGenerateText = (raw: RawClient["server.generate"]) => (input: Gene
|
||||
|
||||
const adaptGroupGenerate = (raw: RawClient["server.generate"]) => ({ text: EndpointGenerateText(raw) })
|
||||
|
||||
type VoiceTranscribeRequest = Parameters<RawClient["server.voice"]["voice.transcribe"]>[0]
|
||||
const EndpointVoiceTranscribe = (raw: RawClient["server.voice"]) => (input: VoiceTranscribeInput) =>
|
||||
preserveEffect<VoiceTranscribeOutput>()(
|
||||
raw["voice.transcribe"]({
|
||||
query: { mediaType: input["mediaType"] },
|
||||
payload: input["payload"],
|
||||
} as VoiceTranscribeRequest).pipe(
|
||||
Effect.mapError(mapClientError),
|
||||
Effect.map((value) => value.data),
|
||||
),
|
||||
)
|
||||
|
||||
const EndpointVoiceSpeech = (raw: RawClient["server.voice"]) => (input: VoiceSpeechInput) =>
|
||||
preserveStream<VoiceSpeechOutput>()(
|
||||
Stream.unwrap(
|
||||
raw["voice.speech"]({ payload: { text: input["text"] } }).pipe(
|
||||
Effect.mapError(mapClientError),
|
||||
Effect.map((stream) => stream.pipe(Stream.mapError(mapClientError))),
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
const adaptGroupVoice = (raw: RawClient["server.voice"]) => ({
|
||||
transcribe: EndpointVoiceTranscribe(raw),
|
||||
speech: EndpointVoiceSpeech(raw),
|
||||
})
|
||||
|
||||
const EndpointProviderList = (raw: RawClient["server.provider"]) => (input?: ProviderListInput) =>
|
||||
preserveEffect<ProviderListOutput>()(
|
||||
raw["provider.list"]({ query: { location: input?.["location"] } }).pipe(Effect.mapError(mapClientError)),
|
||||
@@ -1550,6 +1592,7 @@ const adaptClient = (raw: RawClient) => ({
|
||||
message: adaptGroupMessage(raw["server.message"]),
|
||||
model: adaptGroupModel(raw["server.model"]),
|
||||
generate: adaptGroupGenerate(raw["server.generate"]),
|
||||
voice: adaptGroupVoice(raw["server.voice"]),
|
||||
provider: adaptGroupProvider(raw["server.provider"]),
|
||||
integration: adaptGroupIntegration(raw["server.integration"]),
|
||||
mcp: adaptGroupMcp(raw["server.mcp"]),
|
||||
|
||||
@@ -33,6 +33,8 @@ import type {
|
||||
SessionRemoveOutput,
|
||||
SessionForkInput,
|
||||
SessionForkOutput,
|
||||
SessionCompanionInput,
|
||||
SessionCompanionOutput,
|
||||
SessionSwitchAgentInput,
|
||||
SessionSwitchAgentOutput,
|
||||
SessionSwitchModelInput,
|
||||
@@ -109,6 +111,10 @@ import type {
|
||||
ModelDefaultOutput,
|
||||
GenerateTextInput,
|
||||
GenerateTextOutput,
|
||||
VoiceTranscribeInput,
|
||||
VoiceTranscribeOutput,
|
||||
VoiceSpeechInput,
|
||||
VoiceSpeechOutput,
|
||||
ProviderListInput,
|
||||
ProviderListOutput,
|
||||
ProviderGetInput,
|
||||
@@ -656,6 +662,17 @@ export function make(options: ClientOptions) {
|
||||
},
|
||||
requestOptions,
|
||||
).then((value) => value.data),
|
||||
companion: (input: SessionCompanionInput, requestOptions?: RequestOptions) =>
|
||||
request<{ readonly data: SessionCompanionOutput }>(
|
||||
{
|
||||
method: "POST",
|
||||
path: `/api/experimental/session/${encodeURIComponent(input.sessionID)}/companion`,
|
||||
successStatus: 200,
|
||||
declaredStatuses: [400, 401, 404],
|
||||
empty: false,
|
||||
},
|
||||
requestOptions,
|
||||
).then((value) => value.data),
|
||||
switchAgent: (input: SessionSwitchAgentInput, requestOptions?: RequestOptions) =>
|
||||
request<SessionSwitchAgentOutput>(
|
||||
{
|
||||
@@ -1141,6 +1158,34 @@ export function make(options: ClientOptions) {
|
||||
requestOptions,
|
||||
).then((value) => value.data),
|
||||
},
|
||||
voice: {
|
||||
transcribe: (input: VoiceTranscribeInput, requestOptions?: RequestOptions) =>
|
||||
request<{ readonly data: VoiceTranscribeOutput }>(
|
||||
{
|
||||
method: "POST",
|
||||
path: `/api/experimental/voice/transcribe`,
|
||||
query: { mediaType: input["mediaType"] },
|
||||
body: input["payload"],
|
||||
successStatus: 200,
|
||||
declaredStatuses: [400, 401, 503],
|
||||
empty: false,
|
||||
binaryBody: true,
|
||||
},
|
||||
requestOptions,
|
||||
).then((value) => value.data),
|
||||
speech: (input: VoiceSpeechInput, requestOptions?: RequestOptions): AsyncIterable<VoiceSpeechOutput> =>
|
||||
sse<VoiceSpeechOutput>(
|
||||
{
|
||||
method: "POST",
|
||||
path: `/api/experimental/voice/speech`,
|
||||
body: { text: input["text"] },
|
||||
successStatus: 200,
|
||||
declaredStatuses: [400, 401, 503],
|
||||
empty: false,
|
||||
},
|
||||
requestOptions,
|
||||
),
|
||||
},
|
||||
provider: {
|
||||
list: (input?: ProviderListInput, requestOptions?: RequestOptions) =>
|
||||
request<ProviderListOutput>(
|
||||
|
||||
@@ -30,6 +30,8 @@ export type PluginFeatures = { server?: true; tui?: true; rpc?: true }
|
||||
|
||||
export type PluginState = { status: "active" } | { status: "failed"; error: string; ref?: string }
|
||||
|
||||
export type SessionKind = "companion"
|
||||
|
||||
export type SessionForkBoundary = { type: "before"; messageID: string } | { type: "through"; messageID: string }
|
||||
|
||||
export type MoneyUSD = number
|
||||
@@ -231,6 +233,10 @@ export type MoneyUSDPerMillionTokens = number
|
||||
|
||||
export type GenerateTextResponse = { data: { text: string } }
|
||||
|
||||
export type VoiceTranscribeResponse = { data: { text: string } }
|
||||
|
||||
export type VoiceSpeechFormat = { type: "mp3" } | { type: "pcm"; sampleRate: number; channels: number }
|
||||
|
||||
export type IntegrationCommandMethod = { id: string; type: "command"; label: string; command: Array<string> }
|
||||
|
||||
export type IntegrationEnvMethod = { type: "env"; names: Array<string> }
|
||||
@@ -1469,6 +1475,12 @@ export type ModelCost = {
|
||||
cache: { read: MoneyUSDPerMillionTokens; write: MoneyUSDPerMillionTokens }
|
||||
}
|
||||
|
||||
export type VoiceSpeechEvent =
|
||||
| { type: "format"; format: VoiceSpeechFormat }
|
||||
| { type: "audio"; data: string }
|
||||
| { type: "done" }
|
||||
| { type: "error"; message: string }
|
||||
|
||||
export type ConnectionInfo = ConnectionCredentialInfo | ConnectionEnvInfo
|
||||
|
||||
export type McpServer = {
|
||||
@@ -1957,6 +1969,7 @@ export type SessionPermissions = {
|
||||
export type SessionInfo = {
|
||||
id: string
|
||||
parentID?: string
|
||||
kind?: SessionKind
|
||||
fork?: { sessionID: string; boundary: SessionForkBoundary }
|
||||
projectID: string
|
||||
agent?: string
|
||||
@@ -1986,6 +1999,7 @@ export type SessionCreated = {
|
||||
location: LocationRef
|
||||
subpath?: string
|
||||
parentID?: string
|
||||
kind?: SessionKind
|
||||
slug: string
|
||||
title?: string
|
||||
agent?: string
|
||||
@@ -2052,6 +2066,16 @@ export type ConfigEntry =
|
||||
media?: {
|
||||
image?: { auto_resize?: boolean; max_width?: number; max_height?: number; max_base64_bytes?: number }
|
||||
}
|
||||
voice?: {
|
||||
transcription?: { model: string | { providerID: string; model: string; variant?: string }; language?: string }
|
||||
speech?: {
|
||||
model: string | { providerID: string; model: string; variant?: string }
|
||||
voice?: string
|
||||
language?: string
|
||||
speed?: number
|
||||
instructions?: string
|
||||
}
|
||||
}
|
||||
tool_output?: { max_lines?: number; max_bytes?: number }
|
||||
mcp?: {
|
||||
timeout?: { startup?: number; catalog?: number; execution?: number }
|
||||
@@ -2975,6 +2999,7 @@ export type SessionImportInput = {
|
||||
readonly info: {
|
||||
readonly id: string
|
||||
readonly parentID?: string
|
||||
readonly kind?: "companion"
|
||||
readonly fork?: {
|
||||
readonly sessionID: string
|
||||
readonly boundary:
|
||||
@@ -3292,6 +3317,7 @@ export type SessionImportInput = {
|
||||
readonly info: {
|
||||
readonly id: string
|
||||
readonly parentID?: string
|
||||
readonly kind?: "companion"
|
||||
readonly fork?: {
|
||||
readonly sessionID: string
|
||||
readonly boundary:
|
||||
@@ -3609,6 +3635,7 @@ export type SessionImportInput = {
|
||||
readonly info: {
|
||||
readonly id: string
|
||||
readonly parentID?: string
|
||||
readonly kind?: "companion"
|
||||
readonly fork?: {
|
||||
readonly sessionID: string
|
||||
readonly boundary:
|
||||
@@ -3950,6 +3977,10 @@ export type SessionForkInput = {
|
||||
|
||||
export type SessionForkOutput = { data: SessionInfo }["data"]
|
||||
|
||||
export type SessionCompanionInput = { readonly sessionID: { readonly sessionID: string }["sessionID"] }
|
||||
|
||||
export type SessionCompanionOutput = { data: SessionInfo }["data"]
|
||||
|
||||
export type SessionSwitchAgentInput = {
|
||||
readonly sessionID: { readonly sessionID: string }["sessionID"]
|
||||
readonly agent: { readonly agent: string }["agent"]
|
||||
@@ -5483,6 +5514,17 @@ export type GenerateTextInput = {
|
||||
|
||||
export type GenerateTextOutput = GenerateTextResponse["data"]
|
||||
|
||||
export type VoiceTranscribeInput = {
|
||||
readonly mediaType: { readonly mediaType: string }["mediaType"]
|
||||
readonly payload: globalThis.Uint8Array
|
||||
}
|
||||
|
||||
export type VoiceTranscribeOutput = VoiceTranscribeResponse["data"]
|
||||
|
||||
export type VoiceSpeechInput = { readonly text: { readonly text: string }["text"] }
|
||||
|
||||
export type VoiceSpeechOutput = VoiceSpeechEvent
|
||||
|
||||
export type ProviderListInput = {
|
||||
readonly location?: { readonly location?: { readonly directory?: string | undefined } | undefined }["location"]
|
||||
}
|
||||
|
||||
@@ -13,6 +13,7 @@ test("exposes every standard HTTP API group", () => {
|
||||
"message",
|
||||
"model",
|
||||
"generate",
|
||||
"voice",
|
||||
"provider",
|
||||
"integration",
|
||||
"mcp",
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
{
|
||||
"version": "7",
|
||||
"dialect": "sqlite",
|
||||
"id": "e98619ef-c334-45e3-8c48-9932ed1ef290",
|
||||
"prevIds": ["be60f352-8da1-40e1-8d70-dc41121cfbc5"],
|
||||
"id": "421b5a69-27cf-4f13-ba81-85503aa81323",
|
||||
"prevIds": ["e98619ef-c334-45e3-8c48-9932ed1ef290"],
|
||||
"ddl": [
|
||||
{
|
||||
"name": "account_state",
|
||||
@@ -1110,6 +1110,16 @@
|
||||
"entityType": "columns",
|
||||
"table": "session_v2"
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"notNull": false,
|
||||
"autoincrement": false,
|
||||
"default": null,
|
||||
"generated": null,
|
||||
"name": "kind",
|
||||
"entityType": "columns",
|
||||
"table": "session_v2"
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"notNull": false,
|
||||
|
||||
@@ -207,6 +207,7 @@ export function normalize(input: unknown): Result {
|
||||
username: Info.fields.username,
|
||||
snapshots: Info.fields.snapshots,
|
||||
media: Info.fields.media,
|
||||
voice: Info.fields.voice,
|
||||
tool_output: Info.fields.tool_output,
|
||||
websearch: Info.fields.websearch,
|
||||
worktree: Info.fields.worktree,
|
||||
|
||||
+2
@@ -47,6 +47,7 @@ import m44 from "./migration/20260819222447_session_viewed_state.js"
|
||||
import m45 from "./migration/20260823191254_nullable_workspace_binding.js"
|
||||
import m46 from "./migration/20260910120000_clear_v1_session_permission.js"
|
||||
import m47 from "./migration/20260923013825_project_time_active.js"
|
||||
import m48 from "./migration/20260925102936_session_kind.js"
|
||||
|
||||
export const migrations = [
|
||||
m00,
|
||||
@@ -97,4 +98,5 @@ export const migrations = [
|
||||
m45,
|
||||
m46,
|
||||
m47,
|
||||
m48,
|
||||
] satisfies DatabaseMigration.Migration[]
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
import { Effect } from "effect"
|
||||
import type { DatabaseMigration } from "../migration.js"
|
||||
|
||||
const migration: DatabaseMigration.Migration = {
|
||||
id: "20260925102936_session_kind",
|
||||
up(tx) {
|
||||
return Effect.gen(function* () {
|
||||
yield* tx.run(`ALTER TABLE \`session_v2\` ADD \`kind\` text;`)
|
||||
})
|
||||
},
|
||||
}
|
||||
|
||||
export default migration
|
||||
@@ -185,6 +185,7 @@ const schema: Omit<DatabaseMigration.Migration, "id"> = {
|
||||
\`project_id\` text NOT NULL,
|
||||
\`workspace_id\` text,
|
||||
\`parent_id\` text,
|
||||
\`kind\` text,
|
||||
\`fork_session_id\` text,
|
||||
\`fork_boundary\` text,
|
||||
\`slug\` text NOT NULL,
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
import { LLMClient, RequestExecutor } from "@opencode/ai/route"
|
||||
import { SpeechClient } from "@opencode/ai/speech-client"
|
||||
import { TranscriptionClient } from "@opencode/ai/transcription-client"
|
||||
import { Socket } from "effect/unstable/socket"
|
||||
import { makeGlobalNode } from "@opencode/util/effect/app-node"
|
||||
import { httpClient } from "@opencode/util/effect/app-node-platform"
|
||||
@@ -12,6 +14,18 @@ export const requestExecutor = makeGlobalNode({
|
||||
|
||||
export const llmClient = makeGlobalNode({ service: LLMClient.Service, layer: LLMClient.layer, deps: [requestExecutor] })
|
||||
|
||||
export const speechClient = makeGlobalNode({
|
||||
service: SpeechClient.Service,
|
||||
layer: SpeechClient.layer,
|
||||
deps: [requestExecutor],
|
||||
})
|
||||
|
||||
export const transcriptionClient = makeGlobalNode({
|
||||
service: TranscriptionClient.Service,
|
||||
layer: TranscriptionClient.layer,
|
||||
deps: [requestExecutor],
|
||||
})
|
||||
|
||||
export const webSocketConstructor = makeGlobalNode({
|
||||
service: Socket.WebSocketConstructor,
|
||||
layer: WebSocketConstructor.layer,
|
||||
|
||||
@@ -13,6 +13,7 @@ import { Formatter } from "./formatter.js"
|
||||
import { FileSystem } from "./filesystem.js"
|
||||
import { FileSystemSearch } from "./filesystem/search.js"
|
||||
import { Generate } from "./generate.js"
|
||||
import { Voice } from "./voice.js"
|
||||
import { Form } from "./form.js"
|
||||
import { Image } from "./image.js"
|
||||
import { LocationWatcher } from "./filesystem/location-watcher.js"
|
||||
@@ -97,6 +98,7 @@ const nodes = [
|
||||
InstructionEntry.node,
|
||||
Form.node,
|
||||
Generate.node,
|
||||
Voice.node,
|
||||
ReadToolFileSystem.node,
|
||||
McpTool.node,
|
||||
SessionInstructions.node,
|
||||
|
||||
@@ -23,6 +23,20 @@ Guidelines:
|
||||
|
||||
Complete the user's search request efficiently and report your findings clearly.`
|
||||
|
||||
const PROMPT_COMPANION = `You are the companion of an OpenCode coding session, called the main session. The user talks with you, in text or by voice, about what the main session is doing while it keeps working.
|
||||
|
||||
Your job:
|
||||
- Answer questions about the main session: what it is doing, why, what it changed, and what it waits for.
|
||||
- Steer it when the user asks. Use main_send to pass instructions, corrections, or follow-up work. Write each prompt as a clear, self-contained instruction for the main session's agent, not as a transcript of the user's words.
|
||||
- Use main_interrupt only when the user wants the main session to stop.
|
||||
- Check things yourself with read, grep, and glob when that is faster than asking.
|
||||
- You may run read-only git commands with the shell tool: git status, git diff, git log, git show, and plain git branch. Run each one alone, without pipes, redirects, or --output.
|
||||
- You cannot edit files or run other commands. Ask the main session to do that with main_send.
|
||||
|
||||
Each user turn may start with a <main-session> snapshot. Trust it for the current status. Call main_read or main_status when you need more detail.
|
||||
|
||||
Your replies may be read aloud. Keep them short and conversational: a few sentences, no headings or tables, and code only when the user asks for it. When you steer the main session, say briefly what you sent.`
|
||||
|
||||
const PROMPT_TITLE = `You are a title generator. You output ONLY a thread title. Nothing else.
|
||||
|
||||
<task>
|
||||
@@ -130,6 +144,45 @@ export const Plugin = define({
|
||||
)
|
||||
})
|
||||
|
||||
editor.update(Agent.ID.make("companion"), (item) => {
|
||||
const externalDirectories = item.permissions.filter(
|
||||
(rule) => rule.action === "external_directory" && rule.effect === "allow",
|
||||
)
|
||||
item.name = Agent.Name.make("Companion")
|
||||
item.description = "Talks with the user about a main session and steers it."
|
||||
item.system = PROMPT_COMPANION
|
||||
item.mode = "primary"
|
||||
item.hidden = true
|
||||
// Clients show no permission prompts for companions, so every rule allows or denies.
|
||||
item.permissions.push(
|
||||
{ action: "*", resource: "*", effect: "deny" },
|
||||
{ action: "grep", resource: "*", effect: "allow" },
|
||||
{ action: "glob", resource: "*", effect: "allow" },
|
||||
{ action: "webfetch", resource: "*", effect: "allow" },
|
||||
{ action: "websearch", resource: "*", effect: "allow" },
|
||||
{ action: "read", resource: "*", effect: "allow" },
|
||||
{ action: "read", resource: "*.env", effect: "deny" },
|
||||
{ action: "read", resource: "*.env.*", effect: "deny" },
|
||||
{ action: "read", resource: "*.env.example", effect: "allow" },
|
||||
{ action: "main_*", resource: "*", effect: "allow" },
|
||||
...["status", "diff", "log", "show"].map((command) => ({
|
||||
action: "shell",
|
||||
resource: `git ${command} *`,
|
||||
effect: "allow" as const,
|
||||
})),
|
||||
// Exact forms only: `git branch <name>` creates a branch, even after -v.
|
||||
...["", " --show-current", " -a", " -r", " -v", " -vv"].map((flags) => ({
|
||||
action: "shell",
|
||||
resource: `git branch${flags}`,
|
||||
effect: "allow" as const,
|
||||
})),
|
||||
// Redirects and --output write files.
|
||||
{ action: "shell", resource: "*>*", effect: "deny" },
|
||||
{ action: "shell", resource: "*--output*", effect: "deny" },
|
||||
...externalDirectories,
|
||||
)
|
||||
})
|
||||
|
||||
editor.update(Agent.ID.make("compaction"), (item) => {
|
||||
item.name = Agent.Name.make("Compaction")
|
||||
item.mode = "primary"
|
||||
|
||||
@@ -79,6 +79,7 @@ import { ReadTool } from "../tool/plugin/read.js"
|
||||
import { ShellTool } from "../tool/plugin/shell.js"
|
||||
import { SkillTool } from "../tool/plugin/skill.js"
|
||||
import { SubagentTool } from "../tool/plugin/subagent.js"
|
||||
import { CompanionTools } from "../tool/plugin/companion.js"
|
||||
import { Tool } from "../tool.js"
|
||||
import { ToolOutput } from "../tool-output.js"
|
||||
import { WebFetchTool } from "../tool/plugin/webfetch.js"
|
||||
@@ -243,6 +244,7 @@ const pre = [
|
||||
ShellTool.Plugin,
|
||||
SkillTool.Plugin,
|
||||
SubagentTool.Plugin,
|
||||
CompanionTools.Plugin,
|
||||
WebFetchTool.Plugin,
|
||||
WebSearchTool.Plugin,
|
||||
WriteTool.Plugin,
|
||||
|
||||
@@ -17,7 +17,7 @@ import { SessionProjector } from "./session/projector.js"
|
||||
import { SessionMessageTable } from "./session/sql.js"
|
||||
import { SessionSchema } from "./session/schema.js"
|
||||
import { RelativePath } from "./schema.js"
|
||||
import { Agent } from "@opencode/schema/agent"
|
||||
import { Agent } from "./agent.js"
|
||||
import type { Permission } from "@opencode/schema/permission"
|
||||
import { App } from "./app.js"
|
||||
import { Slug } from "./util/slug.js"
|
||||
@@ -52,6 +52,7 @@ import {
|
||||
DestinationUnavailableError,
|
||||
} from "./session/move.js"
|
||||
import { SessionModelTransport } from "./session/model-transport.js"
|
||||
import { Plugin } from "./plugin/service.js"
|
||||
import { llmClient } from "./effect/app-node-platform.js"
|
||||
import { Snapshot } from "./snapshot.js"
|
||||
import { Session } from "./session/session.js"
|
||||
@@ -81,6 +82,7 @@ export type ListInput = SessionStore.ListInput
|
||||
|
||||
type CreateBaseInput = {
|
||||
id?: SessionSchema.ID
|
||||
kind?: SessionSchema.Kind
|
||||
title?: string
|
||||
agent?: Agent.ID
|
||||
model?: Model.Ref
|
||||
@@ -119,6 +121,8 @@ export interface Interface {
|
||||
readonly data: SessionSchema.Info[]
|
||||
}>
|
||||
readonly create: (input: CreateInput) => Effect.Effect<SessionSchema.Info, NotFoundError>
|
||||
/** The main Session's companion, created on first use. A companion is its own companion. */
|
||||
readonly companion: (sessionID: SessionSchema.ID) => Effect.Effect<SessionSchema.Info, NotFoundError>
|
||||
readonly fork: (
|
||||
input: ForkInput,
|
||||
) => Effect.Effect<SessionSchema.Info, NotFoundError | MessageNotFoundError | ForkEmptyError>
|
||||
@@ -266,6 +270,7 @@ const layer = Layer.effect(
|
||||
version: app.version,
|
||||
projectID: project.id,
|
||||
parentID: input.parentID,
|
||||
kind: input.kind,
|
||||
location,
|
||||
subpath: RelativePath.make(path.relative(project.directory, location.directory).replaceAll("\\", "/")),
|
||||
title: input.title,
|
||||
@@ -304,6 +309,25 @@ const layer = Layer.effect(
|
||||
// TODO: Restore recorded sessions onto replacement synchronized workspaces in a future API slice.
|
||||
return yield* result.get(sessionID).pipe(Effect.orDie)
|
||||
}),
|
||||
companion: Effect.fn("Session.companion")(function* (sessionID) {
|
||||
const main = yield* result.get(sessionID)
|
||||
if (main.kind === "companion") return main
|
||||
const existing = (yield* store.list({ parentID: sessionID })).find((child) => child.kind === "companion")
|
||||
if (existing) return existing
|
||||
const agent = yield* Plugin.awaitActivation.pipe(
|
||||
Effect.andThen(Agent.Service.use((agents) => agents.get(Agent.ID.make("companion")))),
|
||||
instances.provide(main),
|
||||
)
|
||||
return yield* result.create({
|
||||
parentID: sessionID,
|
||||
kind: "companion",
|
||||
title: "Companion",
|
||||
agent: Agent.ID.make("companion"),
|
||||
model: agent?.model ?? main.model,
|
||||
// Main-session approvals must not widen what the companion may do.
|
||||
permissions: [],
|
||||
})
|
||||
}),
|
||||
fork: Effect.fn("Session.fork")(function* (input) {
|
||||
const parent = yield* result.get(input.sessionID)
|
||||
const boundary = yield* db
|
||||
|
||||
@@ -19,6 +19,7 @@ export function fromRow(row: typeof SessionTable.$inferSelect): SessionSchema.In
|
||||
projectID: Project.ID.make(row.project_id),
|
||||
title: row.title ?? undefined,
|
||||
parentID: row.parent_id ? SessionSchema.ID.make(row.parent_id) : undefined,
|
||||
kind: row.kind ?? undefined,
|
||||
fork:
|
||||
row.fork_session_id && row.fork_boundary
|
||||
? {
|
||||
|
||||
@@ -444,6 +444,7 @@ const layer = Layer.effectDiscard(
|
||||
project_id: event.data.projectID,
|
||||
workspace_id: event.data.location.workspaceID ? Workspace.ID.make(event.data.location.workspaceID) : null,
|
||||
parent_id: event.data.parentID,
|
||||
kind: event.data.kind,
|
||||
slug: event.data.slug,
|
||||
directory: event.data.location.directory,
|
||||
path: event.data.subpath,
|
||||
|
||||
@@ -29,6 +29,7 @@ export const SessionTable = sqliteTable(
|
||||
.references(() => ProjectTable.id, { onDelete: "cascade" }),
|
||||
workspace_id: text().$type<Workspace.ID>(),
|
||||
parent_id: text().$type<SessionSchema.ID>(),
|
||||
kind: text().$type<Session.Kind>(),
|
||||
fork_session_id: text().$type<SessionSchema.ID>(),
|
||||
fork_boundary: text({ mode: "json" }).$type<Session.ForkBoundary>(),
|
||||
slug: text().notNull(),
|
||||
|
||||
@@ -92,6 +92,7 @@ const layer = Layer.effect(
|
||||
{
|
||||
sessionID,
|
||||
parentID: input.data.info.parentID,
|
||||
kind: input.data.info.kind,
|
||||
slug: Slug.create(),
|
||||
version: app.version,
|
||||
projectID: project.id,
|
||||
|
||||
@@ -0,0 +1,287 @@
|
||||
export * as CompanionTools from "./companion.js"
|
||||
|
||||
import { Message, ToolFailure } from "@opencode/ai"
|
||||
import type { Context } from "@opencode/plugin/effect/plugin"
|
||||
import type { SessionHooks } from "@opencode/plugin/effect/session"
|
||||
import { Agent } from "@opencode/schema/agent"
|
||||
import { SessionInbox } from "@opencode/schema/session-inbox"
|
||||
import { SessionMessage } from "@opencode/schema/session-message"
|
||||
import { Tool } from "@opencode/schema/tool"
|
||||
import { Effect, Predicate, Schema } from "effect"
|
||||
import { Permission } from "../../permission.js"
|
||||
import { Session } from "../../session.js"
|
||||
import { SessionSchema } from "../../session/schema.js"
|
||||
import { ShellTool } from "./shell.js"
|
||||
|
||||
export const agent = Agent.ID.make("companion")
|
||||
|
||||
const names = ["main_status", "main_read", "main_send", "main_cancel", "main_interrupt"]
|
||||
|
||||
// User config rules apply after the companion's own, but never beyond these actions.
|
||||
const actions = new Set(["read", "grep", "glob", "webfetch", "websearch", "shell", "external_directory", ...names])
|
||||
|
||||
// One read-only git command with no shell operators. The shell parser also drops redirects that
|
||||
// follow `&&` or `||` from permission resources, so the companion's rules alone cannot enforce this.
|
||||
const readOnlyGit = /^git (?:(?:status|diff|log|show)(?: [^;&|<>$`\n]*)?|branch(?: --show-current| -a| -r| -vv?)?)$/
|
||||
|
||||
const ReadInput = Schema.Struct({
|
||||
limit: Schema.optionalKey(Schema.Int.check(Schema.isBetween({ minimum: 1, maximum: 50 }))).annotate({
|
||||
description: "Number of most recent messages to return. Defaults to 12.",
|
||||
}),
|
||||
})
|
||||
|
||||
const SendInput = Schema.Struct({
|
||||
text: Schema.String.check(Schema.isMinLength(1)).annotate({
|
||||
description: "A clear, self-contained instruction for the main session's agent.",
|
||||
}),
|
||||
delivery: Schema.optionalKey(SessionInbox.Delivery).annotate({
|
||||
description:
|
||||
"steer (default) delivers at the main session's next safe boundary, or starts it when idle. queue waits until the main session finishes its current work.",
|
||||
}),
|
||||
})
|
||||
|
||||
const CancelInput = Schema.Struct({
|
||||
inboxID: SessionMessage.ID.annotate({ description: "ID of a pending inbox item from main_status." }),
|
||||
})
|
||||
|
||||
const InterruptInput = Schema.Struct({
|
||||
resume: Schema.optionalKey(Schema.Boolean).annotate({
|
||||
description: "Resume pending steering prompts after the interrupt. Defaults to false.",
|
||||
}),
|
||||
})
|
||||
|
||||
/** Tools that let a companion observe and steer its main Session. Only companion Sessions may call them. */
|
||||
export const Plugin = {
|
||||
id: "opencode.tool.companion",
|
||||
effect: Effect.fn("CompanionTools.Plugin")(function* (ctx: Context) {
|
||||
const sessions = yield* Session.Service
|
||||
const permission = yield* Permission.Service
|
||||
|
||||
const mainOf = Effect.fn("CompanionTools.mainOf")(function* (sessionID: SessionSchema.ID) {
|
||||
const self = yield* sessions
|
||||
.get(sessionID)
|
||||
.pipe(Effect.mapError((error) => new ToolFailure({ message: `Session not found: ${sessionID}`, error })))
|
||||
if (self.kind !== "companion" || !self.parentID)
|
||||
return yield* new ToolFailure({ message: "Only a companion session can use this tool" })
|
||||
return yield* sessions
|
||||
.get(self.parentID)
|
||||
.pipe(
|
||||
Effect.mapError((error) => new ToolFailure({ message: `Main session not found: ${self.parentID}`, error })),
|
||||
)
|
||||
})
|
||||
|
||||
const status = Effect.fn("CompanionTools.status")(function* (main: SessionSchema.Info) {
|
||||
const [active, inbox, permissions] = yield* Effect.all(
|
||||
[sessions.active, sessions.inbox(main.id).pipe(Effect.orElseSucceed(() => [])), permission.forSession(main.id)],
|
||||
{ concurrency: "unbounded" },
|
||||
)
|
||||
const pending = inbox.flatMap((item) =>
|
||||
item.type === "user" || item.type === "synthetic"
|
||||
? [`- ${item.id} [${item.delivery}] ${clip(item.payload.text, 200)}`]
|
||||
: [],
|
||||
)
|
||||
const asked = permissions.map(
|
||||
(request) => `- ${request.action}: ${clip(request.message ?? request.resources.join(", "), 200)}`,
|
||||
)
|
||||
return [
|
||||
`Main session: ${main.title ?? "Untitled"} (${main.id})`,
|
||||
`Status: ${active.has(main.id) ? "running" : "idle"}${main.outcome ? ` (last run ${main.outcome})` : ""}`,
|
||||
`Agent: ${main.agent ?? "default"}${main.model ? ` · Model: ${main.model.providerID}/${main.model.id}` : ""}`,
|
||||
pending.length > 0 ? ["Pending inbox:", ...pending].join("\n") : "Pending inbox: none",
|
||||
asked.length > 0 ? ["Waiting for user permission:", ...asked].join("\n") : undefined,
|
||||
]
|
||||
.filter((line) => line !== undefined)
|
||||
.join("\n")
|
||||
})
|
||||
|
||||
const recent = Effect.fn("CompanionTools.recent")(function* (main: SessionSchema.Info, limit: number) {
|
||||
const messages = yield* sessions
|
||||
.messages({ sessionID: main.id, order: "desc", limit })
|
||||
.pipe(Effect.mapError((error) => new ToolFailure({ message: "Unable to read the main session", error })))
|
||||
return messages.toReversed().flatMap(describe)
|
||||
})
|
||||
|
||||
yield* ctx.tool
|
||||
.transform((draft) => {
|
||||
draft.namespace({
|
||||
name: "main",
|
||||
description: "Observe and steer the main session this companion is attached to.",
|
||||
})
|
||||
draft.add({
|
||||
name: "status",
|
||||
description:
|
||||
"Report whether the main session is running, which prompts are waiting in its inbox, and whether it waits for a permission decision.",
|
||||
input: Schema.Struct({}),
|
||||
options: { namespace: "main", codemode: false },
|
||||
execute: (_input, context) =>
|
||||
mainOf(context.sessionID).pipe(
|
||||
Effect.flatMap(status),
|
||||
Effect.map((content) => ({ content })),
|
||||
),
|
||||
})
|
||||
draft.add({
|
||||
name: "read",
|
||||
description:
|
||||
"Read the main session's most recent messages as compact text, including the step that is still running.",
|
||||
input: ReadInput,
|
||||
options: { namespace: "main", codemode: false },
|
||||
execute: (input, context) =>
|
||||
Effect.gen(function* () {
|
||||
const main = yield* mainOf(context.sessionID)
|
||||
const lines = yield* recent(main, input.limit ?? 12)
|
||||
return { content: lines.length > 0 ? lines.join("\n\n") : "The main session has no messages yet." }
|
||||
}),
|
||||
})
|
||||
draft.add({
|
||||
name: "send",
|
||||
description:
|
||||
"Send a prompt to the main session's agent. Use it to steer, correct, or give the main session follow-up work. The user sees the prompt in the main session.",
|
||||
input: SendInput,
|
||||
options: { namespace: "main", codemode: false },
|
||||
execute: (input, context) =>
|
||||
Effect.gen(function* () {
|
||||
const main = yield* mainOf(context.sessionID)
|
||||
const delivery = input.delivery ?? "steer"
|
||||
const item = yield* sessions
|
||||
.prompt({
|
||||
sessionID: main.id,
|
||||
text: input.text,
|
||||
delivery,
|
||||
metadata: { source: "companion", companionID: context.sessionID },
|
||||
})
|
||||
.pipe(
|
||||
Effect.mapError((error) => new ToolFailure({ message: "Unable to prompt the main session", error })),
|
||||
)
|
||||
return {
|
||||
content: `Sent to the main session as ${delivery} (inbox item ${item.id}).`,
|
||||
metadata: { inboxID: item.id, delivery },
|
||||
}
|
||||
}),
|
||||
})
|
||||
draft.add({
|
||||
name: "cancel",
|
||||
description: "Cancel a prompt that is still waiting in the main session's inbox.",
|
||||
input: CancelInput,
|
||||
options: { namespace: "main", codemode: false },
|
||||
execute: (input, context) =>
|
||||
Effect.gen(function* () {
|
||||
const main = yield* mainOf(context.sessionID)
|
||||
yield* sessions
|
||||
.cancelInbox({ sessionID: main.id, inboxID: input.inboxID })
|
||||
.pipe(
|
||||
Effect.mapError(
|
||||
(error) => new ToolFailure({ message: `Unable to cancel inbox item ${input.inboxID}`, error }),
|
||||
),
|
||||
)
|
||||
return { content: `Cancelled inbox item ${input.inboxID}.` }
|
||||
}),
|
||||
})
|
||||
draft.add({
|
||||
name: "interrupt",
|
||||
description:
|
||||
"Stop the main session's current work. Use it only when the user wants the main session to stop.",
|
||||
input: InterruptInput,
|
||||
options: { namespace: "main", codemode: false },
|
||||
execute: (input, context) =>
|
||||
Effect.gen(function* () {
|
||||
const main = yield* mainOf(context.sessionID)
|
||||
const interrupted = yield* sessions.interrupt(main.id, { resume: input.resume === true })
|
||||
return {
|
||||
content: interrupted ? "Interrupted the main session." : "The main session was already idle.",
|
||||
metadata: { interrupted },
|
||||
}
|
||||
}),
|
||||
})
|
||||
})
|
||||
.pipe(Effect.orDie)
|
||||
|
||||
// No client shows companion permission prompts, so asks become denials too.
|
||||
yield* ctx.permission.hook("evaluate", (event) =>
|
||||
Effect.sync(() => {
|
||||
if (event.agent !== agent) return
|
||||
if (!actions.has(event.action)) {
|
||||
event.effect = "deny"
|
||||
event.message = `The companion cannot use ${event.action}`
|
||||
return
|
||||
}
|
||||
if (event.effect !== "ask") return
|
||||
event.effect = "deny"
|
||||
event.message = "The companion cannot ask for permission"
|
||||
}),
|
||||
)
|
||||
|
||||
yield* ctx.tool.hook("execute.before", (event) => {
|
||||
if (event.agent !== agent || event.tool !== ShellTool.name || !Predicate.isObject(event.input)) return Effect.void
|
||||
const command = typeof event.input.command === "string" ? event.input.command.trim() : ""
|
||||
if (readOnlyGit.test(command) && !command.includes("--output")) return Effect.void
|
||||
return Effect.fail(
|
||||
new Tool.Error({ message: "The companion can only run one read-only git command, without shell operators" }),
|
||||
)
|
||||
})
|
||||
|
||||
// Companion tools stay out of every other agent's catalog.
|
||||
const hide = (event: SessionHooks["context"]) =>
|
||||
Effect.sync(() => {
|
||||
if (event.agent === agent) return
|
||||
names.forEach((name) => delete event.tools[name])
|
||||
})
|
||||
yield* ctx.session.hook("compaction", hide)
|
||||
yield* ctx.session.hook("generate", hide)
|
||||
yield* ctx.session.hook("context", (event) =>
|
||||
Effect.gen(function* () {
|
||||
yield* hide(event)
|
||||
if (event.agent !== agent || event.messages.at(-1)?.role !== "user") return
|
||||
const self = yield* sessions.get(event.sessionID).pipe(Effect.orElseSucceed(() => undefined))
|
||||
if (self?.kind !== "companion" || !self.parentID) return
|
||||
const main = yield* sessions.get(self.parentID).pipe(Effect.orElseSucceed(() => undefined))
|
||||
if (!main) return
|
||||
const digest = yield* Effect.all([status(main), recent(main, 6)]).pipe(Effect.orElseSucceed(() => undefined))
|
||||
if (!digest) return
|
||||
// Unpersisted and placed before the newest prompt, so earlier turns keep their cached prefix.
|
||||
event.messages.splice(
|
||||
event.messages.length - 1,
|
||||
0,
|
||||
Message.user(
|
||||
[
|
||||
"<main-session>",
|
||||
digest[0],
|
||||
"",
|
||||
"Recent activity:",
|
||||
...digest[1].map((line) => clip(line, 600)),
|
||||
"</main-session>",
|
||||
].join("\n"),
|
||||
),
|
||||
)
|
||||
}),
|
||||
)
|
||||
}),
|
||||
}
|
||||
|
||||
function describe(message: SessionMessage.Info): string[] {
|
||||
if (message.type === "user")
|
||||
return [`[user${message.metadata?.source === "companion" ? " via companion" : ""}] ${clip(message.text, 1500)}`]
|
||||
if (message.type === "synthetic") return [`[note] ${clip(message.text, 600)}`]
|
||||
if (message.type === "shell") return [`[shell ${message.status}] ${clip(message.command, 300)}`]
|
||||
if (message.type === "compaction") return [`[compaction ${message.status}]`]
|
||||
if (message.type === "idle") return [`[main session went idle: ${message.outcome}]`]
|
||||
if (message.type === "agent-switched") return [`[switched to agent ${message.agent}]`]
|
||||
if (message.type !== "assistant") return []
|
||||
const content = message.content.flatMap((part) => {
|
||||
if (part.type === "text") return part.text.trim() ? [`[assistant] ${clip(part.text, 1500)}`] : []
|
||||
if (part.type === "reasoning") return []
|
||||
const input = part.state.status === "streaming" ? part.state.input : JSON.stringify(part.state.input)
|
||||
const result =
|
||||
part.state.status === "completed" || part.state.status === "error"
|
||||
? (part.state.content ?? []).flatMap((item) => (item.type === "text" ? [item.text] : [])).join("\n")
|
||||
: ""
|
||||
return [`[tool ${part.name} ${part.state.status}] ${clip(input, 200)}${result ? `\n→ ${clip(result, 300)}` : ""}`]
|
||||
})
|
||||
if (message.error) content.push(`[error] ${clip(message.error.message, 300)}`)
|
||||
return content
|
||||
}
|
||||
|
||||
function clip(text: string, limit: number) {
|
||||
const flat = text.trim()
|
||||
if (flat.length <= limit) return flat
|
||||
return `${flat.slice(0, limit)}…`
|
||||
}
|
||||
@@ -165,6 +165,8 @@ export const Plugin = {
|
||||
return yield* new ToolFailure({
|
||||
message: `Session ${existing.id} is not a child of the current session`,
|
||||
})
|
||||
if (existing?.kind === "companion")
|
||||
return yield* new ToolFailure({ message: `Session ${existing.id} is a companion, not a subagent` })
|
||||
const override = input.model === undefined ? undefined : yield* resolveModel(input.model)
|
||||
// Continuing with a different agent switches the child, mirroring create semantics
|
||||
// where an explicit model wins over the agent's configured model, which wins over the inherited one.
|
||||
|
||||
@@ -0,0 +1,194 @@
|
||||
export * as Voice from "./voice.js"
|
||||
|
||||
import {
|
||||
Media,
|
||||
Speech,
|
||||
SpeechEvent,
|
||||
Transcription,
|
||||
type AIError,
|
||||
type SpeechModel,
|
||||
type TranscriptionModel,
|
||||
} from "@opencode/ai"
|
||||
import { SpeechClient } from "@opencode/ai/speech-client"
|
||||
import { TranscriptionClient } from "@opencode/ai/transcription-client"
|
||||
import type { ConfigModel } from "@opencode/schema/config/model"
|
||||
import { makeLocationNode } from "@opencode/util/effect/app-node"
|
||||
import { Context, Effect, Layer, Schema, Stream } from "effect"
|
||||
import { Config } from "./config.js"
|
||||
import { speechClient, transcriptionClient } from "./effect/app-node-platform.js"
|
||||
import { Integration } from "./integration.js"
|
||||
import { Plugin } from "./plugin.js"
|
||||
import { Provider } from "./provider.js"
|
||||
|
||||
export class UnavailableError extends Schema.TaggedError<UnavailableError>()("Voice.UnavailableError", {
|
||||
message: Schema.String,
|
||||
service: Schema.optional(Schema.String),
|
||||
}) {}
|
||||
|
||||
/** How `audio` chunks decode. Raw PCM is interleaved little-endian 16-bit samples. */
|
||||
export type Format =
|
||||
| { readonly type: "mp3" }
|
||||
| { readonly type: "pcm"; readonly sampleRate: number; readonly channels: number }
|
||||
|
||||
export type SpeechChunk =
|
||||
| { readonly type: "format"; readonly format: Format }
|
||||
| { readonly type: "audio"; readonly chunk: Uint8Array }
|
||||
|
||||
export interface Interface {
|
||||
readonly transcribe: (input: {
|
||||
readonly audio: Uint8Array
|
||||
readonly mediaType: string
|
||||
}) => Effect.Effect<string, UnavailableError>
|
||||
/** Fails before streaming when voice is not configured. The stream starts with one `format` chunk. */
|
||||
readonly speak: (input: {
|
||||
readonly text: string
|
||||
}) => Effect.Effect<Stream.Stream<SpeechChunk, UnavailableError>, UnavailableError>
|
||||
}
|
||||
|
||||
export class Service extends Context.Service<Service, Interface>()("@opencode/Voice") {}
|
||||
|
||||
type Facade = {
|
||||
readonly speech?: (modelID: string) => SpeechModel
|
||||
readonly transcription?: (modelID: string) => TranscriptionModel
|
||||
}
|
||||
|
||||
type Settings = { readonly apiKey?: string; readonly baseURL?: string }
|
||||
|
||||
// Speech and transcription models are not in the model catalog, so each provider package maps to its facade directly.
|
||||
const facades: Record<string, () => Promise<{ readonly configure: (settings: Settings) => Facade }>> = {
|
||||
"@opencode/ai/providers/openai": () => import("@opencode/ai/providers/openai"),
|
||||
"@opencode/ai/providers/google": () => import("@opencode/ai/providers/google"),
|
||||
"@opencode/ai/providers/xai": () => import("@opencode/ai/providers/xai"),
|
||||
"@opencode/ai/providers/elevenlabs": () => import("@opencode/ai/providers/elevenlabs"),
|
||||
"@opencode/ai/providers/deepgram": () => import("@opencode/ai/providers/deepgram"),
|
||||
"@opencode/ai/providers/assemblyai": () => import("@opencode/ai/providers/assemblyai"),
|
||||
}
|
||||
|
||||
// Gemini speech is PCM only; every other supported provider streams MP3.
|
||||
const pcm: Record<string, Format> = {
|
||||
"@opencode/ai/providers/google": { type: "pcm", sampleRate: 24000, channels: 1 },
|
||||
}
|
||||
|
||||
export const layer = Layer.effect(
|
||||
Service,
|
||||
Effect.gen(function* () {
|
||||
const config = yield* Config.Service
|
||||
const providers = yield* Provider.Service
|
||||
const integrations = yield* Integration.Service
|
||||
const speech = yield* SpeechClient.Service
|
||||
const transcription = yield* TranscriptionClient.Service
|
||||
const plugins = yield* Plugin.Service
|
||||
|
||||
const facade = Effect.fn("Voice.facade")(function* (selection: ConfigModel.Selection) {
|
||||
// Config providers register during plugin activation, which may still be running on a fresh Location.
|
||||
yield* plugins.awaitActivation
|
||||
const provider = yield* providers.get(selection.providerID)
|
||||
const connection = yield* integrations.connection.active(
|
||||
provider?.integrationID ?? Integration.ID.make(selection.providerID),
|
||||
)
|
||||
const credential = connection ? yield* integrations.connection.resolve(connection) : undefined
|
||||
const specifier = provider?.package || `@opencode/ai/providers/${selection.providerID}`
|
||||
const load = facades[specifier]
|
||||
if (!load)
|
||||
return yield* new UnavailableError({
|
||||
message: `Provider ${selection.providerID} has no speech or transcription support`,
|
||||
service: selection.providerID,
|
||||
})
|
||||
// OAuth logins target chat backends that do not serve audio endpoints; fall back to API-key environment variables.
|
||||
const settings: Settings =
|
||||
credential?.type === "oauth"
|
||||
? {}
|
||||
: {
|
||||
apiKey: credential?.key ?? stringSetting(provider?.settings?.apiKey),
|
||||
baseURL: stringSetting(provider?.settings?.baseURL),
|
||||
}
|
||||
const { configure } = yield* Effect.promise(load)
|
||||
return { specifier, facade: configure(settings) }
|
||||
})
|
||||
|
||||
const selected = Effect.fn("Voice.selected")(function* <Key extends "transcription" | "speech">(key: Key) {
|
||||
const voice = Config.latest(yield* config.entries(), "voice")
|
||||
const selection = voice?.[key]
|
||||
if (!selection)
|
||||
return yield* new UnavailableError({
|
||||
message: `Configure voice.${key}.model to use ${key === "speech" ? "text-to-speech" : "speech-to-text"}`,
|
||||
})
|
||||
return selection
|
||||
})
|
||||
|
||||
const unavailable = (providerID: string) => (error: AIError) =>
|
||||
new UnavailableError({ message: describe(error), service: providerID })
|
||||
|
||||
const transcribe: Interface["transcribe"] = Effect.fn("Voice.transcribe")(
|
||||
function* (input) {
|
||||
const selection = yield* selected("transcription")
|
||||
const loaded = yield* facade(selection.model)
|
||||
if (!loaded.facade.transcription)
|
||||
return yield* new UnavailableError({
|
||||
message: `Provider ${selection.model.providerID} does not support transcription`,
|
||||
service: selection.model.providerID,
|
||||
})
|
||||
const response = yield* Transcription.generate({
|
||||
model: loaded.facade.transcription(selection.model.model),
|
||||
audio: Media.bytes(input.audio, input.mediaType),
|
||||
language: selection.language,
|
||||
}).pipe(
|
||||
Effect.provideService(TranscriptionClient.Service, transcription),
|
||||
Effect.mapError(unavailable(selection.model.providerID)),
|
||||
)
|
||||
return response.text.trim()
|
||||
},
|
||||
Effect.catchTag("Integration.Authorization", () => Effect.fail(credentialsUnavailable)),
|
||||
)
|
||||
|
||||
const speak: Interface["speak"] = Effect.fn("Voice.speak")(
|
||||
function* (input) {
|
||||
const selection = yield* selected("speech")
|
||||
const loaded = yield* facade(selection.model)
|
||||
if (!loaded.facade.speech)
|
||||
return yield* new UnavailableError({
|
||||
message: `Provider ${selection.model.providerID} does not support speech`,
|
||||
service: selection.model.providerID,
|
||||
})
|
||||
const format = pcm[loaded.specifier] ?? { type: "mp3" as const }
|
||||
const audio = Speech.stream({
|
||||
model: loaded.facade.speech(selection.model.model),
|
||||
text: input.text,
|
||||
voice: selection.voice,
|
||||
format: format.type,
|
||||
speed: selection.speed,
|
||||
language: selection.language,
|
||||
instructions: selection.instructions,
|
||||
}).pipe(
|
||||
Stream.provideService(SpeechClient.Service, speech),
|
||||
Stream.filter(SpeechEvent.is.audioDelta),
|
||||
Stream.map((event): SpeechChunk => ({ type: "audio", chunk: event.chunk })),
|
||||
Stream.mapError(unavailable(selection.model.providerID)),
|
||||
)
|
||||
return Stream.succeed<SpeechChunk>({ type: "format", format }).pipe(Stream.concat(audio))
|
||||
},
|
||||
Effect.catchTag("Integration.Authorization", () => Effect.fail(credentialsUnavailable)),
|
||||
)
|
||||
|
||||
return Service.of({ transcribe, speak })
|
||||
}),
|
||||
)
|
||||
|
||||
const credentialsUnavailable = new UnavailableError({ message: "Voice provider credentials are unavailable" })
|
||||
|
||||
// Unrecognized provider error bodies leave only "HTTP 403" in the message; the body usually says why.
|
||||
function describe(error: AIError) {
|
||||
const body = error.reason.body?.trim()
|
||||
if (!body || body.includes(error.message)) return error.message
|
||||
return `${error.message}: ${body.slice(0, 500)}`
|
||||
}
|
||||
|
||||
function stringSetting(value: unknown) {
|
||||
return typeof value === "string" && value !== "" ? value : undefined
|
||||
}
|
||||
|
||||
export const node = makeLocationNode({
|
||||
service: Service,
|
||||
layer,
|
||||
deps: [Config.node, Provider.node, Integration.node, Plugin.node, speechClient, transcriptionClient],
|
||||
})
|
||||
@@ -180,11 +180,42 @@ describe("Agent", () => {
|
||||
expect(agents.map((item) => String(item.id)).sort()).toEqual([
|
||||
"build",
|
||||
"compaction",
|
||||
"companion",
|
||||
"explore",
|
||||
"general",
|
||||
"summary",
|
||||
"title",
|
||||
])
|
||||
const companion = (yield* agent.get(Agent.ID.make("companion")))?.permissions ?? []
|
||||
expect(Permission.evaluate("main_send", "*", companion).effect).toBe("allow")
|
||||
expect(Permission.evaluate("read", "src/index.ts", companion).effect).toBe("allow")
|
||||
expect(Permission.evaluate("read", ".env", companion).effect).toBe("deny")
|
||||
expect(Permission.evaluate("edit", "src/index.ts", companion).effect).toBe("deny")
|
||||
expect(Permission.evaluate("shell", "ls", companion).effect).toBe("deny")
|
||||
const shell = (commands: string[]) =>
|
||||
Object.fromEntries(
|
||||
commands.map((command) => [command, Permission.evaluate("shell", command, companion).effect]),
|
||||
)
|
||||
const readOnly = [
|
||||
"git status",
|
||||
"git diff --stat HEAD~1",
|
||||
"git log --oneline -5",
|
||||
"git show HEAD",
|
||||
"git branch -v",
|
||||
]
|
||||
expect(shell(readOnly)).toEqual(Object.fromEntries(readOnly.map((command) => [command, "allow"])))
|
||||
const writes = [
|
||||
"git commit -m x",
|
||||
"git checkout main",
|
||||
"git branch feature",
|
||||
"git branch -v feature",
|
||||
"git branch -D main",
|
||||
"git diff > patch.txt",
|
||||
"git log --output=log.txt",
|
||||
"git -c core.pager=sh log",
|
||||
"git statusx",
|
||||
]
|
||||
expect(shell(writes)).toEqual(Object.fromEntries(writes.map((command) => [command, "deny"])))
|
||||
expect((yield* agent.get(Agent.defaultID))?.system).toBeUndefined()
|
||||
const permissions = (yield* agent.get(Agent.defaultID))?.permissions ?? []
|
||||
const compaction = yield* agent.get(Agent.ID.make("compaction"))
|
||||
|
||||
@@ -70,6 +70,17 @@ describe("ConfigNormalize", () => {
|
||||
expect(result.agents?.reviewer?.system).toBe("Use V2")
|
||||
})
|
||||
|
||||
test("keeps voice model selections", () => {
|
||||
const result = normalized({
|
||||
voice: { transcription: { model: "xai/grok-voice-transcribe-2.0" }, speech: { model: "google/tts" } },
|
||||
})
|
||||
expect(result.encoded.voice).toEqual({
|
||||
transcription: { model: { providerID: "xai", model: "grok-voice-transcribe-2.0" } },
|
||||
speech: { model: { providerID: "google", model: "tts" } },
|
||||
})
|
||||
expect(result.diagnostics).toEqual([])
|
||||
})
|
||||
|
||||
test("canonicalizes transformed native values through decode then encode", () => {
|
||||
const result = normalized({ warming: { interval: "4 minutes", duration: "30 minutes" } })
|
||||
expect(result.encoded.warming).toEqual({ interval: "240000 millis", duration: "1800000 millis" })
|
||||
|
||||
@@ -486,6 +486,11 @@ describe("LocationServiceMap", () => {
|
||||
"edit",
|
||||
"glob",
|
||||
"grep",
|
||||
"main_cancel",
|
||||
"main_interrupt",
|
||||
"main_read",
|
||||
"main_send",
|
||||
"main_status",
|
||||
"patch",
|
||||
"question",
|
||||
"read",
|
||||
@@ -505,6 +510,11 @@ describe("LocationServiceMap", () => {
|
||||
"edit",
|
||||
"glob",
|
||||
"grep",
|
||||
"main_cancel",
|
||||
"main_interrupt",
|
||||
"main_read",
|
||||
"main_send",
|
||||
"main_status",
|
||||
"patch",
|
||||
"question",
|
||||
"read",
|
||||
|
||||
@@ -0,0 +1,284 @@
|
||||
import { describe, expect } from "bun:test"
|
||||
import { Effect, Layer } from "effect"
|
||||
import { AppNodeBuilder } from "@opencode/core/effect/app-node-builder"
|
||||
import { LayerNode } from "@opencode/util/effect/layer-node"
|
||||
import { Global } from "@opencode/util/global"
|
||||
import { makeGlobalNode, makeLocationNode } from "@opencode/util/effect/app-node"
|
||||
import { Database } from "@opencode/core/database/database"
|
||||
import { Bus } from "@opencode/core/bus"
|
||||
import { Location } from "@opencode/core/location"
|
||||
import { AbsolutePath } from "@opencode/core/schema"
|
||||
import { Job } from "@opencode/core/job"
|
||||
import { LocationServiceMap } from "@opencode/core/location-service-map"
|
||||
import { Session } from "@opencode/core/session"
|
||||
import { SessionExecution } from "@opencode/core/session/execution"
|
||||
import { Plugin } from "@opencode/core/plugin"
|
||||
import { PluginHooks } from "@opencode/core/plugin/hooks"
|
||||
import { PluginSupervisor } from "@opencode/core/plugin/supervisor"
|
||||
import { Permission } from "@opencode/core/permission"
|
||||
import { CompanionTools } from "@opencode/core/tool/plugin/companion"
|
||||
import { Tool } from "@opencode/core/tool"
|
||||
import { Agent } from "@opencode/core/agent"
|
||||
import { Model } from "@opencode/schema/model"
|
||||
import { Provider } from "@opencode/schema/provider"
|
||||
import { tmpdir } from "./fixture/tmpdir"
|
||||
import { tempGlobalLayer } from "./fixture/global"
|
||||
import { offlineModels } from "./fixture/models"
|
||||
import { testEffect } from "./lib/effect"
|
||||
import { executeTool, registerToolPlugin, toolIdentity } from "./lib/tool"
|
||||
|
||||
// Prompts are admitted durably; nothing needs to run for these tools to observe the inbox.
|
||||
const executionNode = makeGlobalNode({
|
||||
service: SessionExecution.Service,
|
||||
layer: Layer.succeed(
|
||||
SessionExecution.Service,
|
||||
SessionExecution.Service.of({
|
||||
active: Effect.succeed(new Set()),
|
||||
isActive: () => Effect.succeed(false),
|
||||
resume: () => Effect.void,
|
||||
wake: () => Effect.void,
|
||||
interrupt: () => Effect.succeed(false),
|
||||
awaitIdle: () => Effect.void,
|
||||
}),
|
||||
),
|
||||
deps: [],
|
||||
})
|
||||
|
||||
const companionPluginSupervisor = makeLocationNode({
|
||||
name: "test/companion-plugins",
|
||||
layer: Layer.effectDiscard(
|
||||
Effect.gen(function* () {
|
||||
const hooks = yield* PluginHooks.Service
|
||||
yield* registerToolPlugin(
|
||||
CompanionTools.Plugin,
|
||||
{
|
||||
permission: {
|
||||
hook: (name, callback) => hooks.register("permission", name, callback),
|
||||
list: () => Effect.die("unused permission.list"),
|
||||
get: () => Effect.die("unused permission.get"),
|
||||
reply: () => Effect.die("unused permission.reply"),
|
||||
},
|
||||
},
|
||||
(name, callback) => hooks.register("tool", name, callback),
|
||||
)
|
||||
}),
|
||||
),
|
||||
deps: [Permission.node, Session.node, Tool.node, PluginHooks.node],
|
||||
})
|
||||
|
||||
const it = testEffect(
|
||||
AppNodeBuilder.build(LayerNode.group([Database.node, Bus.node, Job.node, Session.node, LocationServiceMap.node]), [
|
||||
SessionExecution.node.replace(executionNode),
|
||||
Global.node.replace(tempGlobalLayer),
|
||||
offlineModels,
|
||||
PluginSupervisor.node.replace(companionPluginSupervisor),
|
||||
]),
|
||||
)
|
||||
|
||||
const withDirectory = <A, E, R>(body: (location: Location.Ref) => Effect.Effect<A, E, R>) =>
|
||||
Effect.acquireRelease(
|
||||
Effect.promise(() => tmpdir()),
|
||||
(dir) => Effect.promise(() => dir[Symbol.asyncDispose]()),
|
||||
).pipe(Effect.flatMap((dir) => body(Location.Ref.make({ directory: AbsolutePath.make(dir.path) }))))
|
||||
|
||||
const registry = (location: Location.Ref) =>
|
||||
Effect.gen(function* () {
|
||||
const locations = yield* LocationServiceMap.Service
|
||||
yield* Plugin.Service.use((plugins) => plugins.awaitActivation).pipe(Effect.provide(locations.get(location)))
|
||||
return yield* Tool.Service.pipe(Effect.provide(locations.get(location)))
|
||||
})
|
||||
|
||||
const call = (name: string, input: Record<string, unknown>) => ({
|
||||
type: "tool-call" as const,
|
||||
id: `call-${name}`,
|
||||
name,
|
||||
input,
|
||||
})
|
||||
|
||||
describe("CompanionTools", () => {
|
||||
it.live("gives each main session one companion", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const main = yield* sessions.create({ location, title: "Main" })
|
||||
const companion = yield* sessions.companion(main.id)
|
||||
expect(companion).toMatchObject({ parentID: main.id, kind: "companion", agent: "companion", permissions: [] })
|
||||
expect((yield* sessions.companion(main.id)).id).toBe(companion.id)
|
||||
expect((yield* sessions.companion(companion.id)).id).toBe(companion.id)
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
it.live("uses the companion agent's model before the main session's model", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const locations = yield* LocationServiceMap.Service
|
||||
const model = Model.Ref.make({ providerID: Provider.ID.make("main"), id: Model.ID.make("large") })
|
||||
const fast = Model.Ref.make({ providerID: Provider.ID.make("fast"), id: Model.ID.make("small") })
|
||||
|
||||
const fallback = yield* sessions.companion((yield* sessions.create({ location, model })).id)
|
||||
expect(fallback.model).toMatchObject(model)
|
||||
|
||||
yield* Agent.Service.use((agents) =>
|
||||
agents.transform((editor) =>
|
||||
editor.update(Agent.ID.make("companion"), (agent) => {
|
||||
agent.model = { ...fast }
|
||||
}),
|
||||
),
|
||||
).pipe(Effect.provide(locations.get(location)))
|
||||
const configured = yield* sessions.companion((yield* sessions.create({ location, model })).id)
|
||||
expect(configured.model).toMatchObject(fast)
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
it.live("denies what config rules would let the companion ask for or edit", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const locations = yield* LocationServiceMap.Service
|
||||
const main = yield* sessions.create({ location, agent: Agent.ID.make("reviewer") })
|
||||
const companion = yield* sessions.companion(main.id)
|
||||
yield* registry(location)
|
||||
const effects = yield* Effect.gen(function* () {
|
||||
const agents = yield* Agent.Service
|
||||
yield* agents.transform((editor) =>
|
||||
[Agent.ID.make("reviewer"), Agent.ID.make("companion")].forEach((id) =>
|
||||
editor.update(id, (agent) => {
|
||||
agent.permissions.push(
|
||||
{ action: "shell", resource: "*", effect: "ask" },
|
||||
{ action: "edit", resource: "*", effect: "allow" },
|
||||
)
|
||||
}),
|
||||
),
|
||||
)
|
||||
const permission = yield* Permission.Service
|
||||
return yield* Effect.forEach([main, companion], (session) =>
|
||||
Effect.forEach(
|
||||
[
|
||||
{ action: "shell", resources: ["npm test"] },
|
||||
{ action: "edit", resources: ["src/index.ts"] },
|
||||
],
|
||||
(input) =>
|
||||
permission
|
||||
.ask({ sessionID: session.id, agent: session.agent, ...input })
|
||||
.pipe(Effect.map((result) => result.effect)),
|
||||
),
|
||||
)
|
||||
}).pipe(Effect.provide(locations.get(location)))
|
||||
|
||||
expect(effects).toEqual([
|
||||
["ask", "allow"],
|
||||
["deny", "deny"],
|
||||
])
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
it.live("steers the main session from its companion", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const main = yield* sessions.create({ location, title: "Main" })
|
||||
const companion = yield* sessions.companion(main.id)
|
||||
const tools = yield* registry(location)
|
||||
|
||||
const sent = yield* executeTool(tools, {
|
||||
sessionID: companion.id,
|
||||
...toolIdentity,
|
||||
call: call("main_send", { text: "Run the tests before committing", delivery: "queue" }),
|
||||
})
|
||||
expect(sent.status).toBe("completed")
|
||||
const inbox = yield* sessions.inbox(main.id)
|
||||
expect(inbox).toEqual([
|
||||
expect.objectContaining({
|
||||
type: "user",
|
||||
delivery: "queue",
|
||||
payload: expect.objectContaining({
|
||||
text: "Run the tests before committing",
|
||||
metadata: { source: "companion", companionID: companion.id },
|
||||
}),
|
||||
}),
|
||||
])
|
||||
|
||||
const status = yield* executeTool(tools, {
|
||||
sessionID: companion.id,
|
||||
...toolIdentity,
|
||||
call: call("main_status", {}),
|
||||
})
|
||||
expect(status.content).toEqual([
|
||||
{ type: "text", text: expect.stringContaining("Run the tests before committing") },
|
||||
])
|
||||
|
||||
const cancelled = yield* executeTool(tools, {
|
||||
sessionID: companion.id,
|
||||
...toolIdentity,
|
||||
call: call("main_cancel", { inboxID: inbox[0]?.id }),
|
||||
})
|
||||
expect(cancelled.status).toBe("completed")
|
||||
expect(yield* sessions.inbox(main.id)).toEqual([])
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
it.live("limits companion shell commands to one read-only git command", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const companion = yield* sessions.companion((yield* sessions.create({ location })).id)
|
||||
const tools = yield* registry(location)
|
||||
const run = (agent: string, command: string) =>
|
||||
executeTool(tools, {
|
||||
sessionID: companion.id,
|
||||
...toolIdentity,
|
||||
agent: Agent.ID.make(agent),
|
||||
call: call("shell", { command }),
|
||||
}).pipe(Effect.map((result) => result.error?.message))
|
||||
|
||||
// Config rules such as `git *` allow these; the hook refuses them anyway.
|
||||
const refused = [
|
||||
"git status && git log > notes.txt",
|
||||
"git diff --output=patch.txt",
|
||||
"git commit -m x",
|
||||
"git reset --hard",
|
||||
"git branch -D main",
|
||||
"git status; touch notes.txt",
|
||||
"git log $(touch notes.txt)",
|
||||
"git show `touch notes.txt`",
|
||||
"git log | tee notes.txt",
|
||||
"git status\ntouch notes.txt",
|
||||
"gh pr comment 1 --body x",
|
||||
]
|
||||
for (const command of refused)
|
||||
expect(yield* run("companion", command)).toBe(
|
||||
"The companion can only run one read-only git command, without shell operators",
|
||||
)
|
||||
// Allowed commands and other agents reach tool lookup; this registry has no shell tool.
|
||||
for (const command of ["git status", "git log --oneline -5 --format='%h %s'", "git branch -v", "git diff "])
|
||||
expect(yield* run("companion", command)).toContain('No tool named "shell"')
|
||||
expect(yield* run("build", "git log > notes.txt")).toContain('No tool named "shell"')
|
||||
}),
|
||||
),
|
||||
)
|
||||
|
||||
it.live("refuses callers that are not companions", () =>
|
||||
withDirectory((location) =>
|
||||
Effect.gen(function* () {
|
||||
const sessions = yield* Session.Service
|
||||
const main = yield* sessions.create({ location, title: "Main" })
|
||||
const tools = yield* registry(location)
|
||||
const result = yield* executeTool(tools, {
|
||||
sessionID: main.id,
|
||||
...toolIdentity,
|
||||
call: call("main_send", { text: "hello" }),
|
||||
})
|
||||
expect(result).toMatchObject({
|
||||
status: "error",
|
||||
error: { message: expect.stringContaining("Only a companion session") },
|
||||
})
|
||||
expect(yield* sessions.inbox(main.id)).toEqual([])
|
||||
}),
|
||||
),
|
||||
)
|
||||
})
|
||||
@@ -636,6 +636,15 @@ describe("SubagentTool", () => {
|
||||
message: `Session ${unrelated.id} is not a child of the current session`,
|
||||
},
|
||||
})
|
||||
const companion = yield* sessions.companion(parent.id)
|
||||
expect(yield* call(companion.id, "call-companion")).toEqual({
|
||||
status: "error",
|
||||
error: {
|
||||
type: "tool.execution",
|
||||
message: `Session ${companion.id} is a companion, not a subagent`,
|
||||
},
|
||||
})
|
||||
expect(yield* sessions.get(companion.id)).toMatchObject({ agent: "companion" })
|
||||
expect(yield* call(switched.id, "call-switched-child")).toMatchObject({
|
||||
status: "completed",
|
||||
metadata: { sessionID: switched.id, status: "completed" },
|
||||
|
||||
@@ -35,6 +35,7 @@ const session = (
|
||||
project_id: Project.ID.global,
|
||||
workspace_id: null,
|
||||
parent_id: null,
|
||||
kind: null,
|
||||
fork_session_id: null,
|
||||
fork_boundary: null,
|
||||
slug: "test",
|
||||
|
||||
@@ -48,6 +48,106 @@
|
||||
"summary": "Get server info"
|
||||
}
|
||||
},
|
||||
"/api/pair": {
|
||||
"post": {
|
||||
"tags": ["server"],
|
||||
"operationId": "server.pair",
|
||||
"parameters": [],
|
||||
"security": [],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "PairingCode",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/PairingCode"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "InvalidRequestError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/InvalidRequestErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"401": {
|
||||
"description": "UnauthorizedError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"description": "Create a short-lived, single-use code for a /auth/connect/:code pairing link.",
|
||||
"summary": "Create pairing code"
|
||||
}
|
||||
},
|
||||
"/auth/connect/{code}": {
|
||||
"get": {
|
||||
"tags": ["server"],
|
||||
"operationId": "server.connect",
|
||||
"parameters": [
|
||||
{
|
||||
"name": "code",
|
||||
"in": "path",
|
||||
"schema": {
|
||||
"type": "string"
|
||||
},
|
||||
"required": true
|
||||
}
|
||||
],
|
||||
"security": [],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "PairingSession",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/PairingSession"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "InvalidRequestError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/InvalidRequestErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"401": {
|
||||
"description": "UnauthorizedError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"anyOf": [
|
||||
{
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
},
|
||||
{
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"description": "Redeem a pairing code. Browsers receive a session cookie and a redirect to the web app; requests that accept JSON receive a session token to use as the password.",
|
||||
"summary": "Redeem pairing code"
|
||||
}
|
||||
},
|
||||
"/api/location": {
|
||||
"get": {
|
||||
"tags": ["location"],
|
||||
@@ -1695,6 +1795,75 @@
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/experimental/session/{sessionID}/companion": {
|
||||
"post": {
|
||||
"tags": ["session"],
|
||||
"operationId": "experimental.session.companion",
|
||||
"parameters": [
|
||||
{
|
||||
"name": "sessionID",
|
||||
"in": "path",
|
||||
"schema": {
|
||||
"type": "string",
|
||||
"pattern": "^ses"
|
||||
},
|
||||
"required": true
|
||||
}
|
||||
],
|
||||
"security": [],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "Success",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"data": {
|
||||
"$ref": "#/components/schemas/Session.Info"
|
||||
}
|
||||
},
|
||||
"required": ["data"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "InvalidRequestError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/InvalidRequestErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"401": {
|
||||
"description": "UnauthorizedError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"404": {
|
||||
"description": "SessionNotFoundError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/SessionNotFoundErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"description": "Return the session's companion: a child session that talks with the user about the main session and can steer it. Creates the companion on first use.",
|
||||
"summary": "Get session companion"
|
||||
}
|
||||
},
|
||||
"/api/session/{sessionID}/agent": {
|
||||
"post": {
|
||||
"tags": ["session"],
|
||||
@@ -5255,6 +5424,229 @@
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/experimental/voice/transcribe": {
|
||||
"post": {
|
||||
"tags": ["voice"],
|
||||
"operationId": "experimental.voice.transcribe",
|
||||
"parameters": [
|
||||
{
|
||||
"name": "mediaType",
|
||||
"in": "query",
|
||||
"schema": {
|
||||
"type": "string",
|
||||
"description": "Media type of the uploaded audio, such as audio/wav."
|
||||
},
|
||||
"required": true
|
||||
}
|
||||
],
|
||||
"security": [],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "VoiceTranscribeResponse",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/VoiceTranscribeResponse"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "InvalidRequestError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/InvalidRequestErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"401": {
|
||||
"description": "UnauthorizedError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"503": {
|
||||
"description": "ServiceUnavailableError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/ServiceUnavailableErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"description": "Transcribe one recorded utterance with the configured voice.transcription model.",
|
||||
"summary": "Transcribe speech",
|
||||
"requestBody": {
|
||||
"content": {
|
||||
"application/octet-stream": {
|
||||
"schema": {
|
||||
"type": "string",
|
||||
"format": "binary"
|
||||
}
|
||||
}
|
||||
},
|
||||
"required": true
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/experimental/voice/speech": {
|
||||
"post": {
|
||||
"tags": ["voice"],
|
||||
"operationId": "experimental.voice.speech",
|
||||
"parameters": [],
|
||||
"security": [],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "Success",
|
||||
"content": {
|
||||
"text/event-stream": {
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"id": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
]
|
||||
},
|
||||
"event": {
|
||||
"type": "string"
|
||||
},
|
||||
"data": {
|
||||
"$ref": "#/components/schemas/VoiceSpeechEventEncoded"
|
||||
}
|
||||
},
|
||||
"required": ["id", "event", "data"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"x-effect-stream": {
|
||||
"encoding": "sse",
|
||||
"causeSchema": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"_tag": {
|
||||
"type": "string",
|
||||
"enum": ["Fail"]
|
||||
},
|
||||
"error": {
|
||||
"not": {}
|
||||
}
|
||||
},
|
||||
"required": ["_tag", "error"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"_tag": {
|
||||
"type": "string",
|
||||
"enum": ["Die"]
|
||||
},
|
||||
"defect": {}
|
||||
},
|
||||
"required": ["_tag", "defect"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"_tag": {
|
||||
"type": "string",
|
||||
"enum": ["Interrupt"]
|
||||
},
|
||||
"fiberId": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "number"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"required": ["_tag", "fiberId"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"errorSchema": {
|
||||
"not": {}
|
||||
},
|
||||
"failureEvent": "effect/httpapi/stream/failure"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "InvalidRequestError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/InvalidRequestErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"401": {
|
||||
"description": "UnauthorizedError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/UnauthorizedErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"503": {
|
||||
"description": "ServiceUnavailableError",
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"$ref": "#/components/schemas/ServiceUnavailableErrorEncoded"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"description": "Stream spoken audio for text with the configured voice.speech model. One format event precedes the audio chunks; done or error ends the stream.",
|
||||
"summary": "Synthesize speech",
|
||||
"requestBody": {
|
||||
"content": {
|
||||
"application/json": {
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"text": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["text"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"required": true
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/provider": {
|
||||
"get": {
|
||||
"tags": ["provider"],
|
||||
@@ -13003,6 +13395,47 @@
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"voice": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"transcription": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"model": {
|
||||
"$ref": "#/components/schemas/Config.ModelSelectionEncoded"
|
||||
},
|
||||
"language": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["model"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"speech": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"model": {
|
||||
"$ref": "#/components/schemas/Config.ModelSelectionEncoded"
|
||||
},
|
||||
"voice": {
|
||||
"type": "string"
|
||||
},
|
||||
"language": {
|
||||
"type": "string"
|
||||
},
|
||||
"speed": {
|
||||
"type": "number"
|
||||
},
|
||||
"instructions": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["model"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"tool_output": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -13374,6 +13807,33 @@
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"Config.ModelSelectionEncoded": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string",
|
||||
"pattern": "^[^/#]+\\/[^#]+(?:#[^#]+)?$"
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"providerID": {
|
||||
"type": "string",
|
||||
"pattern": "^[^/#]+$"
|
||||
},
|
||||
"model": {
|
||||
"type": "string",
|
||||
"pattern": "^[^#]+$"
|
||||
},
|
||||
"variant": {
|
||||
"type": "string",
|
||||
"pattern": "^[^#]+$"
|
||||
}
|
||||
},
|
||||
"required": ["providerID", "model"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
},
|
||||
"Config.Patch": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -15691,6 +16151,29 @@
|
||||
"Money.USDPerMillionTokens": {
|
||||
"type": "number"
|
||||
},
|
||||
"PairingCode": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"code": {
|
||||
"type": "string"
|
||||
},
|
||||
"expires_in": {
|
||||
"type": "integer"
|
||||
}
|
||||
},
|
||||
"required": ["code", "expires_in"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"PairingSession": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"token": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["token"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"Permission.Effect": {
|
||||
"type": "string",
|
||||
"enum": ["allow", "deny", "ask"]
|
||||
@@ -17150,6 +17633,9 @@
|
||||
"type": "string",
|
||||
"pattern": "^ses"
|
||||
},
|
||||
"kind": {
|
||||
"$ref": "#/components/schemas/Session.Kind"
|
||||
},
|
||||
"fork": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -17227,6 +17713,14 @@
|
||||
"required": ["id", "projectID", "cost", "tokens", "time", "location"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"Session.Kind": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string",
|
||||
"enum": ["companion"]
|
||||
}
|
||||
]
|
||||
},
|
||||
"Session.Message.AgentSelected": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -18635,6 +19129,9 @@
|
||||
"exit": {
|
||||
"type": "number"
|
||||
},
|
||||
"signal": {
|
||||
"type": "string"
|
||||
},
|
||||
"metadata": {
|
||||
"type": "object"
|
||||
},
|
||||
@@ -18911,6 +19408,27 @@
|
||||
"type": "string",
|
||||
"enum": ["working", "branch", "committed"]
|
||||
},
|
||||
"VoiceSpeechEventEncoded": {
|
||||
"type": "string",
|
||||
"contentMediaType": "application/json"
|
||||
},
|
||||
"VoiceTranscribeResponse": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"data": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"text": {
|
||||
"type": "string"
|
||||
}
|
||||
},
|
||||
"required": ["text"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"required": ["data"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"WebSearch.Provider": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -19097,6 +19615,10 @@
|
||||
"name": "generate",
|
||||
"description": "Experimental one-shot generation routes."
|
||||
},
|
||||
{
|
||||
"name": "voice",
|
||||
"description": "Experimental speech-to-text and text-to-speech routes."
|
||||
},
|
||||
{
|
||||
"name": "provider",
|
||||
"description": "Experimental provider routes."
|
||||
|
||||
@@ -2,6 +2,7 @@ import { Context } from "effect"
|
||||
import { HttpApi, HttpApiGroup, HttpApiMiddleware, OpenApi } from "effect/unstable/httpapi"
|
||||
import { SchemaErrorMiddleware } from "./middleware/schema-error.js"
|
||||
import { GenerateGroup } from "./groups/generate.js"
|
||||
import { VoiceGroup } from "./groups/voice.js"
|
||||
import { MessageGroup } from "./groups/message.js"
|
||||
import { ModelGroup } from "./groups/model.js"
|
||||
import { ProviderGroup } from "./groups/provider.js"
|
||||
@@ -91,6 +92,7 @@ type ApiGroups<
|
||||
| typeof MigrationGroup
|
||||
| typeof WorktreeGroup
|
||||
| typeof GenerateGroup
|
||||
| typeof VoiceGroup
|
||||
| typeof PersistentPtyGroup
|
||||
| typeof CredentialGroup
|
||||
| LocationGroups<LocationId>
|
||||
@@ -161,6 +163,7 @@ const makeApiFromGroup = <
|
||||
.add(MessageGroup)
|
||||
.add(ModelGroup.middleware(locationMiddleware))
|
||||
.add(GenerateGroup)
|
||||
.add(VoiceGroup)
|
||||
.add(ProviderGroup.middleware(locationMiddleware))
|
||||
.add(IntegrationGroup.middleware(locationMiddleware))
|
||||
.add(McpGroup.middleware(locationMiddleware))
|
||||
|
||||
@@ -43,6 +43,7 @@ export const groupNames = {
|
||||
"server.message": "message",
|
||||
"server.model": "model",
|
||||
"server.generate": "generate",
|
||||
"server.voice": "voice",
|
||||
"server.provider": "provider",
|
||||
"server.integration": "integration",
|
||||
"server.websearch": "websearch",
|
||||
|
||||
@@ -322,6 +322,20 @@ export const makeSessionGroup = <
|
||||
}),
|
||||
),
|
||||
)
|
||||
.add(
|
||||
HttpApiEndpoint.post("session.companion", "/api/experimental/session/:sessionID/companion", {
|
||||
params: { sessionID: Session.ID },
|
||||
success: Schema.Struct({ data: PublicSessionInfo }),
|
||||
error: SessionNotFoundError,
|
||||
}).annotateMerge(
|
||||
OpenApi.annotations({
|
||||
identifier: "experimental.session.companion",
|
||||
summary: "Get session companion",
|
||||
description:
|
||||
"Return the session's companion: a child session that talks with the user about the main session and can steer it. Creates the companion on first use.",
|
||||
}),
|
||||
),
|
||||
)
|
||||
.add(
|
||||
HttpApiEndpoint.post("session.switchAgent", "/api/session/:sessionID/agent", {
|
||||
params: { sessionID: Session.ID },
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
import { PositiveInt } from "@opencode/schema/schema"
|
||||
import { Schema } from "effect"
|
||||
import { HttpApiEndpoint, HttpApiGroup, HttpApiSchema, OpenApi } from "effect/unstable/httpapi"
|
||||
import { ServiceUnavailableError } from "../errors.js"
|
||||
|
||||
const SpeechFormat = Schema.Union([
|
||||
Schema.Struct({ type: Schema.Literal("mp3") }),
|
||||
Schema.Struct({ type: Schema.Literal("pcm"), sampleRate: PositiveInt, channels: PositiveInt }).annotate({
|
||||
description: "Interleaved little-endian 16-bit samples.",
|
||||
}),
|
||||
]).annotate({ identifier: "VoiceSpeechFormat" })
|
||||
|
||||
const SpeechEvent = Schema.Union([
|
||||
Schema.Struct({ type: Schema.Literal("format"), format: SpeechFormat }),
|
||||
Schema.Struct({
|
||||
type: Schema.Literal("audio"),
|
||||
data: Schema.String.annotate({ description: "Base64-encoded audio bytes." }),
|
||||
}),
|
||||
Schema.Struct({ type: Schema.Literal("done") }),
|
||||
Schema.Struct({ type: Schema.Literal("error"), message: Schema.String }),
|
||||
]).annotate({ identifier: "VoiceSpeechEvent" })
|
||||
|
||||
export const VoiceGroup = HttpApiGroup.make("server.voice")
|
||||
.add(
|
||||
HttpApiEndpoint.post("voice.transcribe", "/api/experimental/voice/transcribe", {
|
||||
query: {
|
||||
mediaType: Schema.String.annotate({ description: "Media type of the uploaded audio, such as audio/wav." }),
|
||||
},
|
||||
payload: Schema.Uint8Array.pipe(HttpApiSchema.asUint8Array()),
|
||||
success: Schema.Struct({
|
||||
data: Schema.Struct({ text: Schema.String }),
|
||||
}).annotate({ identifier: "VoiceTranscribeResponse" }),
|
||||
error: ServiceUnavailableError,
|
||||
}).annotateMerge(
|
||||
OpenApi.annotations({
|
||||
identifier: "experimental.voice.transcribe",
|
||||
summary: "Transcribe speech",
|
||||
description: "Transcribe one recorded utterance with the configured voice.transcription model.",
|
||||
}),
|
||||
),
|
||||
)
|
||||
.add(
|
||||
HttpApiEndpoint.post("voice.speech", "/api/experimental/voice/speech", {
|
||||
payload: Schema.Struct({ text: Schema.String }),
|
||||
success: HttpApiSchema.StreamSse({ data: SpeechEvent }),
|
||||
error: ServiceUnavailableError,
|
||||
}).annotateMerge(
|
||||
OpenApi.annotations({
|
||||
identifier: "experimental.voice.speech",
|
||||
summary: "Synthesize speech",
|
||||
description:
|
||||
"Stream spoken audio for text with the configured voice.speech model. One format event precedes the audio chunks; done or error ends the stream.",
|
||||
}),
|
||||
),
|
||||
)
|
||||
.annotateMerge(
|
||||
OpenApi.annotations({
|
||||
title: "voice",
|
||||
description: "Experimental speech-to-text and text-to-speech routes.",
|
||||
}),
|
||||
)
|
||||
@@ -18,6 +18,7 @@ import { ConfigProvider } from "./config/provider.js"
|
||||
import { ConfigReference } from "./config/reference.js"
|
||||
import { ConfigWebSearch } from "./config/websearch.js"
|
||||
import { ConfigToolOutput } from "./config/tool-output.js"
|
||||
import { ConfigVoice } from "./config/voice.js"
|
||||
import { ConfigWatcher } from "./config/watcher.js"
|
||||
import { ConfigWarming } from "./config/warming.js"
|
||||
import { ConfigWorktree } from "./config/worktree.js"
|
||||
@@ -72,6 +73,9 @@ export class Info extends Schema.Class<Info>("Config.Info")({
|
||||
media: ConfigMedia.Info.pipe(optional).annotate({
|
||||
description: "Media processing configuration",
|
||||
}),
|
||||
voice: ConfigVoice.Info.pipe(optional).annotate({
|
||||
description: "Speech-to-text and text-to-speech models for voice conversations",
|
||||
}),
|
||||
tool_output: ConfigToolOutput.Info.pipe(optional).annotate({
|
||||
description: "Tool output truncation thresholds",
|
||||
}),
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
export * as ConfigVoice from "./voice.js"
|
||||
|
||||
import { Schema } from "effect"
|
||||
import { optional } from "../schema.js"
|
||||
import { ConfigModel } from "./model.js"
|
||||
|
||||
export class Transcription extends Schema.Class<Transcription>("Config.Voice.Transcription")({
|
||||
model: ConfigModel.Selection.annotate({ description: "Speech-to-text model as provider/model" }),
|
||||
language: Schema.String.pipe(optional).annotate({ description: "Spoken language hint, such as en" }),
|
||||
}) {}
|
||||
|
||||
export class Speech extends Schema.Class<Speech>("Config.Voice.Speech")({
|
||||
model: ConfigModel.Selection.annotate({ description: "Text-to-speech model as provider/model" }),
|
||||
voice: Schema.String.pipe(optional).annotate({ description: "Provider-native voice name or ID" }),
|
||||
language: Schema.String.pipe(optional),
|
||||
speed: Schema.Finite.pipe(optional),
|
||||
instructions: Schema.String.pipe(optional).annotate({
|
||||
description: "Delivery instructions for providers that accept them, such as OpenAI",
|
||||
}),
|
||||
}) {}
|
||||
|
||||
export class Info extends Schema.Class<Info>("Config.Voice")({
|
||||
transcription: Transcription.pipe(optional),
|
||||
speech: Speech.pipe(optional),
|
||||
}) {}
|
||||
@@ -25,6 +25,7 @@ import { TokenUsage } from "./token-usage.js"
|
||||
import { SessionInbox } from "./session-inbox.js"
|
||||
import { Project } from "./project.js"
|
||||
import { SessionFork } from "./session-fork.js"
|
||||
import { SessionKind } from "./session-kind.js"
|
||||
import { Permission } from "./permission.js"
|
||||
|
||||
export { FileAttachment }
|
||||
@@ -57,6 +58,7 @@ export const Created = Event.durable({
|
||||
location: Location.Ref,
|
||||
subpath: RelativePath.pipe(optional),
|
||||
parentID: SessionID.pipe(optional),
|
||||
kind: SessionKind.Kind.pipe(optional),
|
||||
slug: Schema.String,
|
||||
title: Schema.String.pipe(optional),
|
||||
agent: Agent.ID.pipe(optional),
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
export * as SessionKind from "./session-kind.js"
|
||||
|
||||
import { Schema } from "effect"
|
||||
|
||||
/** A role fixed when the Session is created. Absent means an ordinary Session. */
|
||||
export const Kind = Schema.Literals(["companion"]).annotate({ identifier: "Session.Kind" })
|
||||
export type Kind = typeof Kind.Type
|
||||
@@ -14,10 +14,14 @@ import { Permission } from "./permission.js"
|
||||
import { TokenUsage } from "./token-usage.js"
|
||||
import { Revert } from "./session-revert.js"
|
||||
import { SessionFork } from "./session-fork.js"
|
||||
import { SessionKind } from "./session-kind.js"
|
||||
|
||||
export const ID = SessionID
|
||||
export type ID = SessionID
|
||||
|
||||
export const Kind = SessionKind.Kind
|
||||
export type Kind = SessionKind.Kind
|
||||
|
||||
export const Metadata = SessionMetadata
|
||||
export type Metadata = SessionMetadata
|
||||
|
||||
@@ -31,6 +35,7 @@ export interface Info extends Schema.Schema.Type<typeof Info> {}
|
||||
export const Info = Schema.Struct({
|
||||
id: ID,
|
||||
parentID: ID.pipe(optional),
|
||||
kind: Kind.pipe(optional),
|
||||
fork: Schema.Struct({
|
||||
sessionID: ID,
|
||||
boundary: ForkBoundary,
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import { Layer } from "effect"
|
||||
import { GenerateHandler } from "./handlers/generate"
|
||||
import { VoiceHandler } from "./handlers/voice"
|
||||
import { MessageHandler } from "./handlers/message"
|
||||
import { ModelHandler } from "./handlers/model"
|
||||
import { ProviderHandler } from "./handlers/provider"
|
||||
@@ -42,6 +43,7 @@ export const handlers = Layer.mergeAll(
|
||||
MessageHandler,
|
||||
ModelHandler,
|
||||
GenerateHandler,
|
||||
VoiceHandler,
|
||||
ProviderHandler,
|
||||
IntegrationHandler,
|
||||
WebSearchHandler,
|
||||
|
||||
@@ -234,6 +234,16 @@ export const SessionHandler = HttpApiBuilder.group(Api, "server.session", (handl
|
||||
}
|
||||
}),
|
||||
)
|
||||
.handle(
|
||||
"session.companion",
|
||||
Effect.fn(function* (ctx) {
|
||||
return {
|
||||
data: yield* session
|
||||
.companion(ctx.params.sessionID)
|
||||
.pipe(Effect.catchTag("Session.NotFoundError", missingSession)),
|
||||
}
|
||||
}),
|
||||
)
|
||||
.handle(
|
||||
"session.switchAgent",
|
||||
Effect.fn(function* (ctx) {
|
||||
|
||||
@@ -0,0 +1,49 @@
|
||||
import { Location } from "@opencode/core/location"
|
||||
import { LocationServiceMap } from "@opencode/core/location-services"
|
||||
import { AbsolutePath } from "@opencode/core/schema"
|
||||
import { Voice } from "@opencode/core/voice"
|
||||
import { ServiceUnavailableError } from "@opencode/protocol/errors"
|
||||
import { Global } from "@opencode/util/global"
|
||||
import { Effect, Encoding, Stream } from "effect"
|
||||
import { HttpApiBuilder } from "effect/unstable/httpapi"
|
||||
import { Api } from "../api"
|
||||
|
||||
const unavailable = (error: Voice.UnavailableError) =>
|
||||
new ServiceUnavailableError({ message: error.message, service: error.service ?? "voice" })
|
||||
|
||||
export const VoiceHandler = HttpApiBuilder.group(Api, "server.voice", (handlers) =>
|
||||
Effect.gen(function* () {
|
||||
const global = yield* Global.Service
|
||||
const locations = yield* LocationServiceMap.Service
|
||||
// Voice models come from the base configuration, like one-shot generation.
|
||||
const services = locations.get(Location.Ref.make({ directory: AbsolutePath.make(global.config) }))
|
||||
return handlers
|
||||
.handle(
|
||||
"voice.transcribe",
|
||||
Effect.fn("server.voice.transcribe")(function* (request) {
|
||||
const voice = yield* Voice.Service
|
||||
const text = yield* voice
|
||||
.transcribe({ audio: request.payload, mediaType: request.query.mediaType })
|
||||
.pipe(Effect.mapError(unavailable))
|
||||
return { data: { text } }
|
||||
}, Effect.provide(services)),
|
||||
)
|
||||
.handle(
|
||||
"voice.speech",
|
||||
Effect.fn("server.voice.speech")(function* (request) {
|
||||
const voice = yield* Voice.Service
|
||||
const audio = yield* voice.speak({ text: request.payload.text }).pipe(Effect.mapError(unavailable))
|
||||
return audio.pipe(
|
||||
Stream.map((chunk) =>
|
||||
chunk.type === "format" ? chunk : { type: "audio" as const, data: Encoding.encodeBase64(chunk.chunk) },
|
||||
),
|
||||
Stream.concat(Stream.succeed({ type: "done" as const })),
|
||||
// Provider failures arrive after the response started, so they travel as a terminal event.
|
||||
Stream.catchTag("Voice.UnavailableError", (error) =>
|
||||
Stream.succeed({ type: "error" as const, message: error.message }),
|
||||
),
|
||||
)
|
||||
}, Effect.provide(services)),
|
||||
)
|
||||
}),
|
||||
)
|
||||
@@ -4,7 +4,8 @@ import { readFile } from "node:fs/promises"
|
||||
let audio: Audio | null | undefined
|
||||
const sounds = new Map<string, Promise<AudioSound | null>>()
|
||||
|
||||
function getAudio() {
|
||||
// One engine per process: OpenTUI allows a single capture owner, and all playback shares one mixer.
|
||||
export function getAudio() {
|
||||
if (audio !== undefined) return audio
|
||||
try {
|
||||
const next = Audio.create({ autoStart: false })
|
||||
|
||||
@@ -13,7 +13,18 @@ type Experiment = {
|
||||
// In-flight features anyone can opt into. Each entry is temporary: an
|
||||
// experiment either graduates (delete the entry, make the behavior
|
||||
// unconditional) or dies (delete the entry and the branch it gated).
|
||||
export const experiments: Experiment[] = []
|
||||
export const experiments: Experiment[] = [
|
||||
{
|
||||
id: "companion",
|
||||
title: "Companion",
|
||||
description: "Talk with a companion, by text or voice, that observes and steers the current session",
|
||||
},
|
||||
{
|
||||
id: "companion_barge_in",
|
||||
title: "Companion barge-in",
|
||||
description: "Keep the microphone open while the companion speaks, so you can talk over and interrupt it",
|
||||
},
|
||||
]
|
||||
|
||||
export function DialogExperiments() {
|
||||
const config = useConfig()
|
||||
|
||||
@@ -127,6 +127,10 @@ export const Definitions = {
|
||||
"session.background": keybind("ctrl+b", "Background blocking session tools"),
|
||||
"session.compact": keybind("<leader>c", "Compact the session"),
|
||||
"session.aside": keybind("none", "Ask a side question"),
|
||||
"session.companion": keybind("<leader>o", "Talk with the session companion"),
|
||||
"companion.talk": keybind("<leader>v,alt+v", "Start or finish a voice message to the companion"),
|
||||
"companion.submit": keybind("return", "Send a message to the companion"),
|
||||
"companion.stop": keybind("escape", "Stop the companion's voice and reply"),
|
||||
"session.cd": keybind("none", "Change working directory"),
|
||||
"session.queued_prompts": keybind("<leader>q", "Manage queued prompts"),
|
||||
"queued_prompt.delete": keybind("ctrl+d", "Delete queued prompt"),
|
||||
|
||||
@@ -125,7 +125,7 @@ export const { use: useSessionTabs, provider: SessionTabsProvider } = createSimp
|
||||
}
|
||||
const family = (sessionID: string) => {
|
||||
const session = root(sessionID)
|
||||
const members = data.session.family(session)
|
||||
const members = data.session.family(session).filter((id) => data.session.get(id)?.kind !== "companion")
|
||||
return members.length > 0 ? members : [session]
|
||||
}
|
||||
const normalize = (value: TabsState) => ({
|
||||
|
||||
@@ -21,7 +21,12 @@ export function PromptFooter(props: {
|
||||
if (!props.sessionID) return 0
|
||||
const count = props.context.data.session
|
||||
.family(props.sessionID)
|
||||
.filter((id) => id !== props.sessionID && props.context.data.session.status(id) === "running").length
|
||||
.filter(
|
||||
(id) =>
|
||||
id !== props.sessionID &&
|
||||
props.context.data.session.get(id)?.kind !== "companion" &&
|
||||
props.context.data.session.status(id) === "running",
|
||||
).length
|
||||
return count ? `${count} subagent${count === 1 ? "" : "s"}` : undefined
|
||||
})
|
||||
const shells = createMemo(() => {
|
||||
|
||||
@@ -0,0 +1,177 @@
|
||||
import type { Audio, AudioCaptureStream, AudioStream } from "@opentui/core"
|
||||
import type { VoiceSpeechFormat } from "@opencode/client"
|
||||
import { getAudio } from "../../../audio"
|
||||
import { concat, energy } from "./barge"
|
||||
|
||||
const TRANSCRIBE_RATE = 16000
|
||||
// About a third of a second at 48 kHz: longer than the speaker-to-microphone delay.
|
||||
const TAP_FRAMES = 16384
|
||||
const TAP_BLOCK = 1024
|
||||
|
||||
export type Recording = {
|
||||
/** Stops capture and returns the utterance as 16 kHz mono WAV, or undefined when nothing was captured. */
|
||||
readonly stop: () => Promise<Uint8Array | undefined>
|
||||
readonly cancel: () => void
|
||||
}
|
||||
|
||||
export async function record(onLevel: (level: number) => void): Promise<Recording> {
|
||||
const audio = getAudio()
|
||||
if (!audio) throw new Error("Audio is unavailable in this terminal")
|
||||
const capture = await audio.openCapture({ channels: 1 })
|
||||
const chunks: Float32Array[] = []
|
||||
const reader = capture.readable.getReader()
|
||||
const pump = (async () => {
|
||||
while (true) {
|
||||
const next = await reader.read()
|
||||
if (next.done) return
|
||||
chunks.push(next.value)
|
||||
onLevel(rms(next.value))
|
||||
}
|
||||
})().catch(() => undefined)
|
||||
return {
|
||||
async stop() {
|
||||
await finish(capture, pump)
|
||||
const samples = concat(chunks)
|
||||
if (samples.length < capture.sampleRate / 4) return undefined
|
||||
return encode(samples, capture.sampleRate)
|
||||
},
|
||||
cancel() {
|
||||
void finish(capture, pump)
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/** Opens the microphone during playback. Each chunk comes with the loudest block of recent playback. */
|
||||
export async function monitor(onChunk: (samples: Float32Array, sampleRate: number, playback: number) => void) {
|
||||
const audio = getAudio()
|
||||
if (!audio) throw new Error("Audio is unavailable in this terminal")
|
||||
if (!audio.enableTap(TAP_FRAMES)) throw new Error("Audio playback cannot be monitored")
|
||||
const capture = await audio.openCapture({ channels: 1 }).catch((error) => {
|
||||
audio.disableTap()
|
||||
return Promise.reject(error)
|
||||
})
|
||||
const reader = capture.readable.getReader()
|
||||
const pump = (async () => {
|
||||
while (true) {
|
||||
const next = await reader.read()
|
||||
if (next.done) return
|
||||
onChunk(next.value, capture.sampleRate, playbackLevel(audio))
|
||||
}
|
||||
})().catch(() => undefined)
|
||||
return {
|
||||
async close() {
|
||||
await finish(capture, pump)
|
||||
audio.disableTap()
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
function playbackLevel(audio: Audio) {
|
||||
const tap = audio.readTapFrames(TAP_FRAMES, 1)
|
||||
if (!tap) return 0
|
||||
return Array.from({ length: Math.ceil(tap.framesRead / TAP_BLOCK) }, (_, index) =>
|
||||
energy(tap.frames.subarray(index * TAP_BLOCK, Math.min(tap.framesRead, (index + 1) * TAP_BLOCK))),
|
||||
).reduce((loudest, level) => Math.max(loudest, level), 0)
|
||||
}
|
||||
|
||||
/** Encodes captured samples as 16 kHz mono WAV for transcription. */
|
||||
export function encode(samples: Float32Array, sampleRate: number) {
|
||||
return wav(resample(samples, sampleRate, TRANSCRIBE_RATE), TRANSCRIBE_RATE)
|
||||
}
|
||||
|
||||
async function finish(capture: AudioCaptureStream, pump: Promise<unknown>) {
|
||||
capture.stop()
|
||||
await pump
|
||||
capture.dispose()
|
||||
}
|
||||
|
||||
/** Plays one clip and resolves when playback ends or is stopped. */
|
||||
export async function play(format: VoiceSpeechFormat, body: AsyncIterable<Uint8Array>, signal: AbortSignal) {
|
||||
const audio = getAudio()
|
||||
if (!audio) throw new Error("Audio is unavailable in this terminal")
|
||||
if (!audio.isStarted() && !audio.start()) throw new Error("Audio playback could not start")
|
||||
const stream: AudioStream = await (
|
||||
format.type === "mp3"
|
||||
? audio.playStream(body, { format: "mp3", signal })
|
||||
: audio.playStream(pcm(body), {
|
||||
format: "pcm",
|
||||
sampleFormat: "f32le",
|
||||
sampleRate: format.sampleRate,
|
||||
channels: format.channels === 2 ? 2 : 1,
|
||||
signal,
|
||||
})
|
||||
).catch((error) => Promise.reject(cause(error)))
|
||||
// Aborting the signal disposes the stream, which closes it.
|
||||
await Promise.race([
|
||||
stream.closed,
|
||||
new Promise<never>((_, reject) => stream.on("error", (error) => reject(cause(error)))),
|
||||
])
|
||||
}
|
||||
|
||||
// OpenTUI wraps source failures; the provider's message is more useful than "source failed".
|
||||
function cause(error: unknown) {
|
||||
return error instanceof Error && error.cause instanceof Error ? error.cause : error
|
||||
}
|
||||
|
||||
// Providers return signed 16-bit PCM; OpenTUI plays float PCM.
|
||||
async function* pcm(body: AsyncIterable<Uint8Array>) {
|
||||
let carry = new Uint8Array(0)
|
||||
for await (const chunk of body) {
|
||||
const bytes = carry.length === 0 ? chunk : concatBytes(carry, chunk)
|
||||
const even = bytes.length - (bytes.length % 2)
|
||||
carry = bytes.slice(even)
|
||||
const view = new DataView(bytes.buffer, bytes.byteOffset, even)
|
||||
const samples = new Float32Array(even / 2)
|
||||
for (let index = 0; index < samples.length; index++) samples[index] = view.getInt16(index * 2, true) / 32768
|
||||
yield new Uint8Array(samples.buffer)
|
||||
}
|
||||
}
|
||||
|
||||
/** Level for the meter, from 0 to 1. */
|
||||
export function rms(samples: Float32Array) {
|
||||
return Math.min(1, energy(samples) * 4)
|
||||
}
|
||||
|
||||
function concatBytes(left: Uint8Array, right: Uint8Array) {
|
||||
const result = new Uint8Array(left.length + right.length)
|
||||
result.set(left)
|
||||
result.set(right, left.length)
|
||||
return result
|
||||
}
|
||||
|
||||
// Box-filter decimation: averaging each output window also removes most aliasing.
|
||||
function resample(samples: Float32Array, from: number, to: number) {
|
||||
if (from === to) return samples
|
||||
const ratio = from / to
|
||||
const result = new Float32Array(Math.floor(samples.length / ratio))
|
||||
for (let index = 0; index < result.length; index++) {
|
||||
const start = Math.floor(index * ratio)
|
||||
const end = Math.max(start + 1, Math.min(samples.length, Math.floor((index + 1) * ratio)))
|
||||
let sum = 0
|
||||
for (let cursor = start; cursor < end; cursor++) sum += samples[cursor] ?? 0
|
||||
result[index] = sum / (end - start)
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
function wav(samples: Float32Array, sampleRate: number) {
|
||||
const bytes = new Uint8Array(44 + samples.length * 2)
|
||||
const view = new DataView(bytes.buffer)
|
||||
const ascii = (offset: number, text: string) =>
|
||||
Array.from(text).forEach((char, index) => view.setUint8(offset + index, char.charCodeAt(0)))
|
||||
ascii(0, "RIFF")
|
||||
view.setUint32(4, 36 + samples.length * 2, true)
|
||||
ascii(8, "WAVE")
|
||||
ascii(12, "fmt ")
|
||||
view.setUint32(16, 16, true)
|
||||
view.setUint16(20, 1, true)
|
||||
view.setUint16(22, 1, true)
|
||||
view.setUint32(24, sampleRate, true)
|
||||
view.setUint32(28, sampleRate * 2, true)
|
||||
view.setUint16(32, 2, true)
|
||||
view.setUint16(34, 16, true)
|
||||
ascii(36, "data")
|
||||
view.setUint32(40, samples.length * 2, true)
|
||||
samples.forEach((sample, index) => view.setInt16(44 + index * 2, Math.max(-1, Math.min(1, sample)) * 0x7fff, true))
|
||||
return bytes
|
||||
}
|
||||
@@ -0,0 +1,122 @@
|
||||
const PREROLL_MS = 400
|
||||
const ONSET_MS = 150
|
||||
const END_MS = 700
|
||||
const MAX_MS = 15_000
|
||||
// Raw RMS (about -46 dBFS) below which nothing counts as speech, even in a silent room.
|
||||
const MIN_LEVEL = 0.005
|
||||
const FLOOR_RATIO = 3
|
||||
const ECHO_MARGIN = 2
|
||||
const MIN_PLAYBACK = 0.005
|
||||
|
||||
export type Utterance = { readonly samples: Float32Array; readonly sampleRate: number }
|
||||
|
||||
/**
|
||||
* Detects the user talking over playback. The microphone also hears the
|
||||
* speakers, so a chunk counts as speech only when it is louder than the
|
||||
* noise floor and louder than `coupling` times the recent playback level
|
||||
* would explain. `coupling` follows the echo it hears between words and
|
||||
* rises when a detected utterance turns out to be echo (`rejected`).
|
||||
*/
|
||||
export function createBargeDetector(input: {
|
||||
readonly onOnset: () => void
|
||||
readonly onUtterance: (utterance: Utterance) => void
|
||||
}) {
|
||||
let coupling = 0.5
|
||||
let floor = MIN_LEVEL
|
||||
let preroll: { samples: Float32Array; ms: number }[] = []
|
||||
let prerollMs = 0
|
||||
let run = { ms: 0, ratio: 0 }
|
||||
let onsetRatio = 0
|
||||
let speech: { chunks: Float32Array[]; sampleRate: number; ms: number; silence: number } | undefined
|
||||
|
||||
const finish = () => {
|
||||
if (!speech) return
|
||||
const utterance = { samples: concat(speech.chunks), sampleRate: speech.sampleRate }
|
||||
speech = undefined
|
||||
input.onUtterance(utterance)
|
||||
}
|
||||
|
||||
return {
|
||||
/** Feeds one microphone chunk with the loudest playback level of the last few hundred milliseconds. */
|
||||
push(samples: Float32Array, sampleRate: number, playback: number) {
|
||||
const ms = (samples.length / sampleRate) * 1000
|
||||
const mic = energy(samples)
|
||||
floor = Math.min(Math.max(mic, 1e-4), floor * 1.002)
|
||||
const ratio = playback > MIN_PLAYBACK ? mic / playback : 0
|
||||
const loud = mic > Math.max(MIN_LEVEL, floor * FLOOR_RATIO) && mic > coupling * playback * ECHO_MARGIN
|
||||
|
||||
if (speech) {
|
||||
speech.chunks.push(samples)
|
||||
speech.ms += ms
|
||||
speech.silence = loud ? 0 : speech.silence + ms
|
||||
if (speech.silence >= END_MS || speech.ms >= MAX_MS) finish()
|
||||
return
|
||||
}
|
||||
|
||||
preroll.push({ samples, ms })
|
||||
prerollMs += ms
|
||||
while (preroll.length > 1 && prerollMs - (preroll[0]?.ms ?? 0) >= PREROLL_MS)
|
||||
prerollMs -= preroll.shift()?.ms ?? 0
|
||||
if (!loud) {
|
||||
run = { ms: 0, ratio: 0 }
|
||||
// Follow the echo slowly in both directions, so real speech above it stays detectable.
|
||||
if (ratio > 0) coupling = ratio > coupling ? Math.min(ratio, coupling * 1.01) : Math.max(0.05, coupling * 0.999)
|
||||
return
|
||||
}
|
||||
run = { ms: run.ms + ms, ratio: Math.max(run.ratio, ratio) }
|
||||
if (run.ms < ONSET_MS) return
|
||||
onsetRatio = run.ratio
|
||||
speech = { chunks: preroll.map((chunk) => chunk.samples), sampleRate, ms: prerollMs, silence: 0 }
|
||||
preroll = []
|
||||
prerollMs = 0
|
||||
run = { ms: 0, ratio: 0 }
|
||||
input.onOnset()
|
||||
},
|
||||
/** Ends the current utterance now. */
|
||||
flush: finish,
|
||||
/** Marks the last utterance as the speakers' own echo, so the same level no longer counts as speech. */
|
||||
rejected() {
|
||||
coupling = Math.min(4, Math.max(coupling, onsetRatio * 0.75))
|
||||
},
|
||||
/** Drops buffered audio; the learned echo level and noise floor stay. */
|
||||
reset() {
|
||||
preroll = []
|
||||
prerollMs = 0
|
||||
run = { ms: 0, ratio: 0 }
|
||||
speech = undefined
|
||||
},
|
||||
get active() {
|
||||
return speech !== undefined
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
export type BargeDetector = ReturnType<typeof createBargeDetector>
|
||||
|
||||
/** Whether a transcript mostly repeats what the speakers just said. */
|
||||
export function echoes(transcript: string, spoken: string) {
|
||||
const heard = words(transcript)
|
||||
const said = new Set(words(spoken))
|
||||
return heard.length > 0 && heard.filter((word) => said.has(word)).length / heard.length >= 0.6
|
||||
}
|
||||
|
||||
function words(text: string) {
|
||||
return text.toLowerCase().match(/[\p{L}\p{N}']+/gu) ?? []
|
||||
}
|
||||
|
||||
/** Root-mean-square amplitude. */
|
||||
export function energy(samples: Float32Array) {
|
||||
if (samples.length === 0) return 0
|
||||
let sum = 0
|
||||
for (const sample of samples) sum += sample * sample
|
||||
return Math.sqrt(sum / samples.length)
|
||||
}
|
||||
|
||||
export function concat(chunks: Float32Array[]) {
|
||||
const result = new Float32Array(chunks.reduce((total, chunk) => total + chunk.length, 0))
|
||||
chunks.reduce((offset, chunk) => {
|
||||
result.set(chunk, offset)
|
||||
return offset + chunk.length
|
||||
}, 0)
|
||||
return result
|
||||
}
|
||||
@@ -0,0 +1,194 @@
|
||||
import { Plugin } from "@opencode/plugin/tui"
|
||||
import { createEffect, createResource, createRoot, createSignal, Show } from "solid-js"
|
||||
import { Spinner } from "../../../component/spinner"
|
||||
import { useConfig } from "../../../config"
|
||||
import { Keymap } from "../../../context/keymap"
|
||||
import { useTheme } from "../../../context/theme"
|
||||
import { errorMessage } from "../../../util/error"
|
||||
import { CompanionPanel } from "./panel"
|
||||
import { createVoice, sentences, speakable } from "./voice"
|
||||
|
||||
const PANEL = "companion"
|
||||
|
||||
export default Plugin.define({
|
||||
id: "opencode.companion",
|
||||
setup(context) {
|
||||
const companions = new Map<string, Promise<string>>()
|
||||
const toastError = (error: unknown) => context.ui.toast.show({ variant: "error", message: errorMessage(error) })
|
||||
// Assistant messages to read aloud: every reply after a spoken prompt, until the user types or talks over it.
|
||||
const [follow, setFollow] = createSignal<{ sessionID: string; known: ReadonlySet<string> }>()
|
||||
const spoken = new Map<string, number>()
|
||||
const [talking, setTalking] = createSignal<string>()
|
||||
const [bargeIn, setBargeIn] = createSignal(false)
|
||||
|
||||
// One lookup per main session, so the panel and the talk key cannot race to create two companions.
|
||||
const lookup = (sessionID: string) => {
|
||||
const cached = companions.get(sessionID)
|
||||
if (cached) return cached
|
||||
const request = context.client.session.companion({ sessionID }).then((info) => info.id)
|
||||
request.catch(() => companions.delete(sessionID))
|
||||
companions.set(sessionID, request)
|
||||
return request
|
||||
}
|
||||
// Session retention evicts the companion's messages with its main session, so every use syncs again.
|
||||
const companionOf = (sessionID: string) =>
|
||||
lookup(sessionID).then(async (companionID) => {
|
||||
await Promise.all([context.data.session.sync(companionID), context.data.session.message.sync(companionID)])
|
||||
return companionID
|
||||
})
|
||||
|
||||
const send = (sessionID: string, text: string, voice: boolean) => {
|
||||
const known = new Set(
|
||||
context.data.session.message
|
||||
.list(sessionID)
|
||||
.flatMap((message) => (message.type === "assistant" ? [message.id] : [])),
|
||||
)
|
||||
setFollow(voice ? { sessionID, known } : undefined)
|
||||
void context.client.session
|
||||
.prompt({ sessionID, text, ...(voice ? { metadata: { voice: true } } : {}) })
|
||||
.catch(toastError)
|
||||
}
|
||||
|
||||
const voice = createVoice({
|
||||
client: context.client,
|
||||
bargeIn,
|
||||
onTranscript: (text, interrupting) => {
|
||||
const sessionID = talking()
|
||||
if (!sessionID) return
|
||||
if (!interrupting) return send(sessionID, text, true)
|
||||
// A prompt admitted before the interrupt would be stranded in the stopped reply's inbox.
|
||||
void stop(sessionID).then(() => send(sessionID, text, true))
|
||||
},
|
||||
onError: toastError,
|
||||
})
|
||||
|
||||
const stop = (sessionID: string | undefined) => {
|
||||
setFollow(undefined)
|
||||
voice.stop()
|
||||
if (!sessionID || context.data.session.status(sessionID) !== "running") return Promise.resolve()
|
||||
return context.client.session.interrupt({ sessionID }).then(() => undefined, toastError)
|
||||
}
|
||||
|
||||
const current = () => {
|
||||
const route = context.ui.router.current()
|
||||
if (route.type === "session") return route.sessionID
|
||||
context.ui.toast.show({ message: "Open a session first", variant: "warning" })
|
||||
}
|
||||
|
||||
const toggle = () => {
|
||||
if (context.ui.panel.current()?.name === PANEL) return context.ui.panel.close()
|
||||
if (!current()) return
|
||||
context.ui.panel.open(PANEL)
|
||||
}
|
||||
|
||||
const talk = async () => {
|
||||
const sessionID = current()
|
||||
if (!sessionID) return
|
||||
if (context.ui.panel.current()?.name !== PANEL) context.ui.panel.open(PANEL)
|
||||
const companionID = await companionOf(sessionID).catch((error) => {
|
||||
toastError(error)
|
||||
return undefined
|
||||
})
|
||||
if (!companionID) return
|
||||
// Talking over the companion stops its reply so the new utterance takes the turn.
|
||||
if (voice.state.status === "idle" || voice.state.status === "speaking") void stop(companionID)
|
||||
setTalking(companionID)
|
||||
voice.toggle()
|
||||
}
|
||||
|
||||
const dispose = createRoot((dispose) => {
|
||||
createEffect(() => {
|
||||
const target = follow()
|
||||
if (!target) return
|
||||
context.data.session.message.list(target.sessionID).forEach((message) => {
|
||||
if (message.type !== "assistant" || target.known.has(message.id)) return
|
||||
const text = speakable(
|
||||
message.content.flatMap((part) => (part.type === "text" ? [part.text] : [])).join("\n\n"),
|
||||
)
|
||||
const next = sentences(text, spoken.get(message.id) ?? 0, message.time.completed !== undefined)
|
||||
spoken.set(message.id, next.end)
|
||||
next.chunks.forEach(voice.speak)
|
||||
})
|
||||
})
|
||||
return dispose
|
||||
})
|
||||
|
||||
context.ui.slot({
|
||||
append: "session.panel",
|
||||
render: (input) => {
|
||||
const [companion] = createResource(
|
||||
() => (input.name === PANEL ? input.sessionID : undefined),
|
||||
(sessionID) => companionOf(sessionID),
|
||||
)
|
||||
return (
|
||||
<Show when={input.name === PANEL}>
|
||||
<CompanionPanel
|
||||
input={input}
|
||||
companionID={companion()}
|
||||
error={companion.error ? errorMessage(companion.error) : undefined}
|
||||
voice={voice}
|
||||
onSubmit={(text) => {
|
||||
voice.stop()
|
||||
// The companion may still be loading when the first message is typed.
|
||||
void companionOf(input.sessionID).then((companionID) => send(companionID, text, false), toastError)
|
||||
}}
|
||||
onStop={() => void stop(companion())}
|
||||
/>
|
||||
</Show>
|
||||
)
|
||||
},
|
||||
})
|
||||
|
||||
context.ui.slot({
|
||||
append: "prompt.footer.status",
|
||||
render: () => {
|
||||
const theme = useTheme()
|
||||
return (
|
||||
<Show when={voice.state.status !== "idle"}>
|
||||
<box flexShrink={0}>
|
||||
<Spinner color={theme.text.feedback.info.base}>{voice.state.status}</Spinner>
|
||||
</box>
|
||||
</Show>
|
||||
)
|
||||
},
|
||||
})
|
||||
|
||||
context.ui.slot({
|
||||
append: "app",
|
||||
render() {
|
||||
const config = useConfig()
|
||||
const enabled = () => config.data.experimental?.companion === true
|
||||
createEffect(() => setBargeIn(enabled() && config.data.experimental?.companion_barge_in === true))
|
||||
Keymap.createLayer(() => ({
|
||||
mode: "global",
|
||||
enabled,
|
||||
commands: [
|
||||
{
|
||||
id: "session.companion",
|
||||
title: "Talk with the companion",
|
||||
description: "Open a persistent side conversation that can observe and steer this session",
|
||||
group: "Session",
|
||||
palette: true,
|
||||
slash: { name: "companion" },
|
||||
run: toggle,
|
||||
},
|
||||
{
|
||||
id: "companion.talk",
|
||||
title: "Talk to the companion by voice",
|
||||
description: "Start listening, or send what you said",
|
||||
group: "Session",
|
||||
palette: true,
|
||||
run: () => void talk(),
|
||||
},
|
||||
],
|
||||
}))
|
||||
return null
|
||||
},
|
||||
})
|
||||
|
||||
return () => {
|
||||
dispose()
|
||||
voice.stop()
|
||||
}
|
||||
},
|
||||
})
|
||||
@@ -0,0 +1,239 @@
|
||||
import type { PanelInput } from "@opencode/plugin/tui/context"
|
||||
import type { SessionMessageAssistant, SessionMessageAssistantTool, SessionMessageInfo } from "@opencode/client"
|
||||
import { TextAttributes, type RGBA, type TextareaRenderable } from "@opentui/core"
|
||||
import { createMemo, createSignal, For, Match, onMount, Show, Switch } from "solid-js"
|
||||
import { Spinner } from "../../../component/spinner"
|
||||
import { useConfig } from "../../../config"
|
||||
import { useData } from "../../../context/data"
|
||||
import { Keymap } from "../../../context/keymap"
|
||||
import { useTheme, useThemes } from "../../../context/theme"
|
||||
import { usePlugin } from "../../../plugin/context"
|
||||
import type { Voice } from "./voice"
|
||||
|
||||
const LEVELS = "▁▂▃▄▅▆▇█"
|
||||
|
||||
export function CompanionPanel(props: {
|
||||
input: PanelInput
|
||||
companionID?: string
|
||||
error?: string
|
||||
voice: Voice
|
||||
onSubmit: (text: string) => void
|
||||
onStop: () => void
|
||||
}) {
|
||||
const theme = useTheme()
|
||||
const data = useData()
|
||||
const config = useConfig().data
|
||||
const shortcuts = Keymap.useShortcuts()
|
||||
const [target, setTarget] = createSignal<TextareaRenderable>()
|
||||
let textarea: TextareaRenderable | undefined
|
||||
|
||||
const messages = createMemo(() => (props.companionID ? data.session.message.list(props.companionID) : []))
|
||||
const running = () => (props.companionID ? data.session.status(props.companionID) === "running" : false)
|
||||
const background = () => (props.input.presentation === "panel" ? theme.background.raised.base : theme.background.base)
|
||||
|
||||
const submit = () => {
|
||||
const text = textarea?.plainText.trim()
|
||||
if (!text || !textarea) return
|
||||
textarea.clear()
|
||||
props.onSubmit(text)
|
||||
}
|
||||
|
||||
Keymap.createLayer(() => ({
|
||||
target,
|
||||
enabled: target() !== undefined,
|
||||
// Submitting must win over the managed textarea's newline binding.
|
||||
priority: 1,
|
||||
commands: [{ id: "companion.submit", title: "Send to companion", group: "Companion", run: submit }],
|
||||
}))
|
||||
Keymap.createLayer(() => ({
|
||||
commands: [{ id: "companion.stop", title: "Stop companion voice", group: "Companion", run: props.onStop }],
|
||||
}))
|
||||
|
||||
onMount(() => {
|
||||
props.input.focus()
|
||||
setTimeout(() => {
|
||||
if (textarea && !textarea.isDestroyed) textarea.focus()
|
||||
}, 1)
|
||||
})
|
||||
|
||||
return (
|
||||
<box flexGrow={1} minHeight={0} paddingLeft={2} paddingRight={2} paddingTop={1} gap={1}>
|
||||
<box flexDirection="row" gap={2} flexShrink={0}>
|
||||
<text attributes={TextAttributes.BOLD} fg={theme.text.base} flexGrow={1}>
|
||||
Companion
|
||||
</text>
|
||||
<VoiceStatus voice={props.voice} running={running()} />
|
||||
</box>
|
||||
<Show when={props.error}>
|
||||
<text fg={theme.text.feedback.error.base} wrapMode="word" flexShrink={0}>
|
||||
{props.error}
|
||||
</text>
|
||||
</Show>
|
||||
<scrollbox
|
||||
flexGrow={1}
|
||||
minHeight={0}
|
||||
stickyScroll
|
||||
stickyStart="bottom"
|
||||
scrollbarOptions={{ visible: false }}
|
||||
contentOptions={{ gap: 1 }}
|
||||
>
|
||||
<Show
|
||||
when={messages().length > 0}
|
||||
fallback={
|
||||
<text fg={theme.text.muted} wrapMode="word">
|
||||
Ask about the main session, or tell the companion what the main session should do next.
|
||||
</text>
|
||||
}
|
||||
>
|
||||
<For each={messages()}>{(message) => <Row message={message} background={background()} />}</For>
|
||||
</Show>
|
||||
</scrollbox>
|
||||
<box flexShrink={0} gap={1} paddingBottom={1}>
|
||||
<textarea
|
||||
minHeight={1}
|
||||
maxHeight={6}
|
||||
wrapMode="word"
|
||||
ref={(value: TextareaRenderable) => {
|
||||
textarea = value
|
||||
setTarget(value)
|
||||
}}
|
||||
placeholder="Talk to the companion"
|
||||
placeholderColor={theme.text.muted}
|
||||
textColor={theme.text.formfield.base}
|
||||
focusedTextColor={theme.text.formfield.base}
|
||||
cursorColor={theme.text.base}
|
||||
cursorStyle={config.cursor}
|
||||
/>
|
||||
<text fg={theme.text.muted} wrapMode="word">
|
||||
{[
|
||||
[shortcuts.get("companion.submit"), "send"],
|
||||
[shortcuts.get("companion.talk"), "talk"],
|
||||
[shortcuts.get("companion.stop"), "stop"],
|
||||
[shortcuts.get("session.companion"), "close"],
|
||||
]
|
||||
.filter((item): item is [string, string] => Boolean(item[0]))
|
||||
.map(([key, label]) => `${key} ${label}`)
|
||||
.join(" · ")}
|
||||
</text>
|
||||
</box>
|
||||
</box>
|
||||
)
|
||||
}
|
||||
|
||||
function VoiceStatus(props: { voice: Voice; running: boolean }) {
|
||||
const theme = useTheme()
|
||||
const color = () => theme.text.feedback.info.base
|
||||
return (
|
||||
<Switch>
|
||||
<Match when={props.voice.state.status === "listening"}>
|
||||
<text fg={color()} flexShrink={0}>
|
||||
● listening {LEVELS[Math.min(LEVELS.length - 1, Math.floor(props.voice.state.level * LEVELS.length))]}
|
||||
</text>
|
||||
</Match>
|
||||
<Match when={props.voice.state.status === "transcribing"}>
|
||||
<Spinner color={color()}>transcribing</Spinner>
|
||||
</Match>
|
||||
<Match when={props.voice.state.status === "speaking"}>
|
||||
<text fg={color()} flexShrink={0}>
|
||||
♪ speaking
|
||||
</text>
|
||||
</Match>
|
||||
<Match when={props.running}>
|
||||
<Spinner color={theme.text.muted}>thinking</Spinner>
|
||||
</Match>
|
||||
</Switch>
|
||||
)
|
||||
}
|
||||
|
||||
function Row(props: { message: SessionMessageInfo; background: RGBA }) {
|
||||
const theme = useTheme()
|
||||
return (
|
||||
<Switch>
|
||||
<Match when={props.message.type === "user" && props.message}>
|
||||
{(message) => (
|
||||
<box flexDirection="row" gap={1} flexShrink={0}>
|
||||
<text fg={theme.text.muted} flexShrink={0}>
|
||||
{message().metadata?.voice === true ? "◉" : "›"}
|
||||
</text>
|
||||
<text fg={theme.text.base} wrapMode="word" flexGrow={1}>
|
||||
{message().text}
|
||||
</text>
|
||||
</box>
|
||||
)}
|
||||
</Match>
|
||||
<Match when={props.message.type === "synthetic" && props.message}>
|
||||
{(message) => (
|
||||
<text fg={theme.text.muted} wrapMode="word" flexShrink={0}>
|
||||
{message().description ?? message().text}
|
||||
</text>
|
||||
)}
|
||||
</Match>
|
||||
<Match when={props.message.type === "assistant" && props.message}>
|
||||
{(message) => <Assistant message={message()} background={props.background} />}
|
||||
</Match>
|
||||
</Switch>
|
||||
)
|
||||
}
|
||||
|
||||
function Assistant(props: { message: SessionMessageAssistant; background: RGBA }) {
|
||||
const theme = useTheme()
|
||||
const syntax = useThemes().currentSyntax
|
||||
const plugins = usePlugin()
|
||||
return (
|
||||
<box flexShrink={0} gap={1}>
|
||||
<For each={props.message.content}>
|
||||
{(part) => (
|
||||
<Switch>
|
||||
<Match when={part.type === "text" && part.text.trim() ? part.text.trim() : undefined}>
|
||||
{(text) => (
|
||||
<markdown
|
||||
syntaxStyle={syntax()}
|
||||
renderNode={plugins.markdown()}
|
||||
content={text()}
|
||||
streaming={props.message.time.completed === undefined}
|
||||
internalBlockMode="top-level"
|
||||
tableOptions={{ style: "grid", cellPaddingX: 1 }}
|
||||
conceal
|
||||
fg={theme.markdown.text}
|
||||
bg={props.background}
|
||||
/>
|
||||
)}
|
||||
</Match>
|
||||
<Match when={part.type === "tool" && part}>
|
||||
{(tool) => (
|
||||
<text
|
||||
fg={tool().state.status === "error" ? theme.text.feedback.error.base : theme.text.muted}
|
||||
wrapMode="word"
|
||||
>
|
||||
{toolLabel(tool())}
|
||||
</text>
|
||||
)}
|
||||
</Match>
|
||||
</Switch>
|
||||
)}
|
||||
</For>
|
||||
<Show when={props.message.error}>
|
||||
{(error) => (
|
||||
<text fg={theme.text.feedback.error.base} wrapMode="word">
|
||||
{error().message}
|
||||
</text>
|
||||
)}
|
||||
</Show>
|
||||
</box>
|
||||
)
|
||||
}
|
||||
|
||||
function toolLabel(tool: SessionMessageAssistantTool) {
|
||||
const input = tool.state.status === "streaming" ? {} : tool.state.input
|
||||
const text = (key: string) => {
|
||||
const value = input[key]
|
||||
return typeof value === "string" ? value : undefined
|
||||
}
|
||||
const pending = tool.state.status === "streaming" || tool.state.status === "running" ? "…" : ""
|
||||
if (tool.name === "main_send") return `→ main (${text("delivery") ?? "steer"}): ${text("text") ?? ""}${pending}`
|
||||
if (tool.name === "main_status") return `· checked the main session${pending}`
|
||||
if (tool.name === "main_read") return `· read the main session${pending}`
|
||||
if (tool.name === "main_interrupt") return `■ interrupted the main session${pending}`
|
||||
if (tool.name === "main_cancel") return `× cancelled a queued prompt${pending}`
|
||||
return `· ${tool.name} ${text("command") ?? text("pattern") ?? text("path") ?? text("url") ?? text("query") ?? ""}${pending}`.trimEnd()
|
||||
}
|
||||
@@ -0,0 +1,245 @@
|
||||
import type { OpenCodeClient, VoiceSpeechFormat } from "@opencode/client"
|
||||
import { createStore } from "solid-js/store"
|
||||
import { encode, monitor, play, record, rms, type Recording } from "./audio"
|
||||
import { createBargeDetector, echoes, type Utterance } from "./barge"
|
||||
|
||||
export type VoiceStatus = "idle" | "listening" | "transcribing" | "speaking"
|
||||
|
||||
// Keeps the microphone open across the short gaps between sentences of one reply.
|
||||
const MONITOR_GRACE_MS = 1500
|
||||
|
||||
/**
|
||||
* Voice loop. By default it is half-duplex: the microphone is closed while
|
||||
* speech plays, so speaker output never reaches the transcript. With
|
||||
* `bargeIn`, the microphone stays open during playback. Talking over the
|
||||
* reply stops it at once. A transcript that is the reply's own echo resumes
|
||||
* it from the interrupted sentence; any other is reported with `interrupting`.
|
||||
*/
|
||||
export function createVoice(input: {
|
||||
readonly client: OpenCodeClient
|
||||
readonly bargeIn: () => boolean
|
||||
readonly onTranscript: (text: string, interrupting: boolean) => void
|
||||
readonly onError: (error: unknown) => void
|
||||
}) {
|
||||
const [state, setState] = createStore({ status: "idle" as VoiceStatus, level: 0 })
|
||||
let recording: Recording | undefined
|
||||
let speech = new AbortController()
|
||||
let playback = Promise.resolve()
|
||||
let queued = 0
|
||||
// Texts in the order their playback started, to recognise echo in a transcript.
|
||||
let played: string[] = []
|
||||
let turn = 0
|
||||
let listener: Promise<Awaited<ReturnType<typeof monitor>> | undefined> | undefined
|
||||
let closing = Promise.resolve()
|
||||
let grace: ReturnType<typeof setTimeout> | undefined
|
||||
let deciding = false
|
||||
// Texts that have not finished playing, and while the user talks over the reply, the ones held back.
|
||||
let unplayed: { text: string }[] = []
|
||||
let held: string[] | undefined
|
||||
|
||||
const detector = createBargeDetector({
|
||||
onOnset: () => {
|
||||
held = unplayed.map((entry) => entry.text)
|
||||
unplayed = []
|
||||
speech.abort()
|
||||
speech = new AbortController()
|
||||
setState({ status: "listening", level: 0 })
|
||||
},
|
||||
onUtterance: (utterance) => void decide(utterance),
|
||||
})
|
||||
|
||||
const transcript = (audio: Uint8Array) =>
|
||||
input.client.voice
|
||||
.transcribe({ mediaType: "audio/wav", payload: audio })
|
||||
.then((result) => result.text.trim())
|
||||
.catch((error) => {
|
||||
input.onError(error)
|
||||
return ""
|
||||
})
|
||||
|
||||
const decide = async (utterance: Utterance) => {
|
||||
const current = turn
|
||||
deciding = true
|
||||
setState({ status: "transcribing", level: 0 })
|
||||
const text = await transcript(encode(utterance.samples, utterance.sampleRate))
|
||||
if (current !== turn) return
|
||||
deciding = false
|
||||
if (text && !echoes(text, played.join(" "))) {
|
||||
stop()
|
||||
input.onTranscript(text, true)
|
||||
return
|
||||
}
|
||||
if (text) detector.rejected()
|
||||
const resumed = held ?? []
|
||||
held = undefined
|
||||
setState("status", "idle")
|
||||
resumed.forEach(speak)
|
||||
if (resumed.length === 0) close()
|
||||
}
|
||||
|
||||
const open = () => {
|
||||
clearTimeout(grace)
|
||||
if (listener || !input.bargeIn()) return
|
||||
detector.reset()
|
||||
listener = monitor((samples, sampleRate, level) => {
|
||||
if (deciding) return
|
||||
detector.push(samples, sampleRate, level)
|
||||
if (detector.active) setState("level", rms(samples))
|
||||
}).catch((error) => {
|
||||
input.onError(error)
|
||||
return undefined
|
||||
})
|
||||
}
|
||||
|
||||
const close = () => {
|
||||
clearTimeout(grace)
|
||||
const current = listener
|
||||
listener = undefined
|
||||
detector.reset()
|
||||
if (current) closing = current.then((opened) => opened?.close()).catch(() => undefined)
|
||||
}
|
||||
|
||||
const listen = async () => {
|
||||
stop()
|
||||
const current = turn
|
||||
setState({ status: "listening", level: 0 })
|
||||
// OpenTUI allows one capture at a time.
|
||||
await closing
|
||||
if (current !== turn) return
|
||||
recording = await record((level) => setState("level", level)).catch((error) => {
|
||||
setState("status", "idle")
|
||||
input.onError(error)
|
||||
return undefined
|
||||
})
|
||||
}
|
||||
|
||||
const transcribe = async () => {
|
||||
const active = recording
|
||||
recording = undefined
|
||||
if (!active) return
|
||||
const current = turn
|
||||
setState({ status: "transcribing", level: 0 })
|
||||
const audio = await active.stop()
|
||||
const text = audio ? await transcript(audio) : ""
|
||||
// Stopping while the utterance transcribes discards it.
|
||||
if (current !== turn) return
|
||||
setState("status", "idle")
|
||||
if (text) input.onTranscript(text, false)
|
||||
}
|
||||
|
||||
const stop = () => {
|
||||
turn++
|
||||
speech.abort()
|
||||
speech = new AbortController()
|
||||
recording?.cancel()
|
||||
recording = undefined
|
||||
close()
|
||||
deciding = false
|
||||
played = []
|
||||
unplayed = []
|
||||
held = undefined
|
||||
setState({ status: "idle", level: 0 })
|
||||
}
|
||||
|
||||
const speak = (text: string) => {
|
||||
if (held) return void held.push(text)
|
||||
const controller = speech
|
||||
const signal = controller.signal
|
||||
const clip = fetchClip(input.client, text, signal)
|
||||
clip.catch(() => undefined)
|
||||
const entry = { text }
|
||||
unplayed.push(entry)
|
||||
queued++
|
||||
if (state.status === "idle") setState("status", "speaking")
|
||||
open()
|
||||
playback = playback
|
||||
.then(async () => {
|
||||
if (signal.aborted) return
|
||||
const loaded = await clip
|
||||
played = [...played.slice(-1), text]
|
||||
await play(loaded.format, loaded.body, signal)
|
||||
})
|
||||
.catch((error) => {
|
||||
if (signal.aborted) return
|
||||
// One failure drops the rest of the reply instead of reporting every sentence.
|
||||
controller.abort()
|
||||
input.onError(error)
|
||||
})
|
||||
.finally(() => {
|
||||
unplayed = unplayed.filter((item) => item !== entry)
|
||||
queued--
|
||||
if (queued > 0) return
|
||||
if (state.status === "speaking") setState("status", "idle")
|
||||
if (detector.active || deciding) return
|
||||
clearTimeout(grace)
|
||||
grace = setTimeout(close, MONITOR_GRACE_MS)
|
||||
})
|
||||
}
|
||||
|
||||
return {
|
||||
state,
|
||||
/** Starts listening, or ends the utterance and transcribes it. */
|
||||
toggle() {
|
||||
if (state.status === "transcribing") return
|
||||
if (detector.active) return detector.flush()
|
||||
if (state.status === "listening") return void transcribe()
|
||||
void listen()
|
||||
},
|
||||
/** Stops listening and speaking. */
|
||||
stop,
|
||||
/** Queues one chunk of text; its audio request starts now so playback has no gap. */
|
||||
speak,
|
||||
}
|
||||
}
|
||||
|
||||
export type Voice = ReturnType<typeof createVoice>
|
||||
|
||||
async function fetchClip(client: OpenCodeClient, text: string, signal: AbortSignal) {
|
||||
const events = client.voice.speech({ text }, { signal })[Symbol.asyncIterator]()
|
||||
const first = await events.next()
|
||||
if (first.done) throw new Error("Speech ended before any audio")
|
||||
if (first.value.type === "error") throw new Error(first.value.message)
|
||||
if (first.value.type !== "format") throw new Error("Speech started without a format")
|
||||
const format: VoiceSpeechFormat = first.value.format
|
||||
async function* body() {
|
||||
while (true) {
|
||||
const next = await events.next()
|
||||
if (next.done || next.value.type === "done") return
|
||||
if (next.value.type === "error") throw new Error(next.value.message)
|
||||
if (next.value.type === "audio") yield new Uint8Array(Buffer.from(next.value.data, "base64"))
|
||||
}
|
||||
}
|
||||
return { format, body: body() }
|
||||
}
|
||||
|
||||
/** Text worth reading aloud: code blocks, including one still streaming, and markdown syntax are dropped. */
|
||||
export function speakable(text: string) {
|
||||
const closed = text.replace(/```[\s\S]*?```/g, " ")
|
||||
const open = closed.indexOf("```")
|
||||
return (open === -1 ? closed : closed.slice(0, open))
|
||||
.replace(/`([^`]*)`/g, "$1")
|
||||
.replace(/\[([^\]]*)\]\([^)]*\)/g, "$1")
|
||||
.replace(/^\s*[-*+]\s+/gm, "")
|
||||
.replace(/[*_#>|]/g, "")
|
||||
.replace(/\s+/g, " ")
|
||||
}
|
||||
|
||||
/**
|
||||
* Splits `text` after `from` into complete sentences. Short sentences merge
|
||||
* with the next one so each speech request carries enough text to sound
|
||||
* natural. When `final`, the remainder is flushed too.
|
||||
*/
|
||||
export function sentences(text: string, from: number, final: boolean) {
|
||||
const boundaries = Array.from(text.slice(from).matchAll(/[.!?…:;]+["')\]]*(?=\s)/g), (match) => {
|
||||
return from + (match.index ?? 0) + match[0].length
|
||||
})
|
||||
const cuts = boundaries.reduce<number[]>((result, boundary) => {
|
||||
const start = result.at(-1) ?? from
|
||||
return text.slice(start, boundary).trim().length >= 40 ? [...result, boundary] : result
|
||||
}, [])
|
||||
const ends = final && text.slice(cuts.at(-1) ?? from).trim() ? [...cuts, text.length] : cuts
|
||||
return {
|
||||
chunks: ends.map((end, index) => text.slice(ends[index - 1] ?? from, end).trim()).filter(Boolean),
|
||||
end: ends.at(-1) ?? from,
|
||||
}
|
||||
}
|
||||
@@ -38,6 +38,8 @@ export default Plugin.define({
|
||||
return
|
||||
}
|
||||
const session = context.data.session.get(sessionID)
|
||||
// The companion panel shows its replies as they stream.
|
||||
if (session?.kind === "companion") return
|
||||
notify(context, sessionID, "Session done", session?.parentID ? "subagent_done" : "done")
|
||||
}
|
||||
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import HomeFooter from "../feature-plugins/home/footer"
|
||||
import PromptBtw from "../feature-plugins/prompt/btw"
|
||||
import PromptFooter from "../feature-plugins/prompt/footer"
|
||||
import SessionCompanion from "../feature-plugins/session/companion"
|
||||
import SidebarContext from "../feature-plugins/sidebar/context"
|
||||
import SidebarFooter from "../feature-plugins/sidebar/footer"
|
||||
import SidebarMcp from "../feature-plugins/sidebar/mcp"
|
||||
@@ -16,6 +17,7 @@ export const builtins = [
|
||||
HomeFooter,
|
||||
PromptFooter,
|
||||
PromptBtw,
|
||||
SessionCompanion,
|
||||
SidebarContext,
|
||||
SidebarMcp,
|
||||
SidebarFooter,
|
||||
|
||||
@@ -37,24 +37,23 @@ export function SubagentsTab(props: { sessionID: string }) {
|
||||
const current = session()
|
||||
if (!current) return []
|
||||
|
||||
const result = sessionFamily<SessionInfo>(data.session.list(), current.id).map(
|
||||
({ session, prefix }): SubagentEntry => {
|
||||
const title = withTimestampedFallback(session)
|
||||
const agentMatch = title.match(/@(\w+) subagent/)
|
||||
return {
|
||||
sessionID: session.id,
|
||||
agent: session.agent
|
||||
? Locale.titlecase(session.agent)
|
||||
: agentMatch
|
||||
? Locale.titlecase(agentMatch[1])
|
||||
: "Subagent",
|
||||
title: agentMatch ? title.replace(agentMatch[0], "").trim() || title : title,
|
||||
status: data.session.status(session.id),
|
||||
current: session.id === route.sessionID,
|
||||
prefix,
|
||||
}
|
||||
},
|
||||
)
|
||||
const subagents = data.session.list().filter((session) => session.kind !== "companion")
|
||||
const result = sessionFamily<SessionInfo>(subagents, current.id).map(({ session, prefix }): SubagentEntry => {
|
||||
const title = withTimestampedFallback(session)
|
||||
const agentMatch = title.match(/@(\w+) subagent/)
|
||||
return {
|
||||
sessionID: session.id,
|
||||
agent: session.agent
|
||||
? Locale.titlecase(session.agent)
|
||||
: agentMatch
|
||||
? Locale.titlecase(agentMatch[1])
|
||||
: "Subagent",
|
||||
title: agentMatch ? title.replace(agentMatch[0], "").trim() || title : title,
|
||||
status: data.session.status(session.id),
|
||||
current: session.id === route.sessionID,
|
||||
prefix,
|
||||
}
|
||||
})
|
||||
|
||||
return result.filter((entry) => (store.active ? entry.status === "running" : entry.status !== "running"))
|
||||
})
|
||||
|
||||
@@ -190,7 +190,9 @@ export function Session(props: {
|
||||
onCleanup(() => setEpilogue())
|
||||
const descendantSessionIDs = createMemo(() => {
|
||||
if (session()?.parentID) return []
|
||||
return data.session.family(route.sessionID).filter((id) => id !== route.sessionID)
|
||||
return data.session
|
||||
.family(route.sessionID)
|
||||
.filter((id) => id !== route.sessionID && data.session.get(id)?.kind !== "companion")
|
||||
})
|
||||
const permissions = createMemo(() => {
|
||||
if (session()?.parentID) return []
|
||||
@@ -2251,6 +2253,9 @@ function UserMessage(props: { message: SessionMessageUser }) {
|
||||
flexShrink={0}
|
||||
>
|
||||
<text fg={theme.text.base}>{props.message.text}</text>
|
||||
<Show when={props.message.metadata?.source === "companion"}>
|
||||
<text fg={theme.text.muted}>via companion</text>
|
||||
</Show>
|
||||
<Show when={skills().length}>
|
||||
<box flexDirection="row" paddingTop={1} gap={1} flexWrap="wrap">
|
||||
<For each={skills()}>
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
import { describe, expect, test } from "bun:test"
|
||||
import { createBargeDetector, echoes, type Utterance } from "../../src/feature-plugins/session/companion/barge"
|
||||
import { sentences, speakable } from "../../src/feature-plugins/session/companion/voice"
|
||||
|
||||
const RATE = 48000
|
||||
const FRAMES = 1024
|
||||
|
||||
function detector() {
|
||||
const events: string[] = []
|
||||
const utterances: Utterance[] = []
|
||||
const barge = createBargeDetector({
|
||||
onOnset: () => events.push("onset"),
|
||||
onUtterance: (utterance) => {
|
||||
events.push("utterance")
|
||||
utterances.push(utterance)
|
||||
},
|
||||
})
|
||||
// A constant chunk's RMS is its level.
|
||||
const push = (count: number, mic: number, playback: number) =>
|
||||
Array.from({ length: count }).forEach(() => barge.push(new Float32Array(FRAMES).fill(mic), RATE, playback))
|
||||
return { barge, events, utterances, push }
|
||||
}
|
||||
|
||||
describe("companion barge-in", () => {
|
||||
test("ignores echo at the level playback explains", () => {
|
||||
const run = detector()
|
||||
run.push(200, 0.08, 0.2)
|
||||
expect(run.events).toEqual([])
|
||||
})
|
||||
|
||||
test("detects speech over playback and keeps the audio before its onset", () => {
|
||||
const run = detector()
|
||||
run.push(40, 0.05, 0.2)
|
||||
run.push(7, 0.3, 0.2)
|
||||
expect(run.events).toEqual([])
|
||||
run.push(13, 0.3, 0.2)
|
||||
expect(run.events).toEqual(["onset"])
|
||||
// Speech ends once the microphone hears only the turned-down echo for long enough.
|
||||
run.push(32, 0.012, 0.05)
|
||||
expect(run.events).toEqual(["onset"])
|
||||
run.push(2, 0.012, 0.05)
|
||||
expect(run.events).toEqual(["onset", "utterance"])
|
||||
// 19 chunks of preroll (the first 400 ms, ending at the onset), 12 more of speech, 33 of silence.
|
||||
expect(run.utterances[0]?.samples.length).toBe((19 + 12 + 33) * FRAMES)
|
||||
expect(run.barge.active).toBe(false)
|
||||
})
|
||||
|
||||
test("stops treating rejected echo as speech", () => {
|
||||
const run = detector()
|
||||
run.push(20, 0.15, 0.1)
|
||||
run.barge.flush()
|
||||
expect(run.events).toEqual(["onset", "utterance"])
|
||||
run.barge.rejected()
|
||||
run.push(200, 0.15, 0.1)
|
||||
expect(run.events).toEqual(["onset", "utterance"])
|
||||
})
|
||||
|
||||
test("recognises a transcript of the reply being spoken", () => {
|
||||
const spoken = "The main session is running the test suite right now."
|
||||
expect(echoes("the main session is running the tests", spoken)).toBe(true)
|
||||
expect(echoes("Stop, tell it to use bun instead.", spoken)).toBe(false)
|
||||
expect(echoes("", spoken)).toBe(false)
|
||||
})
|
||||
})
|
||||
|
||||
describe("companion speech text", () => {
|
||||
test("drops code blocks and markdown syntax", () => {
|
||||
expect(speakable("Run `bun test` in **core**.\n\n```ts\nconst a = 1\n```\n- See [docs](https://x).")).toBe(
|
||||
"Run bun test in core. See docs.",
|
||||
)
|
||||
})
|
||||
|
||||
test("stops before a code block that is still streaming", () => {
|
||||
expect(speakable("Here is the fix:\n```ts\nconst a")).toBe("Here is the fix: ")
|
||||
})
|
||||
|
||||
test("cuts only complete sentences until the reply finishes", () => {
|
||||
const text = "The main session is running the test suite right now. It already fixed the parser and"
|
||||
const partial = sentences(text, 0, false)
|
||||
expect(partial.chunks).toEqual(["The main session is running the test suite right now."])
|
||||
const final = sentences(text, partial.end, true)
|
||||
expect(final.chunks).toEqual(["It already fixed the parser and"])
|
||||
expect(final.end).toBe(text.length)
|
||||
})
|
||||
|
||||
test("merges short sentences so each request sounds natural", () => {
|
||||
expect(sentences("Sure. I sent that to the main session as a steer. Done.", 0, true).chunks).toEqual([
|
||||
"Sure. I sent that to the main session as a steer.",
|
||||
"Done.",
|
||||
])
|
||||
})
|
||||
|
||||
test("does not split decimals", () => {
|
||||
expect(sentences("Coverage went from 81.5 to 84.2 percent after the change", 0, false).chunks).toEqual([])
|
||||
})
|
||||
})
|
||||
+5
-4
@@ -19,10 +19,11 @@ Generated clients follow the assembled public `HttpApi`. GitHub issues own activ
|
||||
|
||||
## Current Contracts
|
||||
|
||||
| Document | Job |
|
||||
| ----------------------- | --------------------------------------------------------------------------------------- |
|
||||
| [Session](./session.md) | Explain prompt admission, execution, instructions, compaction, and recovery boundaries. |
|
||||
| [Tools](./tools.md) | Explain tool construction, registration, execution, and outcome laws. |
|
||||
| Document | Job |
|
||||
| --------------------------- | --------------------------------------------------------------------------------------- |
|
||||
| [Session](./session.md) | Explain prompt admission, execution, instructions, compaction, and recovery boundaries. |
|
||||
| [Tools](./tools.md) | Explain tool construction, registration, execution, and outcome laws. |
|
||||
| [Companion](./companion.md) | Explain companion sessions, main-session tools, and the voice cascade. |
|
||||
|
||||
## Decision Records
|
||||
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
# V2 Companion
|
||||
|
||||
Status: **Experimental.** The TUI gates the feature behind the `companion` experiment. Protocol routes live under `/api/experimental/`.
|
||||
|
||||
A companion is a persistent side conversation about one session, called the main session. The user talks to the companion in text or by voice. The companion reads what the main session is doing and can steer it. `/btw` stays a one-shot, tool-less question; a companion is a full session.
|
||||
|
||||
## Companion Sessions
|
||||
|
||||
A companion is a child session of the main session with `kind: "companion"`. The kind is an irreducible fact recorded on `session.created`; it cannot change later. The companion runs through the normal Session runner, so history, tools, compaction, interruption, and persistence behave as they do for every other session.
|
||||
|
||||
Each main session has at most one companion. `Session.companion(sessionID)` returns the existing companion or creates one. It creates the companion with:
|
||||
|
||||
- the built-in hidden `companion` agent,
|
||||
- the `companion` agent's configured model, or the main session's model when none is configured,
|
||||
- no inherited session permissions, so main-session approvals do not widen what the companion may do.
|
||||
|
||||
The model is chosen once, when the companion is created. To give companions a faster model than the main session, configure the agent like any other:
|
||||
|
||||
```json
|
||||
{
|
||||
"agents": {
|
||||
"companion": { "model": "google/gemini-3.8-flash" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Rules that depend on the kind:
|
||||
|
||||
- The `subagent` tool refuses to continue a companion session.
|
||||
- Clients do not treat companions as subagents. They leave them out of subagent pickers, subagent counts, and completion notifications.
|
||||
- Removing the main session removes its companion like any other child.
|
||||
|
||||
## Agent and Tools
|
||||
|
||||
The `companion` agent is `primary` and `hidden`. It denies every action, then allows read-only file access (`read`, `grep`, `glob`), web lookups, read-only git commands, and the `main_*` tools. The companion does not edit files or run other shell commands; that would make two agents write to one worktree.
|
||||
|
||||
The companion may run these shell commands:
|
||||
|
||||
- `git status`, `git diff`, `git log`, and `git show`, with any arguments;
|
||||
- `git branch` alone, or with exactly one of `--show-current`, `-a`, `-r`, `-v`, or `-vv`. Other forms are denied because `git branch <name>` creates a branch, even after `-v`.
|
||||
|
||||
Git commands that contain `>` or `--output` are denied, because both write files.
|
||||
|
||||
Configured permission rules apply after the companion's own rules, so they can allow more within the capped actions, for example an external directory or `.env` reads. Rules such as `git *` or `edit` have no effect, because two hooks cap the companion regardless of configuration:
|
||||
|
||||
- A tool hook runs a companion shell command only when it is exactly one of the git commands above and contains no `;`, `&`, `|`, `<`, `>`, `$`, backtick, newline, or `--output`. This also covers redirects after `&&` or `||`, which the shell parser leaves out of permission resources.
|
||||
- A permission hook denies every action except `read`, `grep`, `glob`, `webfetch`, `websearch`, `shell`, `external_directory`, and the `main_*` tools. It also turns every `ask` result into `deny`: the companion never asks for permission, because clients do not show permission prompts for companions.
|
||||
|
||||
The `main_*` tools act only on the calling companion's parent. They never accept a session ID from the model, and they fail when the caller is not a companion. Other agents do not see them.
|
||||
|
||||
| Tool | Effect |
|
||||
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `main_status` | Report whether the main session is running, its pending inbox items, and its pending permission requests. |
|
||||
| `main_read` | Return recent main-session messages as compact text, including the in-flight step. |
|
||||
| `main_send` | Admit a prompt into the main session. `steer` delivers at the next safe boundary; `queue` waits until the main session is idle. The prompt carries `metadata.source = "companion"`. |
|
||||
| `main_cancel` | Cancel a pending main-session inbox item. |
|
||||
| `main_interrupt` | Interrupt the main session, optionally resuming pending steers. |
|
||||
|
||||
Steers are sent without confirmation. They appear in the main timeline like any prompt, so the user can see what the companion sent. The TUI marks user messages whose `metadata.source` is `"companion"` with a muted `via companion` label, so they do not look like the user typed them.
|
||||
|
||||
Before each companion request whose tail is a user message, a context hook inserts a short, unpersisted digest of the main session before the newest prompt: the `main_status` report and the six most recent main-session messages in `main_read` form. Most questions then need no tool call, which keeps voice replies fast.
|
||||
|
||||
## Voice
|
||||
|
||||
Voice is two stateless server operations backed by `@opencode/ai` speech and transcription models. The server owns provider credentials; clients own capture, playback, and turn-taking.
|
||||
|
||||
| Operation | Contract |
|
||||
| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Transcribe | The client uploads one recorded utterance as the binary request body with its media type. The server returns the transcript text. |
|
||||
| Speak | The client sends text. The server streams server-sent events: one `format` event that describes the audio encoding, then base64 `audio` chunks, then `done` or `error`. |
|
||||
|
||||
The `voice` configuration comes from the base (global) configuration, like one-shot generation. It names the models as `provider/model`, because the model catalog has no speech or transcription entries:
|
||||
|
||||
```json
|
||||
{
|
||||
"voice": {
|
||||
"transcription": { "model": "xai/grok-voice-transcribe-2.0" },
|
||||
"speech": { "model": "google/gemini-3.8-flash-lite-tts", "voice": "Kore" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The server resolves the provider's stored API key the same way it does for language models, then builds the provider package's media model directly. Custom providers work when their `package` is a supported provider package. OAuth logins are ignored because they target chat backends; those providers fall back to their API-key environment variable. Supported packages are OpenAI, Google, xAI, ElevenLabs, and Deepgram for speech, and OpenAI, Google, xAI, Deepgram, and AssemblyAI for transcription. Speech requests MP3, except Gemini, which only returns 24 kHz PCM.
|
||||
|
||||
## TUI
|
||||
|
||||
The TUI builtin companion plugin opens the companion in the `session.panel` slot for the current session with `/companion` or `session.companion`. The panel shows the companion transcript and a text input. The plugin keeps voice state, audio playback, and the main-to-companion mapping outside the panel component, so speech continues when the panel closes.
|
||||
|
||||
The voice loop is a cascade:
|
||||
|
||||
1. The talk key (`companion.talk`) starts microphone capture. Pressing it again stops capture.
|
||||
2. The client downsamples the capture to 16 kHz mono PCM, wraps it as WAV, and transcribes it.
|
||||
3. The transcript is sent to the companion as a prompt.
|
||||
4. As companion text streams in, the client splits it into sentences outside code blocks and speaks them in order.
|
||||
5. Pressing talk while the companion is thinking or speaking stops playback, interrupts the companion, and starts listening again.
|
||||
|
||||
The stop key (`companion.stop`) stops playback, capture, and the companion's reply. Stopping, or sending a typed message, also discards an utterance that is still being transcribed.
|
||||
|
||||
By default the microphone is closed while the companion speaks, so speaker output does not feed back into the transcript.
|
||||
|
||||
### Barge-in
|
||||
|
||||
The `companion_barge_in` experiment keeps the microphone open while the companion speaks, and for a short grace period after each clip, so the user can talk over a reply. There is no echo cancellation. Instead:
|
||||
|
||||
1. For each microphone chunk, the client reads the last third of a second of mixed playback from the OpenTUI tap. The loudest block of that window is the playback level.
|
||||
2. A chunk counts as speech when it is louder than both a noise floor and the playback level times a learned echo coupling, with a margin. The coupling follows the echo the microphone hears between words.
|
||||
3. After 150 ms of speech, the client stops playback at once and holds back the unfinished sentences and any that stream in later. It records the utterance, including the 400 ms before the onset, until 700 ms without speech. Pressing talk ends the utterance early.
|
||||
4. The client transcribes the utterance. When most of its words repeat the two clips spoken last, the utterance is echo: the coupling rises, and the reply resumes from the start of the interrupted sentence. An empty transcript also resumes the reply.
|
||||
5. Otherwise the client drops the held sentences, interrupts the companion, waits for the interrupt, and then sends the transcript as the next voice prompt. A prompt sent before the interrupt would wait in the stopped reply's inbox.
|
||||
|
||||
This works well with headphones and reasonably with laptop speakers. Loud speakers can still cause false interruptions until there is real echo cancellation.
|
||||
Reference in New Issue
Block a user