Skip to main content
Documentation

AI reference

Provider routing, context compaction, chat-history wrappers, and built-in tools.

On this page
import {
  // Provider
  createDeepSpaceAI, streamDeepSpaceAgent,
  listDeepSpaceAgentModels, resolveDeepSpaceAgentModel,
  // Compaction
  prepareMessagesWithCompaction, truncateOldToolResults, applySlidingWindow,
  capToolResultSize, totalChars,
  turnsToCoreMessages, buildUiParts, unwrapToolOutput,
  makeDefaultSummarizer,
  DEFAULT_CONTEXT_CONFIG,
  // Chat history (DO tools API wrappers)
  getChat, createChat, updateChat, deleteChatCascade,
  loadMessages, appendMessage,
  // Schemas
  AI_CHATS_SCHEMA, AI_MESSAGES_SCHEMA,
  // Built-in tools
  BUILT_IN_TOOLS, applyAiToolDefaults, DEFAULT_QUERY_LIMIT,
} from 'deepspace/worker'

import type {
  DeepSpaceAIEnv, DeepSpaceAIOptions, DeepSpaceModelFactory,
  ChatContextConfig, ChatTurn, Summarizer,
  ChatRow, ChatMessageRow,
  ToolSchema,
} from 'deepspace/worker'
ts

createDeepSpaceAI(env, provider, options?)#

Returns a Vercel AI SDK 7 model factory routed through the DeepSpace proxy.

function createDeepSpaceAI(
  env: DeepSpaceAIEnv,
  provider: 'anthropic' | 'openai' | 'cerebras',
  options?: DeepSpaceAIOptions,
): DeepSpaceModelFactory

type DeepSpaceModelFactory = (modelId: string) => LanguageModel

interface DeepSpaceAIOptions {
  authToken?: string
  sandboxScope?: 'user' | 'app'
}
ts
OptionEffect
authToken (passed)Caller pays - JWT subject is billed
authToken (omitted)Owner pays - falls back to env.APP_OWNER_JWT
sandboxScopeAccess to new sandbox files and containers. Defaults to 'user'; 'app' shares within the app. See sandbox scope.

Use the returned factory with streamText / generateText from the ai package:

import { streamText } from 'ai'
const ai = createDeepSpaceAI(env, 'anthropic', { authToken })
const result = streamText({
  model: ai('claude-opus-5-5'),
  messages,
  tools,
})
ts

Output limits#

For Anthropic models created by createDeepSpaceAI, an omitted maxOutputTokens defaults to 64,000. This is an output allowance for each model request, not a job timeout or a limit on the input document's size.

Set maxOutputTokens on streamText or generateText to override the SDK default:

const result = streamText({
  model: ai('claude-sonnet-5'),
  maxOutputTokens: 96_000,
  messages,
  tools,
})
ts

Choose a value supported by the model; the provider caps it at the model's output ceiling. A higher allowance increases the credits reserved before the request, while billing uses actual usage.

For Anthropic, finishReason: 'length' can mean the output limit or context window was reached. Check rawFinishReason: 'max_tokens' identifies the output limit; 'model_context_window_exceeded' identifies the context window. Increasing maxOutputTokens does not fix a full context window.

Adaptive thinking shares the limit with the response. With Anthropic thinking.type: 'enabled', the provider adds budgetTokens to the output allowance, capped at the model's ceiling. This default applies to Anthropic only; it does not change OpenAI or Cerebras limits.

Agent runner#

streamDeepSpaceAgent(env, options) wraps streamText with the SDK's model catalog, provider routing, and profile limits:

const { result } = streamDeepSpaceAgent(env, {
  profile: 'application',
  authToken,
  instructions: 'Help the user work with their records.',
  messages,
  tools,
  abortSignal: signal,
})
ts

Omitting modelId selects the profile default. An unsupported model throws DeepSpaceAgentModelError. Use listDeepSpaceAgentModels('application') for picker options and resolveDeepSpaceAgentModel(modelId, 'application') to validate a selection before starting a stream. See the model table.

The application profile accepts app-defined tools and limits a turn to 20 tool executions and 21 model steps. streamDeepSpaceAgent owns model, providerOptions, and stopWhen. For a separate sandbox task requiring its own provider options or loop limits, use createDeepSpaceAI with streamText directly.

Context compaction#

prepareMessagesWithCompaction(messages, config, options)#

Pre-stream pipeline that keeps the conversation under the context budget.

function prepareMessagesWithCompaction(
  messages: ChatTurn[],
  config: ChatContextConfig,
  options: {
    summarizer: Summarizer
    cachedSummary?: { text: string; throughId: string }
  },
): Promise<{
  messages: ChatTurn[]
  newSummary?: { text: string; throughId: string }
}>
ts

cachedSummary is the previous turn's summary (if any), anchored to a known message id. When the helper produces a fresh summary, it returns newSummary for persistence - store it on the chat row so the next turn can pass it back as cachedSummary.

Order of operations:

  1. truncateOldToolResults - replace old tool-result payloads with a small marker.
  2. Apply cachedSummary if its throughId is found in the history.
  3. Summarize the older half if still over budget; return as newSummary.
  4. Fall back to applySlidingWindow on summarizer error or missing message ids.

truncateOldToolResults(messages, keepRecent)#

Replaces old tool-result payloads with markers; preserves errors (success: false) and the keepRecent most recent assistant turns intact.

applySlidingWindow(messages, charCap, minKept)#

Drops oldest messages until under charCap, never below minKept. System messages are pinned.

capToolResultSize(result, byteCap)#

Caps individual tool-result payloads with a structured "result too large; narrow your query" error. Preserves a 2KB preview.

totalChars(messages)#

Sum of content + JSON.stringify(parts) lengths.

DEFAULT_CONTEXT_CONFIG#

const DEFAULT_CONTEXT_CONFIG: ChatContextConfig = {
  contextBudget: 240_000,       // chars ≈ 60–80K tokens
  toolResultCap: 30_000,        // bytes per tool result
  keepRecentToolResults: 5,
  minKept: 10,                  // sliding-window floor
}
ts

Sized for 200K+ context models. Lower for shorter-context models.

Format conversions#

function turnsToCoreMessages(turns: ChatTurn[]): ModelMessage[]
function buildUiParts(responseMessages: ModelMessage[]): unknown[]
function unwrapToolOutput(output: unknown): unknown
ts

turnsToCoreMessages converts persisted UI-shape ChatTurn rows into Vercel AI SDK 7 ModelMessages, splitting assistant rows at each tool-call boundary so Anthropic's tool_use → tool_result pairing is preserved.

buildUiParts is the inverse - converts onEnd's responseMessages from all steps into the flat UI-shape parts array we persist on ai-messages rows. response.messages contains only the final step.

unwrapToolOutput unwraps the AI SDK's tagged output ({ type: 'json' | 'text' | ..., value }) into the flat shape we persist.

Summarizers#

makeDefaultSummarizer(env, options?)#

Returns a Claude Haiku 4.5 summarizer.

function makeDefaultSummarizer(
  env: DeepSpaceAIEnv,
  options?: { authToken?: string },
): Summarizer
ts

Omit authToken to bill the owner (compaction as infrastructure cost). Pass the caller's JWT to bill the user (compaction as part of chat cost).

The default summary anchors on the last real message ID in the older half (skipping prior-summary system rows so re-summarization doesn't loop) - preserve that anchoring if you replace it.

Summarizer type#

type Summarizer = (messages: ChatTurn[]) => Promise<string>
ts

Roll your own implementation if you want a different model or strategy.

Chat history helpers (DO tools API wrappers)#

These read and write the ai-chats and ai-messages collections with X-App-Action: 'true' (bypassing user RBAC). The worker is the trust boundary - callers MUST verify chat ownership before invoking write helpers.

function getChat(
  stub: DurableObjectStub,
  chatId: string,
  userId: string,
): Promise<ChatRow | null>

function createChat(
  stub: DurableObjectStub,
  userId: string,
  opts?: { title?: string; model?: string },
): Promise<ChatRow>

function updateChat(
  stub: DurableObjectStub,
  chatId: string,
  userId: string,
  patch: Partial<Pick<ChatRow, 'title' | 'model' | 'compactedSummary' | 'compactedThroughId'>>,
): Promise<void>

function deleteChatCascade(
  stub: DurableObjectStub,
  chatId: string,
  userId: string,
): Promise<void>

function loadMessages(
  stub: DurableObjectStub,
  chatId: string,
  userId: string,
): Promise<ChatMessageRow[]>

function appendMessage(
  stub: DurableObjectStub,
  msg: {
    id: string
    chatId: string
    userId: string
    role: 'user' | 'assistant' | 'system'
    content: string
    parts?: unknown[]
  },
): Promise<void>
ts

appendMessage takes an id field that becomes the new row's recordId on the underlying tools API.

TypeShape
ChatRow{ recordId, id, userId, title, model?, compactedSummary?, compactedThroughId?, createdAt, updatedAt }
ChatMessageRow{ recordId, id, chatId, userId, role, content, parts?, createdAt }

Built-in tools#

const BUILT_IN_TOOLS: ToolSchema[]

interface ToolSchema {
  name: string
  description: string
  params: Record<string, {
    type: 'string' | 'number' | 'boolean' | 'object' | 'array'
    description: string
    required?: boolean
    default?: unknown
  }>
}
ts

BUILT_IN_TOOLS is an array of tool schemas, not a record keyed by name. Each entry declares its parameters as a flat { type, description, required?, default? } map - this is an MCP-like description used by the worker's tools API and by app authors who want to surface SDK tools to an LLM.

The catalog (records, schemas, users, storage, backup, Yjs):

ToolPurpose
records.queryFilter and list records
records.getFetch one record
records.createCreate a record
records.updatePatch a record
records.deleteDelete a record
records.deleteWhereDelete every record matching a filter, one bounded page per call ({ collection, where, limit } → { deleted }; repeat until deleted < limit); same delete permission check as records.delete, and a where key that names no field is refused, not ignored
schema.listEnumerate collection names
schema.describeDescribe one collection's columns and permissions
user.currentLook up the caller's user record
user.listList the users in the room, projected to what the caller may see - full rows for an admin caller, the public-identity projection for everyone else. See the directory from the server side
storage.list / read / write / deleteKey-value storage
backup.create / list / restore / deleteYjs doc backups
yjs.list / getText / setTextCollaborative doc text access

See src/ai/tools.ts in the scaffold for buildSystemPrompt(appName, schemas) and buildReadOnlyTools(executor) - both are app-local references you can edit to customize the assistant's tool surface and system prompt.

applyAiToolDefaults(toolName, params)#

Fills in assistant-only parameter defaults for a built-in tool call, before a model-issued call is dispatched to the tools API.

function applyAiToolDefaults(
  toolName: string,
  params: Record<string, unknown>,
): Record<string, unknown>

const DEFAULT_QUERY_LIMIT = 50
ts

One default: a records.query call with no limit gets limit: DEFAULT_QUERY_LIMIT, so a model-issued unbounded scan can't blow the tool-result byte cap (the model can still raise limit and page). The function is pure - it returns a new params object and never mutates its input.

Call it yourself if you build a custom tool executor over BUILT_IN_TOOLS:

const result = await executeTool(toolName, applyAiToolDefaults(toolName, params))
ts

The default deliberately lives in the AI layer, not in the shared tools dispatch - records.query doubles as the SDK's general record-read primitive (chat history, cron, app actions), and those internal callers must stay unbounded.

See also#