Skip to main content
Documentation

Background jobs

Durable, observable background work that outlives the HTTP response.

On this page

DeepSpace apps include a per-app AppJobRoom Durable Object for background work that can't or shouldn't run inside an HTTP handler - long AI generations, CSV exports, image renders, bulk imports, fan-out side effects. Job rows are persisted in SQLite and broadcast every state change over WebSocket so clients see live progress and can cancel or retry without polling. A Worker restart preserves the job row, but interrupts the running handler; see recovery.

Use this instead of ctx.waitUntil(...). waitUntil is killed 30 seconds after the response goes out - the JobRoom replaces that pattern. For work that runs on a schedule (daily digest, hourly sync), use scheduled tasks; for privileged work that finishes inside the HTTP response, use server actions.

When to use jobs vs. cron vs. server actions#

If the work…Use
Finishes inside the HTTP responseA regular Hono route or server action
Runs on a schedule (daily digest, hourly sync)Scheduled tasks
Is triggered by a user click and may take seconds to more than 15 minutesBackground jobs (this guide)
Is triggered by the worker and needs to outlive the responseBackground jobs (this guide)

Define handlers in src/jobs.ts#

A single runJob function dispatches every job type. It receives the Job row and a JobContext, returns the result on success, and throws to fail.

import type { Job, JobContext } from 'deepspace/worker'

export async function runJob(
  job: Job,
  ctx: JobContext,
  env: Env,
): Promise<unknown | void> {
  if (job.type === 'ai-summarize') {
    const { text } = job.payload as { text: string }
    ctx.progress(0.1, 'starting')

    // Pass ctx.signal so Cancel actually aborts the upstream fetch.
    const summary = await callModel(text, { signal: ctx.signal })

    return { summary, words: summary.split(/\s+/).length }
  }

  if (job.type === 'export-csv') {
    // ... long export, call ctx.progress(p, msg) periodically
  }

  // Unknown type → fail loudly so it shows up in the failed list.
  throw new Error(`Unknown job type: ${job.type}`)
}
ts

The JobContext API#

MemberPurpose
ctx.progress(value, message?)Broadcast progress (0..1). Re-renders any useJobs subscriber.
ctx.signalAbortSignal that fires when the client calls cancel(id) (same isolate). Pass to fetch(url, { signal: ctx.signal }) to abort upstream requests cleanly.
ctx.continue(state, { afterMs? })Checkpoint state and yield to the next alarm tick. Use when the work can be split into resumable parts. Call ctx.continue(...) then return - the next tick invokes onJob again with job.resumeFrom = state.

Worker wiring#

The scaffolded worker.ts already wires AppJobRoom:

export class AppJobRoom extends JobRoom<Env> {
  constructor(state: DurableObjectState, env: Env) {
    super(state, env)
  }
  protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
    return await runJob(job, ctx, this.env)
  }
}
ts

Add job handlers in src/jobs.ts. For handlers that need more than 15 minutes in one run, also configure their type names on AppJobRoom as shown below.

Jobs longer than 15 minutes#

By default, JobRoom waits for handlers inside a Durable Object alarm, which has a 15-minute wall-time limit. If a job waits on a long model call or document conversion that cannot be split into smaller parts, add its type to backgroundJobTypes:

export class AppJobRoom extends JobRoom<Env> {
  constructor(state: DurableObjectState, env: Env) {
    super(state, env, {
      backgroundJobTypes: ['generate-report'],
      authorizeWrite: async (user) => {
        if (user.userId.startsWith('anon-')) return false
        const role = await resolveAppRole(env, user.userId)
        return role === 'member' || role === 'admin'
      },
    })
  }

  protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
    return runJob(job, ctx, this.env)
  }
}
ts

Keep your existing authorization callbacks when adding the setting. The strings are your existing job type names: enqueue 'generate-report' through useJobs or enqueueJob as usual. This works for any app-defined type and needs no new Worker, binding, or route.

For a listed type, the alarm starts the handler and returns. Short alarms keep the room active while the handler waits on external work, so the handler can continue beyond the alarm's 15-minute limit. Jobs still run one at a time; other types keep the default alarm behavior. This setting does not increase Worker CPU or memory limits.

Set a handler deadline#

backgroundJobTypes does not set a timeout. Give the handler a deadline and combine it with ctx.signal so both timeout and user cancellation stop the work:

if (job.type === 'generate-report') {
  const signal = AbortSignal.any([
    ctx.signal,
    AbortSignal.timeout(45 * 60 * 1000),
  ])
  return await generateReport(job.payload, env, signal)
}
ts

Pass the signal through model calls, tool calls, and output downloads. For APIs without an abort option, enforce a timeout in the app's transfer code. Cancellation marks the row immediately, but the queue waits for the current handler to exit before starting another job.

Choose continuous work or checkpoints#

Use backgroundJobTypes when one external operation may take longer than 15 minutes. Use ctx.continue(state) when the task can save progress between smaller parts. Both use the same queue, progress updates, cancellation, and retry settings.

Neither setting preserves a running JavaScript function across a Worker restart. On recovery, JobRoom invokes the handler again if attempts remain, with job.resumeFrom containing the last checkpoint when one was saved. Write handlers so a retry does not duplicate completed side effects.

Enqueueing - two entry points, one queue#

Two enqueue paths write to the same DO row. Pick by where the caller lives.

From the client - useJobs#

The useJobs hook returns a live jobs list and an enqueue function. Every connected subscriber sees the same state transitions in real time.

import { useJobs } from 'deepspace'
import { SCOPE_ID } from '../constants'

function ExportButton() {
  const { enqueue, jobs, cancel, retry } = useJobs(SCOPE_ID)

  return (
    <>
      <button onClick={() => enqueue('export-csv', { filterId: 'q1' }, { maxAttempts: 2 })}>
        Run export
      </button>
      <ul>
        {jobs.map((j) => (
          <li key={j.id}>
            {j.type} - {j.status}
            {j.progress != null && <> ({Math.round(j.progress * 100)}%)</>}
            {j.status === 'running' && <button onClick={() => cancel(j.id)}>Cancel</button>}
            {j.status === 'failed' && <button onClick={() => retry(j.id)}>Retry</button>}
          </li>
        ))}
      </ul>
    </>
  )
}
tsx

enqueue resolves with the jobId once the server acks. jobs is sorted with live/recent first and re-renders on every state change.

From the worker - enqueueJob#

Use this from HTTP routes, server actions, cron handlers, AI routes - anywhere the JobRoom DO isn't the current isolate.

import { enqueueJob } from 'deepspace/worker'

app.post('/api/start-export', async (c) => {
  const auth = await resolveAuth(c.req.raw, c.env)
  if (!auth) return c.json({ error: 'unauthorized' }, 401)

  const jobId = await enqueueJob(
    c.env.JOB_ROOMS,
    `app:${c.env.DEEPSPACE_APP_ID}`,   // immutable app id — names are mutable URL leases
    'export-csv',
    { filterId: '...' },
    { maxAttempts: 2, enqueuedBy: auth.userId },
  )

  return c.json({ jobId })
})
ts

Inside AppJobRoom.onJob(...) itself, call this.enqueue('next-step', payload) to chain follow-up work - the in-isolate call skips the HTTP hop that enqueueJob makes from outside:

export class AppJobRoom extends JobRoom<Env> {
  protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
    const result = await runJob(job, ctx, this.env)
    if (job.type === 'export-csv') this.enqueue('email-export', { jobId: job.id })
    return result
  }
}
ts

Lifecycle and limits#

Every default below is fixed when the DO is constructed. Override them by passing a config object to super(state, env, { ... }) inside AppJobRoom - see the JobRoomConfig reference for every knob.

ConcernDefaultHow to change
Retry on throwNone (maxAttempts: 1)Pass { maxAttempts: N } to enqueue
Retry backoff1 sOverride retryBackoffMs on AppJobRoom config
Terminal-row retention24 hOverride retentionMs on AppJobRoom config
Alarm wall-time limit15 minSplit work with ctx.continue(state), or list its type in backgroundJobTypes
Handler deadline for a listed typeApp-definedCombine a timeout with ctx.signal
Crash recoveryAutoOrphaned running rows older than ~16 min are rescued on a DO wake-up; a handler still active in the same isolate is left running

State machine: queued → running → succeeded | failed | canceled.

Crash-recovery outcomes are deterministic:

  • A rescued running row is retried if attempts remain, otherwise marked failed.
  • A retry starts the handler again; it does not resume an interrupted model request. A saved checkpoint is available in job.resumeFrom.
  • A cancel that lands while the handler is running flips the row to canceled; a return value that arrives afterward (including from another isolate) is discarded rather than overwriting the canceled state.
  • useJobs auto-reconnects on WebSocket drop, so subscribers converge on the recovered state without a refresh.

Outbound calls in handlers#

Handlers run as the app owner, just like scheduled tasks. Use createDeepSpaceAI(env, 'anthropic') for AI calls - it falls back to APP_OWNER_JWT and bills the developer. Pass ctx.signal to every fetch(...) so client cancel aborts cleanly upstream:

const res = await fetch('https://api.example.com/render', {
  method: 'POST',
  body: JSON.stringify({ ... }),
  signal: ctx.signal,
})
ts

Who can enqueue#

The scaffolded AppJobRoom authorizes writes server-side:

export class AppJobRoom extends JobRoom<Env> {
  constructor(state: DurableObjectState, env: Env) {
    super(state, env, {
      authorizeWrite: async (user) => {
        if (user.userId.startsWith('anon-')) return false
        const role = await resolveAppRole(env, user.userId)
        return role === 'member' || role === 'admin'
      },
    })
  }
}
ts

enqueue, cancel, and retry therefore require a verified user whose current app role is member or admin, and the check runs again before every mutation. Hiding the button behind useUser().user?.role === 'admin' or a (protected)/ route is good UX, but it is not what stops an unauthorized enqueue - authorizeWrite is.

Two patterns on top of the default:

  • Paid jobs stay owner-only. If a handler spends owner credits (integrations, AI proxies), tighten authorizeWrite to role === 'admin' - an ordinary member shouldn't be able to spend the owner's credits from the console.
  • A deliberately public producer goes through HTTP, not the socket. Don't loosen authorizeWrite to let anonymous connections write. Instead, expose one app-owned HTTP action that validates a named job type and a bounded payload, rate-limits the caller, and then calls enqueueJob server-side:
app.post('/api/request-summary', async (c) => {
  const { text } = await c.req.json<{ text?: string }>()
  if (typeof text !== 'string' || text.length > 10_000) {
    return c.json({ error: 'invalid_payload' }, 400)
  }
  // Rate-limit here (per IP or per user) before spending anything.
  const jobId = await enqueueJob(
    c.env.JOB_ROOMS,
    `app:${c.env.DEEPSPACE_APP_ID}`,
    'ai-summarize',        // one named type — never a caller-chosen type
    { text },
  )
  return c.json({ jobId })
})
ts

The route owns validation, rate limits, and which job types the public may create; the DO's write role keeps every other path closed.

Testing without waiting for a real upstream#

Two approaches work well:

  1. Use a fast handler in tests. A job type like 'echo' that returns its payload with no I/O lets you assert the full enqueue → run → succeed pipeline in under a second without mocking upstreams.
  2. Hit the enqueue route from a Playwright spec. Render a page that uses useJobs, click the enqueue button, then assert against the rendered status:
test('export job succeeds end-to-end', async ({ page }) => {
  await page.goto('/jobs')
  await page.getByRole('button', { name: /run export/i }).click()
  await expect(
    page.locator('[data-testid="job-row"][data-status="succeeded"]'),
  ).toBeVisible({ timeout: 30_000 })
})
ts

Don't write tests that wait for the 16-minute crash-recovery sweep, and don't manually flip DB rows - use the public enqueue / cancel / retry surface.

Next steps#