Background jobs
Durable, observable background work that outlives the HTTP response.
On this page
DeepSpace apps include a per-app AppJobRoom Durable Object for background work that can't or shouldn't run inside an HTTP handler - long AI generations, CSV exports, image renders, bulk imports, fan-out side effects. Job rows are persisted in SQLite and broadcast every state change over WebSocket so clients see live progress and can cancel or retry without polling. A Worker restart preserves the job row, but interrupts the running handler; see recovery.
Use this instead of ctx.waitUntil(...). waitUntil is killed 30 seconds after the response goes out - the JobRoom replaces that pattern. For work that runs on a schedule (daily digest, hourly sync), use scheduled tasks; for privileged work that finishes inside the HTTP response, use server actions.
When to use jobs vs. cron vs. server actions#
| If the work… | Use |
|---|---|
| Finishes inside the HTTP response | A regular Hono route or server action |
| Runs on a schedule (daily digest, hourly sync) | Scheduled tasks |
| Is triggered by a user click and may take seconds to more than 15 minutes | Background jobs (this guide) |
| Is triggered by the worker and needs to outlive the response | Background jobs (this guide) |
Define handlers in src/jobs.ts#
A single runJob function dispatches every job type. It receives the Job row and a JobContext, returns the result on success, and throws to fail.
import type { Job, JobContext } from 'deepspace/worker'
export async function runJob(
job: Job,
ctx: JobContext,
env: Env,
): Promise<unknown | void> {
if (job.type === 'ai-summarize') {
const { text } = job.payload as { text: string }
ctx.progress(0.1, 'starting')
// Pass ctx.signal so Cancel actually aborts the upstream fetch.
const summary = await callModel(text, { signal: ctx.signal })
return { summary, words: summary.split(/\s+/).length }
}
if (job.type === 'export-csv') {
// ... long export, call ctx.progress(p, msg) periodically
}
// Unknown type → fail loudly so it shows up in the failed list.
throw new Error(`Unknown job type: ${job.type}`)
}
The JobContext API#
| Member | Purpose |
|---|---|
ctx.progress(value, message?) | Broadcast progress (0..1). Re-renders any useJobs subscriber. |
ctx.signal | AbortSignal that fires when the client calls cancel(id) (same isolate). Pass to fetch(url, { signal: ctx.signal }) to abort upstream requests cleanly. |
ctx.continue(state, { afterMs? }) | Checkpoint state and yield to the next alarm tick. Use when the work can be split into resumable parts. Call ctx.continue(...) then return - the next tick invokes onJob again with job.resumeFrom = state. |
Worker wiring#
The scaffolded worker.ts already wires AppJobRoom:
export class AppJobRoom extends JobRoom<Env> {
constructor(state: DurableObjectState, env: Env) {
super(state, env)
}
protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
return await runJob(job, ctx, this.env)
}
}
Add job handlers in src/jobs.ts. For handlers that need more than 15 minutes in one run, also configure their type names on AppJobRoom as shown below.
Jobs longer than 15 minutes#
By default, JobRoom waits for handlers inside a Durable Object alarm, which has a 15-minute wall-time limit. If a job waits on a long model call or document conversion that cannot be split into smaller parts, add its type to backgroundJobTypes:
export class AppJobRoom extends JobRoom<Env> {
constructor(state: DurableObjectState, env: Env) {
super(state, env, {
backgroundJobTypes: ['generate-report'],
authorizeWrite: async (user) => {
if (user.userId.startsWith('anon-')) return false
const role = await resolveAppRole(env, user.userId)
return role === 'member' || role === 'admin'
},
})
}
protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
return runJob(job, ctx, this.env)
}
}
Keep your existing authorization callbacks when adding the setting. The strings are your existing job type names: enqueue 'generate-report' through useJobs or enqueueJob as usual. This works for any app-defined type and needs no new Worker, binding, or route.
For a listed type, the alarm starts the handler and returns. Short alarms keep the room active while the handler waits on external work, so the handler can continue beyond the alarm's 15-minute limit. Jobs still run one at a time; other types keep the default alarm behavior. This setting does not increase Worker CPU or memory limits.
Set a handler deadline#
backgroundJobTypes does not set a timeout. Give the handler a deadline and combine it with ctx.signal so both timeout and user cancellation stop the work:
if (job.type === 'generate-report') {
const signal = AbortSignal.any([
ctx.signal,
AbortSignal.timeout(45 * 60 * 1000),
])
return await generateReport(job.payload, env, signal)
}
Pass the signal through model calls, tool calls, and output downloads. For APIs without an abort option, enforce a timeout in the app's transfer code. Cancellation marks the row immediately, but the queue waits for the current handler to exit before starting another job.
Choose continuous work or checkpoints#
Use backgroundJobTypes when one external operation may take longer than 15 minutes. Use ctx.continue(state) when the task can save progress between smaller parts. Both use the same queue, progress updates, cancellation, and retry settings.
Neither setting preserves a running JavaScript function across a Worker restart. On recovery, JobRoom invokes the handler again if attempts remain, with job.resumeFrom containing the last checkpoint when one was saved. Write handlers so a retry does not duplicate completed side effects.
Enqueueing - two entry points, one queue#
Two enqueue paths write to the same DO row. Pick by where the caller lives.
From the client - useJobs#
The useJobs hook returns a live jobs list and an enqueue function. Every connected subscriber sees the same state transitions in real time.
import { useJobs } from 'deepspace'
import { SCOPE_ID } from '../constants'
function ExportButton() {
const { enqueue, jobs, cancel, retry } = useJobs(SCOPE_ID)
return (
<>
<button onClick={() => enqueue('export-csv', { filterId: 'q1' }, { maxAttempts: 2 })}>
Run export
</button>
<ul>
{jobs.map((j) => (
<li key={j.id}>
{j.type} - {j.status}
{j.progress != null && <> ({Math.round(j.progress * 100)}%)</>}
{j.status === 'running' && <button onClick={() => cancel(j.id)}>Cancel</button>}
{j.status === 'failed' && <button onClick={() => retry(j.id)}>Retry</button>}
</li>
))}
</ul>
</>
)
}
enqueue resolves with the jobId once the server acks. jobs is sorted with live/recent first and re-renders on every state change.
From the worker - enqueueJob#
Use this from HTTP routes, server actions, cron handlers, AI routes - anywhere the JobRoom DO isn't the current isolate.
import { enqueueJob } from 'deepspace/worker'
app.post('/api/start-export', async (c) => {
const auth = await resolveAuth(c.req.raw, c.env)
if (!auth) return c.json({ error: 'unauthorized' }, 401)
const jobId = await enqueueJob(
c.env.JOB_ROOMS,
`app:${c.env.DEEPSPACE_APP_ID}`, // immutable app id — names are mutable URL leases
'export-csv',
{ filterId: '...' },
{ maxAttempts: 2, enqueuedBy: auth.userId },
)
return c.json({ jobId })
})
Inside AppJobRoom.onJob(...) itself, call this.enqueue('next-step', payload) to chain follow-up work - the in-isolate call skips the HTTP hop that enqueueJob makes from outside:
export class AppJobRoom extends JobRoom<Env> {
protected async onJob(job: Job, ctx: JobContext): Promise<unknown> {
const result = await runJob(job, ctx, this.env)
if (job.type === 'export-csv') this.enqueue('email-export', { jobId: job.id })
return result
}
}
Lifecycle and limits#
Every default below is fixed when the DO is constructed. Override them by passing a config object to super(state, env, { ... }) inside AppJobRoom - see the JobRoomConfig reference for every knob.
| Concern | Default | How to change |
|---|---|---|
| Retry on throw | None (maxAttempts: 1) | Pass { maxAttempts: N } to enqueue |
| Retry backoff | 1 s | Override retryBackoffMs on AppJobRoom config |
| Terminal-row retention | 24 h | Override retentionMs on AppJobRoom config |
| Alarm wall-time limit | 15 min | Split work with ctx.continue(state), or list its type in backgroundJobTypes |
| Handler deadline for a listed type | App-defined | Combine a timeout with ctx.signal |
| Crash recovery | Auto | Orphaned running rows older than ~16 min are rescued on a DO wake-up; a handler still active in the same isolate is left running |
State machine: queued → running → succeeded | failed | canceled.
Crash-recovery outcomes are deterministic:
- A rescued
runningrow is retried if attempts remain, otherwise markedfailed. - A retry starts the handler again; it does not resume an interrupted model request. A saved checkpoint is available in
job.resumeFrom. - A cancel that lands while the handler is running flips the row to
canceled; a return value that arrives afterward (including from another isolate) is discarded rather than overwriting the canceled state. useJobsauto-reconnects on WebSocket drop, so subscribers converge on the recovered state without a refresh.
Outbound calls in handlers#
Handlers run as the app owner, just like scheduled tasks. Use createDeepSpaceAI(env, 'anthropic') for AI calls - it falls back to APP_OWNER_JWT and bills the developer. Pass ctx.signal to every fetch(...) so client cancel aborts cleanly upstream:
const res = await fetch('https://api.example.com/render', {
method: 'POST',
body: JSON.stringify({ ... }),
signal: ctx.signal,
})
Who can enqueue#
The scaffolded AppJobRoom authorizes writes server-side:
export class AppJobRoom extends JobRoom<Env> {
constructor(state: DurableObjectState, env: Env) {
super(state, env, {
authorizeWrite: async (user) => {
if (user.userId.startsWith('anon-')) return false
const role = await resolveAppRole(env, user.userId)
return role === 'member' || role === 'admin'
},
})
}
}
enqueue, cancel, and retry therefore require a verified user whose current app role is member or admin, and the check runs again before every mutation. Hiding the button behind useUser().user?.role === 'admin' or a (protected)/ route is good UX, but it is not what stops an unauthorized enqueue - authorizeWrite is.
Two patterns on top of the default:
- Paid jobs stay owner-only. If a handler spends owner credits (integrations, AI proxies), tighten
authorizeWritetorole === 'admin'- an ordinary member shouldn't be able to spend the owner's credits from the console. - A deliberately public producer goes through HTTP, not the socket. Don't loosen
authorizeWriteto let anonymous connections write. Instead, expose one app-owned HTTP action that validates a named job type and a bounded payload, rate-limits the caller, and then callsenqueueJobserver-side:
app.post('/api/request-summary', async (c) => {
const { text } = await c.req.json<{ text?: string }>()
if (typeof text !== 'string' || text.length > 10_000) {
return c.json({ error: 'invalid_payload' }, 400)
}
// Rate-limit here (per IP or per user) before spending anything.
const jobId = await enqueueJob(
c.env.JOB_ROOMS,
`app:${c.env.DEEPSPACE_APP_ID}`,
'ai-summarize', // one named type — never a caller-chosen type
{ text },
)
return c.json({ jobId })
})
The route owns validation, rate limits, and which job types the public may create; the DO's write role keeps every other path closed.
Testing without waiting for a real upstream#
Two approaches work well:
- Use a fast handler in tests. A job type like
'echo'that returns its payload with no I/O lets you assert the full enqueue → run → succeed pipeline in under a second without mocking upstreams. - Hit the enqueue route from a Playwright spec. Render a page that uses
useJobs, click the enqueue button, then assert against the rendered status:
test('export job succeeds end-to-end', async ({ page }) => {
await page.goto('/jobs')
await page.getByRole('button', { name: /run export/i }).click()
await expect(
page.locator('[data-testid="job-row"][data-status="succeeded"]'),
).toBeVisible({ timeout: 30_000 })
})
Don't write tests that wait for the 16-minute crash-recovery sweep, and don't manually flip DB rows - use the public enqueue / cancel / retry surface.
Next steps#
- Worker rooms reference -
JobRoom,Job,JobContext,enqueueJob. - Real-time reference -
useJobsreturn shape. - Scheduled tasks - if the work needs to run on a schedule instead of on demand.
- Server actions - for privileged work that finishes inside the HTTP response.