Job queue

job-queue.mjs · job-worker.mjs · the ready-made table over it: <bg-jobs>

General-purpose background job queue with two storage modes. Pick the one that matches the lifetime your work needs:

This page documents the queue's JavaScript API. For a table with job statuses, counts, retry and delete actions, see <bg-jobs>. Its documentation explains where that component is appropriate and how its actions affect queued work.

Both modes use the same retry timing (see job-backoff.mjs) and events EventTarget. They differ in storage and which tabs can access them. Both return {id, done}. When the job reaches a terminal state, done resolves to {status, result, error, http_status, code}. status describes the job; http_status describes the server response; code identifies the specific error. Use the code when explaining an error, because several errors with different remedies can share one HTTP status.

Design

Read these design constraints before extending the queue; the API signatures alone do not explain them.

Naming: configure the three shared names once

The queue claims three names that are global to the browser origin: an IndexedDB database, a navigator.locks name that elects the single processor, and a BroadcastChannel that carries updates between tabs. This library does not know what to call them — two applications sharing one origin under one set of names would share one queue — so the embedding application names them:

// In the application's own preload module, before anything imports the queue. import { configure } from '/ui/lib/job-queue-names.mjs' configure({ db: 'myapp-jobs', lock: 'myapp-job-processor', channel: 'myapp-jobs', })

Omit a field and it keeps its neutral default (ui-jobs, ui-job-processor, ui-jobs), which is what a page that never calls configure() — a bare harness — runs on. This page is not one of them: it loads the surrounding application's own first-loaded module, so the demos below run under that application's names, and DevTools will show you its database rather than ui-jobs.

Configure names before any module imports the queue, even if no job has been enqueued yet. job-queue.mjs opens the channel while it evaluates, so the names are read at import. That is why configure() lives in a module of its own: importing it from job-queue.mjs would evaluate that module and bind the default before configuration could run. A page that loads a component importing the queue — a jobs table, a queue badge — has already read the names by the time its application entry module runs, so the call belongs in whatever module the page loads first.

A late call throws rather than taking effect. Changing names mid-session would make jobs under the original names unreachable. A persistent job's _params may be the only remaining copy of the requested work (Startup and cleanup). Changing these names in a later release also makes stored jobs unreachable, without an error. Treat all three names as stable parts of the storage format.

Omitting configuration silently uses the default names. An unconfigured queue stores jobs in the default database, which configured pages do not read. This can happen when the browser combines a cached application module with a newer queue module. Reloading consistent code restores access to the original jobs, but jobs created in the default database remain inaccessible to configured pages. Deploy queue changes together with the application code that configures their names.

Transport: the application sends requests and returns responses

A persistent job is rebuilt from a stored {method, path, body} and has to be sent by the application's API transport. The application defines the URL and authentication, and supplies a function for the queue to call on every replay:

// Beside the configure() above, in the module the page loads first. import { configure } from '/ui/lib/job-queue-transport.mjs' configure(async (path, { method, body }) => { const headers = { 'Accept': 'application/json' } if (session.token) headers['Authorization'] = `Bearer ${session.token}` const opts = { method, headers } if (body !== undefined) { headers['Content-Type'] = 'application/json' opts.body = JSON.stringify(body) } return fetch(API_BASE + path, opts) })

Return the unread Response, including non-success statuses. That is what keeps the retry policy in one place: the built-in 'fetch' reconstructor turns !res.ok into an error carrying status and body, and the worker uses that status to distinguish "later" from "never". Returning parsed data would force each application to reimplement the retry policy, potentially handling the same status differently. So do not throw on a non-ok status, and do not read the body — return the response even when the server rejected the request.

The application controls URL construction and authentication. The example above sends a bearer token because that is what its API wants; a cookie, a differently-named header, a signed URL, or no authentication can be implemented by the same transport interface. The queue composes no URL and sets no header of its own — it knows the path its caller enqueued and nothing else about the network.

Read current credentials on every replay. A job may sit in the database for hours and come back after a refresh, so capturing credentials at setup would cause retries to keep using them after expiry. The job could then remain on the 401 branch indefinitely; these demos cannot reproduce an expired application credential. Read credentials inside the function instead of capturing an earlier value.

Replaying a fetch job without a configured transport throws an error. The queue can check this only when a fetch job runs, because transport configuration may legitimately happen any time before replay. A silent no-op would leave real work sitting in the database, retried and abandoned on a widening backoff, without a useful explanation. The lookup therefore throws an error pointing to this page. Volatile enqueue() jobs need no transport at all; they run the closure they were given.

Where jobs run

A job's executor is always a closure, and a closure cannot cross a worker boundary — so every job runs on a main thread. The split is over who schedules it:

Both paths import make_backoff() from job-backoff.mjs so their retry timing is identical rather than merely similar. Live updates reach other tabs through cross-tab-events.mjs (a BroadcastChannel), which is why a volatile job — never persisted — is also never observable elsewhere.

The queue registers as the job-queue service. Stopping it clears the volatile retry wake and closes the channel. The queue still runs after a stop, but its updates then reach only this tab: the worker can answer after the stop, and posting that answer on a closed channel would throw InvalidStateError.

Rebuilding a closure after reload

A persistent record stores _reconstructor (a factory key) and _params, never the closure. register_reconstructor(name, factory) is how app code teaches the queue to rebuild its own job types; 'fetch' is registered in-library and is what makes enqueue_fetch survive a hard reload. A persistent job whose reconstructor is not registered on the page that picks it up cannot run.

Register factories the page must have before anything imports the queue through the configuration instead: configure({ reconstructors: { my_kind: factory } }) in the same call that names the queue. A component that upgrades during parsing can boot the worker, and a job it hands to a page with no factory for it fails as orphaned; importing the queue from the configuring module to register early would open the queue's channel on every page that only wanted its names set.

The job record

The database the application named (version 1), object store jobs, keyed on id.

FieldTypeDescription
idstringcrypto.randomUUID(), primary key
group_idstring|nullGroups related jobs for rollup queries
typestringJob type label, e.g. 'create_section'; also the key a yield is configured by
labelstring|nullA readable name for job surfaces; <bg-jobs> prints it instead of type
statusstringpending → running → done | failed | cancelled; paused while the type's lock is held
resultanyClosure return value on success
errorstring|nullError message on failure
http_statusnumber|nullHTTP status carried by the last attempt that had one; cleared to null when the job succeeds
codestring|nullMachine error code named by that attempt's response body (a JSON body's code field), or set directly on the thrown error by a custom executor. null where the response named none — and on every record written before this field existed, so a reader falls back to http_status rather than assuming one is there
attemptsnumberTimes executed so far
max_attemptsnumberRetries before failed (default 5)
next_retry_atnumberTimestamp; 0 = ready now
created_at / updated_atnumberTimestamps
_reconstructorstring?Factory key for rebuilding the closure after a reload, on any page that registered it
_paramsany?Params handed to that factory

Indexes: by_status, by_group, by_group_status (compound), by_next_retry.

Retry timing and special HTTP statuses

A persistent job is replayed from its stored {method, path, body}, so every attempt sends a byte-identical request. Some HTTP statuses therefore need special handling instead of the ordinary backoff schedule.

A conflict that can be resolved is resolved by the caller — build a new request and enqueue it, or call retry_all() once the state it collided with has changed. A failed job that runs again then fails fast again rather than re-climbing the ladder.

What bounds a run, and what an attempt counts

Only the handshake is bounded. The clock starts when the worker posts run and stops when the page answers run_started. What it covers is looking the record up and rebuilding its executor — a fixed shape, with no per-machine cost in it. After that the work is untimed however long it takes: decoding an hour of audio, uploading over a bad connection, or waiting on a lock it needs. An executor that is slow, or quiet, is not an executor that has stopped, and the queue has no way to tell those apart — so it does not try.

When the clock does expire the queue stops waiting, tells the page to drop whatever it may be holding ({type:'abandon_run', job_id}), and marks the record failed with an error naming the signal that never came. In practice that means the page has no reconstructor registered for this kind of job, or the module holding it never loaded.

An executor that really has stopped is not this clock's business. It holds only its own type, because each type is its own queue, and the jobs page can end it with delete_job(). Whether the work is still moving is something only the executor can know, and it says so by failing.

Every look is taken again while the lanes run. A pass returns paused jobs whose lock nobody holds, and recovers records another page left running, each time round its loop — not once on the way in. A lane that runs for an hour would otherwise hold all of that up for the hour. Recovery skips the jobs this page has in flight, which under the lanes are running records too.

attempts counts the times a job was really tried, and a record never stands above its max_attempts. The ceiling is checked before a job is handed to the page, so no outcome can route around it: a job with nothing left to spend is failed rather than run. Three outcomes cost no attempt, because none of them says anything about the job's own work — a 401, a failure with no network, and an executor that stopped to yield. A run whose page went away does cost one, so a job that crashes its page every time eventually stops rather than coming back for ever.

Preventing duplicate execution across tabs

Every tab spawns its own worker, but only one may process. Each wraps its processing loop in navigator.locks.request(…) under the lock name the application supplied; the others wait, and a closed or crashed tab releases the lease automatically.

Recovering orphaned running jobs must happen inside that lock. A record left running by a killed tab is reset to pending so the next holder picks it up — but run that reset outside the lock and a newly-opened tab can reset a job the current holder is actively executing, and both tabs run it. The reset runs at the start of every pass: a page that outlives the runner of a long job — a jobs page kept open while the tab decoding a take is closed — is the page that picks that job up, however many passes it has already made.

The lanes do not weaken that. A pass still ends only once every lane has drained, so the page running it has nothing in flight at the moment recovery runs. What did change is the ceiling: a recovered job whose attempts are spent is failed rather than started over, because the dead run spent an attempt no outcome ever answered for.

Startup and cleanup

On acquiring the lock the worker resets orphaned running records left by another page to pending, returns paused jobs whose lock nobody holds to the queue, then processes until nothing is ready. Volatile jobs need no recovery step — they were never in the database.

Automatically clean up only successfully completed jobs. Separately from that recovery, cleanup_old() deletes done records older than 7 days — and only those. A failed or cancelled record is work that did not complete, and its _params are the last remaining copy of what the caller asked for, so deleting it can remove the input the application needs to display or retry. Those records are removed only by an explicit delete_job(), or become eligible for cleanup after retry_all() completes them successfully.

Before deleting records with any additional status, check the retention requirements of every caller. A failed request may exist only in the browser because the server does not retain rejected work. Deleting that record could permanently lose the user's input. Those retention requirements belong to the calling applications and cannot be established from this library alone.

Failed records can accumulate without a limit. Applications that create many of them should delete records once they are no longer needed. Keeping them consumes storage, but deleting them may lose work.

IndexedDB transactions auto-commit when no request is pending, so both sweeps queue every put/delete and then await them together. Awaiting each one in turn commits the transaction after the first and the next throws TransactionInactiveError.

Wake triggers

The worker starts processing when a job is enqueued, when the browser fires online, or when the retry timer for the soonest next_retry_at fires.

Inspecting jobs

DevTools → Application → IndexedDB → the configured database → jobs lists every persisted record. The worker's own URL carries that name, so DevTools → Sources → the worker entry says which database this page is on. Volatile jobs appear nowhere; query them with query_all() on the tab that created them.

Inside the running application, <bg-jobs> shows the same set without DevTools — both kinds at once, live — which is usually the faster way to answer "what is this queue doing right now".

Suspending volatile work

await suspend_type(type) pauses pending jobs of that type, including future ones, and sends an AbortSignal to its active executor. It resolves when that executor settles. The executor must stop its actual work before settling and preserve any checkpoints needed to resume. An interrupted attempt does not count as a failure, and the job's done promise stays pending. Other job types continue to run.

resume_type(type) returns paused jobs to the queue. Suspending an already suspended type is idempotent; callers with multiple owners must coordinate their holds. These operations apply to this tab's volatile jobs only. A duplicate outstanding id returns the existing completion promise. The optional label on enqueue supplies a readable name for job surfaces.

Yielding: persistent work that pauses while a lock is held

A persistent job type can be told to yield to a Web Lock. Name the lock in the queue's configuration, configure({ yields: { heavy_decode: 'app-foreground-priority' } }), before anything imports the queue; the worker reads the map from its URL like the database and lock names. While any page on the origin holds that lock exclusively, the worker parks pending jobs of the type as paused — a persisted status, so every tab's <bg-jobs> shows it — and runs other types on. It keeps one shared-mode request on the lock, which the lock manager grants the moment the holder releases, including by the holder's document being discarded, so resumption needs no timer. A page that opens later un-parks what it finds paused under a lock nobody holds.

The foreground work takes the lock exclusively, with steal: true if it must never wait behind the job. The job's executor receives an AbortSignal and may also hold the same lock in shared mode for the length of its work: a steal rejects that request, which is the executor's signal to stop. An executor that stops on request rejects with an error carrying interrupted: true; the worker gives the attempt back, keeps the done promise pending, and parks the job if the lock is now held or re-queues it otherwise. An error carrying terminal: true fails the job at once, like a 403, 409 or 422 — for work whose executor knows that repeating it cannot help.

Debouncing a save

A save made a moment after the last change, such as a position or an autosave, is debounced in the queue rather than with a timer on the page. Each change files its job with delay_ms and deletes the job it replaces, with delete_group or delete_job; a cancelled record is kept until someone deletes it, so cancelling would leave one behind per change. The latest change is then stored the moment it is made and sent once the delay has passed. A page timer holds the change in memory until it fires, so closing the tab inside the delay loses it; and a job handed to the queue as the page closes is lost too, because the page is gone before the queue's worker has stored it. The stored job is sent the next time the queue runs on this device, from any page that holds the lock.

API

ExportSignatureDescription
enqueue (executor, opts?) => { id, done } Run a closure on this tab's main thread. The job lives in memory only — never persisted, never visible to other tabs, lost on reload.
enqueue_fetch ({type?, group_id?, method?, path, body?, max_attempts?, delay_ms?}) => { id, done } Persistent HTTP request. Params stored in IDB; processed by whichever tab holds the lock; rebuilt on reload via the built-in 'fetch' reconstructor and sent through the transport the application supplied, which owns the URL and any credential. path is stored verbatim and replayed unchanged, so it must be what a fresh page would send. delay_ms is passed on as enqueue_persistent describes.
enqueue_persistent ({reconstructor, params?, id?, type?, label?, group_id?, max_attempts?, delay_ms?}) => { id, done } Persistent job of any registered kind. The record stores the reconstructor's name and params; whichever page holds the lock rebuilds the executor from the factory registered under that name, so register it on every page that may run the job. A caller-supplied id is idempotent: while a job under it is open the call returns that job's completion, and a terminal record under it is replaced, which is how one job is retried without retry_all. Every call reaches the queue, including one that answers with a completion this page is already holding: that call asks only that a record exist, so it leaves an open or finished job alone while repairing an id this page believes is claimed and the queue has no record of. Deciding from the page's own memory instead left such an id claimed for the life of the page, and the work could never be queued again. delay_ms stores the job at once and holds it back that long before its first attempt; Debouncing a save is what it is for.
register_reconstructor (name, factory) => void Register a custom executor reconstructor for durable non-fetch jobs. factory(params) => (signal) => Promise<any>; the signal aborts when the job is cancelled while running, when its record is deleted, or when the queue gave up on the handshake (What bounds a run). See Yielding for interrupted and terminal errors.
wake () => void Ask this tab's worker to look for ready work now — for a page that learns of a job from another tab's update and is the one able to run it.
query_group (group_id) => Promise<{total, pending, running, done, failed, cancelled, jobs, all_done}> Counts of jobs sharing group_id, the records themselves in jobs — whole records, _params included, so a caller can show the work a failed job was carrying rather than only that one failed — and a Promise that resolves when every job in the group is terminal.
query_job (job_id) => Promise<Job|null> Fetch a single job's record — a volatile job from this tab's memory, a persistent one from IDB. Includes a done promise (matching the original return); null when no such job exists.
query_all () => Promise<Job[]> Every job this tab can see: its own volatile jobs, plus every persistent record (active + historic) currently in IDB.
cancel_job / cancel_group (id) => void Mark a job (or every job in a group) cancelled. Pending jobs flip immediately; a running job is told to stop through its AbortSignal on whichever tab is running it, and is recorded cancelled whatever it then reports. An executor that ignores the signal runs to completion and its result is discarded. Enqueueing the same id again afterwards files a new job that runs; the cancel does not carry over to it.
delete_job (id) => void Remove a job — from IDB if it is persistent, from this tab's memory if it is volatile. A running job is stopped first: its executor is told to stop wherever it is running, and the run it was holding settles as abandoned. Stopping is cooperative — the built-in 'fetch' executor forwards the signal, so its request is aborted, while a custom executor that ignores its signal keeps going until it settles. Either way the record is gone and nothing is written back. The id is released on every tab: what waited on the job settles as cancelled, and a later enqueue_persistent under the same id files a fresh record rather than answering with the deleted job's completion.
delete_group (group_id) => void Remove every stored job in the group that has not started, as delete_job removes one; what waited on each settles as cancelled. A running job is left to finish. Use it to replace a debounced save (Debouncing a save).
retry_job (id) => void Attempt one job now, whatever it was waiting for. A job waiting out a backoff starts at once instead of at its next_retry_at; one parked behind a lock returns to the queue, and is parked again if the lock is still held; one that spent its attempts gets a fresh ladder. A running job is left alone, because the attempt this asks for is already under way, and an id the queue does not hold is answered with nothing. Use it for a control a person presses about one piece of work: enqueue_persistent cannot serve that, since a record still open is that job and filing it again deliberately changes nothing.
retry_all () => void Reset the stored pending, paused and failed jobs and reprocess them now — and, in this tab, its volatile pending and failed ones. Failed jobs get their attempt counter reset. A running job is left alone; end one with delete_job(). This reaches every job of every type on the origin, so prefer retry_job when the press was about one of them.
events EventTarget Emits 'job-update' with the job record in e.detail. Volatile (closure) updates fire only on the originating tab; persistent updates broadcast to every open tab via BroadcastChannel.

Backoff: min(1000 · 2^(attempt-1), 30000) + 30% jitter, capped at 30 s; max_attempts default 5. A job is marked failed when attempts >= max_attempts, or on the first 403, 409 or 422; earlier failures flip the status back to pending with a scheduled next_retry_at. A 401 on a persistent fetch job never fails it, however often it repeats — the record re-schedules in 60 s. A terminal record is swept after 7 days only when its status is done — failed and cancelled records are kept until delete_job.

1. enqueue round-trip

Run a closure, await its done promise, log the result.

No result yet
Show code

2. events listener

Subscribe to 'job-update' for live status transitions. Each demo job below transitions pending → running → done. Jobs are volatile — never written to IDB, never visible to other tabs.

IdStatusUpdated
Show code

3. query_group rollup + all_done

Group jobs by group_id, then poll the rollup or await all_done for the whole batch.

No batch yet
Show code

4. Failure + retry

The closure throws at random; the queue retries with backoff up to max_attempts. Watch the status flip back to pending between attempts.

No job
Show code

5. cancel_job / delete_job / retry_all

Control existing jobs directly. Cancel a long-running job; delete a terminated one from IDB; reset every failed/pending job and re-run.

IdTypeStatus
Show code