job-queue.mjs
· job-worker.mjs
· the ready-made table over it: <bg-jobs>
General-purpose background job queue with two storage modes. Pick the one that matches the lifetime your work needs:
This page documents the queue's JavaScript API. For a table with job statuses, counts, retry and delete actions, see <bg-jobs>. Its documentation explains where that component is appropriate and how its actions affect queued work.
enqueue(executor, opts) — closure
jobs. Run on this tab's main thread, never written to IndexedDB, observable
only on the originating tab. Lost on reload.
enqueue_fetch(params) — persistent
fetch jobs. Persisted in IDB. Whichever tab currently holds the
navigator.locks lease processes them, and every open tab
receives 'job-update' events via BroadcastChannel.
Survive a hard reload because the executor is rebuilt from the persisted
{method, path, body}.
Both modes use the same retry timing (see job-backoff.mjs) and events EventTarget. They differ in storage and which tabs can access them. Both return {id, done}. When the job reaches a terminal state, done resolves to {status, result, error, http_status, code}. status describes the job; http_status describes the server response; code identifies the specific error. Use the code when explaining an error, because several errors with different remedies can share one HTTP status.
Read these design constraints before extending the queue; the API signatures alone do not explain them.
The queue claims three names that are global to the browser origin: an IndexedDB database,
a navigator.locks name that elects the single processor, and a
BroadcastChannel that carries updates between tabs. This library does not know
what to call them — two applications sharing one origin under one set of names would share
one queue — so the embedding application names them:
Omit a field and it keeps its neutral default (ui-jobs,
ui-job-processor, ui-jobs), which is what a page that never calls
configure() — a bare harness — runs on. This page is not one of them: it loads the
surrounding application's own first-loaded module, so the demos below run under that
application's names, and DevTools will show you its database rather than ui-jobs.
Configure names before any module imports the queue, even if no job has been enqueued yet. job-queue.mjs opens the channel while it evaluates,
so the names are read at import. That is why configure() lives in a module of
its own: importing it from job-queue.mjs would evaluate that module and bind the default before configuration could run. A page that loads a component importing the queue —
a jobs table, a queue badge — has already read the names by the time its application entry
module runs, so the call belongs in whatever module the page loads first.
A late call throws rather than taking effect. Changing names mid-session would make jobs under the original names unreachable. A persistent job's _params may be the only remaining copy of the requested work (Startup and cleanup). Changing these names in a later release also makes stored jobs unreachable, without an error. Treat all three names as stable parts of the storage format.
Omitting configuration silently uses the default names. An unconfigured queue stores jobs in the default database, which configured pages do not read. This can happen when the browser combines a cached application module with a newer queue module. Reloading consistent code restores access to the original jobs, but jobs created in the default database remain inaccessible to configured pages. Deploy queue changes together with the application code that configures their names.
A persistent job is rebuilt from a stored {method, path, body} and has to be sent
by the application's API transport. The application defines the URL and authentication, and supplies a function for the queue to call on every replay:
Return the unread Response, including non-success statuses. That is
what keeps the retry policy in one place: the built-in 'fetch' reconstructor turns
!res.ok into an error carrying status and body, and the worker uses that status to distinguish "later" from "never". Returning parsed data would force each application to reimplement the retry policy, potentially handling the same status differently. So do not throw on a non-ok
status, and do not read the body — return the response even when the server rejected the request.
The application controls URL construction and authentication. The example above sends
a bearer token because that is what its API wants; a cookie, a differently-named header, a signed URL, or no authentication can be implemented by the same transport interface. The queue composes no
URL and sets no header of its own — it knows the path its caller enqueued and
nothing else about the network.
Read current credentials on every replay. A job may sit in the database for hours and come back after a refresh, so capturing credentials at setup would cause retries to keep using them after expiry. The job could then remain on the 401 branch indefinitely; these demos cannot reproduce an expired application credential. Read credentials inside the function instead of capturing an earlier value.
Replaying a fetch job without a configured transport throws an error. The queue can check this only when a fetch job runs, because transport configuration may legitimately happen any time before replay. A silent no-op would leave real work sitting
in the database, retried and abandoned on a widening backoff, without a useful explanation. The lookup therefore throws an error pointing to this page. Volatile enqueue() jobs need no transport at
all; they run the closure they were given.
A job's executor is always a closure, and a closure cannot cross a worker boundary — so every job runs on a main thread. The split is over who schedules it:
enqueue() keeps the closure in this tab's memory and drives it
from _run_volatile_loop() in job-queue.mjs. Nothing is
written to IndexedDB, so nothing survives a reload and no other tab can see it.
enqueue_fetch() writes a record to IndexedDB;
job-worker.mjs owns that database, the retry schedule, and the
cross-tab single-runner lease. When a job is ready the worker posts
{type:'run', job_id} to the main thread, which looks the closure up
in its in-memory map — rebuilding it from the record's
_reconstructor factory if this is a fresh page — runs it, and
replies {type:'run_result', job_id, ok, result|error, status, code}.
The worker records the outcome and resumes. Only fields included in that message survive the closure. Other error details are lost when the catch returns, which is why the response body is
read for its code there rather than persisted whole.
{type:'run_started', job_id}. It does not settle the
run; it tells the worker this page has the job in hand. See
What bounds a run for what happens when it never arrives.
Both paths import make_backoff() from
job-backoff.mjs so their retry
timing is identical rather than merely similar. Live updates reach other tabs
through cross-tab-events.mjs
(a BroadcastChannel), which is why a volatile job — never persisted —
is also never observable elsewhere.
The queue registers as the job-queue service.
Stopping it clears the volatile retry wake and closes the channel. The queue still runs
after a stop, but its updates then reach only this tab: the worker can answer after the
stop, and posting that answer on a closed channel would throw InvalidStateError.
A persistent record stores _reconstructor (a factory key) and
_params, never the closure. register_reconstructor(name,
factory) is how app code teaches the queue to rebuild its own job types;
'fetch' is registered in-library and is what makes
enqueue_fetch survive a hard reload. A persistent job whose
reconstructor is not registered on the page that picks it up cannot run.
Register factories the page must have before anything imports the queue through the
configuration instead: configure({ reconstructors: { my_kind: factory } }) in
the same call that names the queue. A component that upgrades during
parsing can boot the worker, and a job it hands to a page with no factory for it fails as
orphaned; importing the queue from the configuring module to register early would open
the queue's channel on every page that only wanted its names set.
The database the application named (version 1), object store
jobs, keyed on id.
| Field | Type | Description |
|---|---|---|
id | string | crypto.randomUUID(), primary key |
group_id | string|null | Groups related jobs for rollup queries |
type | string | Job type label, e.g. 'create_section'; also the key a yield is configured by |
label | string|null | A readable name for job surfaces; <bg-jobs> prints it instead of type |
status | string | pending → running → done | failed | cancelled; paused while the type's lock is held |
result | any | Closure return value on success |
error | string|null | Error message on failure |
http_status | number|null | HTTP status carried by the last attempt that had one; cleared to null when the job succeeds |
code | string|null | Machine error code named by that attempt's response body (a JSON body's code field), or set directly on the thrown error by a custom executor. null where the response named none — and on every record written before this field existed, so a reader falls back to http_status rather than assuming one is there |
attempts | number | Times executed so far |
max_attempts | number | Retries before failed (default 5) |
next_retry_at | number | Timestamp; 0 = ready now |
created_at / updated_at | number | Timestamps |
_reconstructor | string? | Factory key for rebuilding the closure after a reload, on any page that registered it |
_params | any? | Params handed to that factory |
Indexes: by_status, by_group,
by_group_status (compound), by_next_retry.
A persistent job is replayed from its stored {method, path, body}, so
every attempt sends a byte-identical request. Some HTTP statuses therefore need special handling instead of the ordinary backoff schedule.
next_retry_at = now + min(1000 · 2attempt−1, 30000) plus
up to 30% jitter — so roughly 1s, 2s, 4s, 8s, 16s, then capped near 30s.
status === 401 is an expired token, not a broken job: the record
goes back to pending with next_retry_at = now + 60s,
and that branch never consults max_attempts — so however many
refreshes a job waits out, none of them can flip it to failed.
The attempt is given back too, as it is for a failure with no network below: the
token expiring says nothing about the job's own work, so waiting out a refresh
costs a job none of its max_attempts budget. The built-in 'fetch'
reconstructor copies the HTTP status onto the error, which is what makes this
automatic. It also puts the response text on err.body, from which
the run's code is read — so a custom executor joins in by throwing
an error carrying status plus either body or an
already-named code, and one that throws a plain Error
gets null for both.
navigator.onLine is false, nothing could be reached, so the failure
says nothing about the job: the record goes back to pending, the
attempt is given back, and next_retry_at comes from the ordinary
ladder above — capped, so the job re-checks by itself even if no
online event is delivered, and sooner when one is. A job whose work
outlives a long disconnection is therefore not failed for waiting through it.
This is checked after the terminal branch below, so an executor that has
already decided repeating cannot help still fails at once.
failed on the first attempt,
whatever max_attempts says. Marking it failed immediately lets the caller address the refusal instead of waiting through repeated retries. The set is closed
(TERMINAL_STATUSES) rather than a per-job option: callers must handle the same HTTP status consistently. The record and its
_params are kept (see Startup and cleanup),
so the caller can show what was attempted and offer it back.
failed at max_attempts; a caller that knows its 404 is final should display that outcome itself, without changing the shared set.
A conflict that can be resolved is resolved by the caller — build a new
request and enqueue it, or call retry_all() once the state it collided
with has changed. A failed job that runs again then fails fast again rather than
re-climbing the ladder.
Only the handshake is bounded. The clock starts when the worker posts
run and stops when the page answers run_started. What it
covers is looking the record up and rebuilding its executor — a fixed shape, with no
per-machine cost in it. After that the work is untimed however long it takes:
decoding an hour of audio, uploading over a bad connection, or waiting on a lock it
needs. An executor that is slow, or quiet, is not an executor that has stopped, and
the queue has no way to tell those apart — so it does not try.
When the clock does expire the queue stops waiting, tells the page to drop whatever it
may be holding ({type:'abandon_run', job_id}), and marks the record
failed with an error naming the signal that never came. In practice that
means the page has no reconstructor registered for this kind of job, or the module
holding it never loaded.
An executor that really has stopped is not this clock's business. It
holds only its own type, because each type is its own queue, and the jobs page can end
it with delete_job(). Whether the work is still moving is something only
the executor can know, and it says so by failing.
Every look is taken again while the lanes run. A pass returns paused jobs
whose lock nobody holds, and recovers records another page left running, each
time round its loop — not once on the way in. A lane that runs for an hour would otherwise
hold all of that up for the hour. Recovery skips the jobs this page has in flight, which under
the lanes are running records too.
attempts counts the times a job was really tried, and
a record never stands above its max_attempts. The ceiling is checked
before a job is handed to the page, so no outcome can route around it: a job with
nothing left to spend is failed rather than run. Three outcomes cost no attempt,
because none of them says anything about the job's own work — a 401, a failure with
no network, and an executor that stopped to yield. A run whose page went away
does cost one, so a job that crashes its page every time eventually stops
rather than coming back for ever.
Every tab spawns its own worker, but only one may process. Each wraps its
processing loop in navigator.locks.request(…) under the lock name the
application supplied; the others wait, and a closed or crashed tab
releases the lease automatically.
Recovering orphaned running jobs must happen inside that
lock. A record left running by a killed tab is reset to
pending so the next holder picks it up — but run that reset outside
the lock and a newly-opened tab can reset a job the current holder is
actively executing, and both tabs run it. The reset runs at the start of every
pass: a page that outlives the runner of a long job — a jobs page kept open while the tab
decoding a take is closed — is the page that picks that job up, however many passes it has
already made.
The lanes do not weaken that. A pass still ends only once every lane has
drained, so the page running it has nothing in flight at the moment recovery runs. What did
change is the ceiling: a recovered job whose attempts are spent is failed rather
than started over, because the dead run spent an attempt no outcome ever answered for.
On acquiring the lock the worker resets orphaned running records left by
another page to pending, returns paused jobs whose lock nobody holds to the queue, then
processes until nothing is ready. Volatile jobs need no recovery step — they were never in the database.
Automatically clean up only successfully completed jobs. Separately from
that recovery, cleanup_old() deletes done records older than
7 days — and only those. A failed or cancelled record is
work that did not complete, and its _params are the last remaining copy of
what the caller asked for, so deleting it can remove the input the application needs to display or retry. Those records are removed only by an explicit
delete_job(), or become eligible for cleanup after retry_all() completes them successfully.
Before deleting records with any additional status, check the retention requirements of every caller. A failed request may exist only in the browser because the server does not retain rejected work. Deleting that record could permanently lose the user's input. Those retention requirements belong to the calling applications and cannot be established from this library alone.
Failed records can accumulate without a limit. Applications that create many of them should delete records once they are no longer needed. Keeping them consumes storage, but deleting them may lose work.
IndexedDB transactions auto-commit when no request is pending, so
both sweeps queue every put/delete and then await them
together. Awaiting each one in turn commits the transaction after the first and
the next throws TransactionInactiveError.
The worker starts processing when a job is enqueued, when the browser fires
online, or when the retry timer for the soonest
next_retry_at fires.
DevTools → Application → IndexedDB → the configured database →
jobs lists every persisted record. The worker's own URL carries that name, so
DevTools → Sources → the worker entry says which database this page is on. Volatile jobs
appear nowhere; query them with query_all() on the tab that created them.
Inside the running application, <bg-jobs> shows the
same set without DevTools — both kinds at once, live — which is usually the faster way to answer
"what is this queue doing right now".
await suspend_type(type) pauses pending jobs of that type, including future
ones, and sends an AbortSignal to its active executor. It resolves when that
executor settles. The executor must stop its actual work before settling and preserve any
checkpoints needed to resume. An interrupted attempt does not count as a failure, and the
job's done promise stays pending. Other job types continue to run.
resume_type(type) returns paused jobs to the queue. Suspending an already
suspended type is idempotent; callers with multiple owners must coordinate their holds.
These operations apply to this tab's volatile jobs only. A duplicate outstanding
id returns the existing completion promise. The optional label
on enqueue supplies a readable name for job surfaces.
A persistent job type can be told to yield to a Web Lock. Name the lock in the queue's
configuration, configure({ yields: { heavy_decode: 'app-foreground-priority' } }),
before anything imports the queue; the worker reads the map from its URL like the database
and lock names. While any page on the origin holds that lock exclusively, the worker
parks pending jobs of the type as paused — a persisted status, so every tab's
<bg-jobs> shows it — and runs other types on.
It keeps one shared-mode request on the lock, which the lock manager grants the moment the
holder releases, including by the holder's document being discarded, so resumption needs no
timer. A page that opens later un-parks what it finds paused under a lock nobody holds.
The foreground work takes the lock exclusively, with steal: true if it must
never wait behind the job. The job's executor receives an AbortSignal and may
also hold the same lock in shared mode for the length of its work: a steal rejects
that request, which is the executor's signal to stop. An executor that stops on request
rejects with an error carrying interrupted: true; the worker gives the attempt
back, keeps the done promise pending, and parks the job if the lock is now held
or re-queues it otherwise. An error carrying terminal: true fails the job at once,
like a 403, 409 or 422 — for work whose executor knows that repeating it cannot help.
A save made a moment after the last change, such as a position or an autosave, is debounced
in the queue rather than with a timer on the page. Each change files its job with
delay_ms and deletes the job it replaces, with delete_group or
delete_job; a cancelled record is kept until someone deletes it, so cancelling
would leave one behind per change. The latest change is then stored the moment it is made and sent once
the delay has passed. A page timer holds the change in memory until it fires, so closing the
tab inside the delay loses it; and a job handed to the queue as the page closes is lost too,
because the page is gone before the queue's worker has stored it. The stored job is sent the
next time the queue runs on this device, from any page that holds the lock.
| Export | Signature | Description |
|---|---|---|
enqueue |
(executor, opts?) => { id, done } |
Run a closure on this tab's main thread. The job lives in memory only — never persisted, never visible to other tabs, lost on reload. |
enqueue_fetch |
({type?, group_id?, method?, path, body?, max_attempts?, delay_ms?}) => { id, done } |
Persistent HTTP request. Params stored in IDB; processed by whichever tab holds the lock; rebuilt on reload via the built-in 'fetch' reconstructor and sent through the transport the application supplied, which owns the URL and any credential. path is stored verbatim and replayed unchanged, so it must be what a fresh page would send. delay_ms is passed on as enqueue_persistent describes. |
enqueue_persistent |
({reconstructor, params?, id?, type?, label?, group_id?, max_attempts?, delay_ms?}) => { id, done } |
Persistent job of any registered kind. The record stores the reconstructor's name and params; whichever page holds the lock rebuilds the executor from the factory registered under that name, so register it on every page that may run the job. A caller-supplied id is idempotent: while a job under it is open the call returns that job's completion, and a terminal record under it is replaced, which is how one job is retried without retry_all. Every call reaches the queue, including one that answers with a completion this page is already holding: that call asks only that a record exist, so it leaves an open or finished job alone while repairing an id this page believes is claimed and the queue has no record of. Deciding from the page's own memory instead left such an id claimed for the life of the page, and the work could never be queued again. delay_ms stores the job at once and holds it back that long before its first attempt; Debouncing a save is what it is for. |
register_reconstructor |
(name, factory) => void |
Register a custom executor reconstructor for durable non-fetch jobs. factory(params) => (signal) => Promise<any>; the signal aborts when the job is cancelled while running, when its record is deleted, or when the queue gave up on the handshake (What bounds a run). See Yielding for interrupted and terminal errors. |
wake |
() => void |
Ask this tab's worker to look for ready work now — for a page that learns of a job from another tab's update and is the one able to run it. |
query_group |
(group_id) => Promise<{total, pending, running, done, failed, cancelled, jobs, all_done}> |
Counts of jobs sharing group_id, the records themselves in jobs — whole records, _params included, so a caller can show the work a failed job was carrying rather than only that one failed — and a Promise that resolves when every job in the group is terminal. |
query_job |
(job_id) => Promise<Job|null> |
Fetch a single job's record — a volatile job from this tab's memory, a persistent one from IDB. Includes a done promise (matching the original return); null when no such job exists. |
query_all |
() => Promise<Job[]> |
Every job this tab can see: its own volatile jobs, plus every persistent record (active + historic) currently in IDB. |
cancel_job / cancel_group |
(id) => void |
Mark a job (or every job in a group) cancelled. Pending jobs flip immediately; a running job is told to stop through its AbortSignal on whichever tab is running it, and is recorded cancelled whatever it then reports. An executor that ignores the signal runs to completion and its result is discarded. Enqueueing the same id again afterwards files a new job that runs; the cancel does not carry over to it. |
delete_job |
(id) => void |
Remove a job — from IDB if it is persistent, from this tab's memory if it is volatile. A running job is stopped first: its executor is told to stop wherever it is running, and the run it was holding settles as abandoned. Stopping is cooperative — the built-in 'fetch' executor forwards the signal, so its request is aborted, while a custom executor that ignores its signal keeps going until it settles. Either way the record is gone and nothing is written back. The id is released on every tab: what waited on the job settles as cancelled, and a later enqueue_persistent under the same id files a fresh record rather than answering with the deleted job's completion. |
delete_group |
(group_id) => void |
Remove every stored job in the group that has not started, as delete_job removes one; what waited on each settles as cancelled. A running job is left to finish. Use it to replace a debounced save (Debouncing a save). |
retry_job |
(id) => void |
Attempt one job now, whatever it was waiting for. A job waiting out a backoff starts at once instead of at its next_retry_at; one parked behind a lock returns to the queue, and is parked again if the lock is still held; one that spent its attempts gets a fresh ladder. A running job is left alone, because the attempt this asks for is already under way, and an id the queue does not hold is answered with nothing. Use it for a control a person presses about one piece of work: enqueue_persistent cannot serve that, since a record still open is that job and filing it again deliberately changes nothing. |
retry_all |
() => void |
Reset the stored pending, paused and failed jobs and reprocess them now — and, in this tab, its volatile pending and failed ones. Failed jobs get their attempt counter reset. A running job is left alone; end one with delete_job(). This reaches every job of every type on the origin, so prefer retry_job when the press was about one of them. |
events |
EventTarget | Emits 'job-update' with the job record in e.detail. Volatile (closure) updates fire only on the originating tab; persistent updates broadcast to every open tab via BroadcastChannel. |
Backoff: min(1000 · 2^(attempt-1), 30000) + 30% jitter, capped at 30 s; max_attempts default 5.
A job is marked failed when attempts >= max_attempts, or on the first 403, 409 or 422; earlier failures flip the status back to pending with a scheduled next_retry_at.
A 401 on a persistent fetch job never fails it, however often it repeats — the record re-schedules in 60 s.
A terminal record is swept after 7 days only when its status is done — failed and cancelled records are kept until delete_job.
enqueue round-tripRun a closure, await its done promise, log the result.
events listenerSubscribe to 'job-update' for live status transitions. Each demo job below transitions pending → running → done. Jobs are volatile — never written to IDB, never visible to other tabs.
| Id | Status | Updated |
|---|
query_group rollup + all_doneGroup jobs by group_id, then poll the rollup or await all_done for the whole batch.
The closure throws at random; the queue retries with backoff up to max_attempts. Watch the status flip back to pending between attempts.
cancel_job / delete_job / retry_allControl existing jobs directly. Cancel a long-running job; delete a terminated one from IDB; reset every failed/pending job and re-run.
| Id | Type | Status |
|---|