SpokenWords

What's changed in SpokenWords 14

Changes since SpokenWords 13, through 22 September 2026

The Studio now shows more of the document's original appearance. You can choose how much information appears beside the text, see how a sentence should be read aloud, and follow your reading while you record. Stopping a recording no longer makes you wait: the accurate transcript is prepared in the background, survives a lost connection or a closed tab, and the status bar says what the queue is doing. Playback starts where a sentence's sound begins. Speech recognition can now run on your computer when your team has that service; the computer measures itself, chooses its speech models and says so when it cannot run them. Saving and production messages give clearer explanations when something needs your attention.

This report covers the changes you can see or use. Changes to how the app is built inside are grouped under Behind the scenes. It uses the English names shown in the app. Depending on your role and your team's speech service, some controls may not be available to you.

Read the documentHeadings, tables and pictures in the Studio. Choose your viewTry the reading choices in the example below. Follow your readingSee your place and words that did not match. Keep recordingAccurate transcription finishes in the background.

The pictures below use example text to show what changed.

Reading in the Studio

BeforeStudio

A walk by the river

Dr. Green walked along the river.

The water was cold.

The route

Place

Distance

Bridge

2 km

NowStudio

A walk by the river

Dr. Green walked along the river. The water was cold.

The route
PlaceDistance
Bridge2 km

12

Before, each sentence had its own line. Now the Studio keeps the document's paragraphs, headings and tables together.

Read the document with its formatting

The reading area now shows the document's paragraphs, headings, lists, tables, notes and other formatting together. A chapter the server has no imported document for still shows one sentence per line. The command Original document in a new window has been removed; you read and work with the document in the Studio.

Status dots remain useful in every view

Follow links, go to a page and search

Links in the document now work in the reading area. A link in the text is marked with a dotted underline. You can follow a footnote or a link to another part of the publication. Website links open in another tab, leaving the Studio open. With the keyboard, Up and Down move through the status dots and the links in the order they appear in the text, and Enter follows the link you are on.

Open Table of contents from the handle at the left edge of the text, then choose Go to page… and enter the printed page number. You can also type # followed by the number in the search box. This searches page numbers in the open chapter. If the open chapter prints no page numbers, Go to page… is switched off and gives the reason: This chapter has no page numbers.

If search finds words in source material that the open chapter cannot draw, the find bar says so and the chapter line carries the match. Search results therefore no longer appear to vanish simply because that source fragment has no visible manuscript element.

A chapter that will not open leaves you where you were

If the next chapter's content never arrives, because the connection is gone or the server refuses, the Studio used to claim you had moved while the manuscript, the take list and the overview stayed behind, and told you nothing. You now stay on the chapter you were reading, the chapter tree puts its highlight back, playback and your reading position are untouched, and the status bar says The chapter could not be opened. You are still on the one you were reading. The word chapter follows the publication: an article or a section is named as such.

Reading settings and Studio panels

View
Spoken form
A walk by the riverRecording
Book

1. The journey

2. A walk by the river

3. Going home

Go to page…

12

A walk by the river

Dr.(Doctor) Green walked along the river. The water was cold.

He stopped by the bridge.

He had walked 2 km(two kilometres).

Continuous Record Microphone
Bookmarks Post-processing Reading settings Annotations, in the editor's view
Try the choices above. Clean hides the status dots and the recording marks; Word marks hides only the dots. Spoken form adds or hides the reading instructions. In the Studio, find these choices under the gear icon at the right edge, or press V. The picture shows the rail as someone who records and manages the production sees it.

Choose what you see while reading

Under Reading settings, the new View choices are:

A separate Spoken form choice controls the reading instructions shown with the text:

Spoken form also applies to structural reading. Announcements and silence markers follow the Spoken form choice. Number announcements inside list items appear only with Full, while authored announcements appear with Authored as expected.

These choices are remembered between visits. They change what you see, not what the finished audio should say.

Find the panels at the right edge

The Studio's panels at the right edge now open from the icons there; the chapter list keeps its own handle on the left. Choose an icon to open its panel or switch to another one. You get the panels your work needs: Reading settings always, Bookmarks while you record or review, Annotations whenever your role's view allows editing, whatever the article's current stage, and Post-processing if you manage the production and are not in the editor's view. Annotations never appears beside Bookmarks or Post-processing, because it belongs to the editor's view alone.

The panel you last had open opens again the next time you come back to the Studio, as soon as your work offers that panel again. If you close the panel before you leave, it stays closed when you return. Panels stay hidden until the Studio is ready, so no empty panel frame appears while it loads.

Use the gear icon for Reading settings, or press V. Bookmarks, Post-processing and Reading settings are no longer listed in the ⋯ menu, and Annotations no longer has its own tab on the edge of the reading area. Keyboard shortcuts and the available production actions remain in the menu.

The panels at the right edge can also be pushed open or closed with a finger. Clicking the rail button and pushing the panel move it through the same open and closed states. A focus or reveal inside a closed panel no longer drags the whole Studio sideways.

The status bar floats over the manuscript

The status bar floats over the bottom of the manuscript. A new message no longer reserves a band or resizes the reading area. The manuscript has enough scroll clearance to move its final lines above the floating message.

Fewer controls in focus mode

Focus mode now hides the controls above the waveform as well as the surrounding panels and toolbars. The waveform and connection messages remain visible. Press V to open Reading settings even while the panel icons are hidden. Close settings to return to reading.

Reading instructions and editing

Manuscript

Dr.(Doctor) Green walked along the river.

How this is read aloud
How sentence 1 is read aloud
Spoken form

Doctor Green walked along the river.


Reading instructions

Read as: Doctor

Set in the editor
The document can say “Dr.” while the spoken form says “Doctor”. Open the sentence's menu to see the words and instructions together.

See how a sentence is read aloud

Choose How this is read aloud from a sentence's menu to see its Spoken form. This shows alternative words to read and lists its pronunciation, emphasis, pauses and other reading instructions.

When the information is available, instructions say whether they came from the source document, the editor or the Studio. A sentence with nothing to read now distinguishes Marked to be skipped during narration from No words to read aloud.

Words that did not match shows the expected words alongside what speech recognition heard. It also says when nothing was heard for a word.

Saved instructions come back correctly

Editing choices match what can be saved

Emphasis now offers Light emphasis, Normal emphasis and Strong emphasis. Extra stress has been removed.

Percent and Abbreviation have been removed from Reading form because those choices could not be saved correctly. Alternative reading, Spell out, Year and Currency remain.

Editing tools now explain when a sentence cannot accept an instruction. This includes text marked to be skipped, page breaks and announcements that cannot be edited here. The toolbar and sentence menu apply the same rules. If a save is refused, the message gives the reason instead of only saying that saving failed.

Recording and listening

Recording

The water was cold.

He stopped by the bridge.

Words that did not match

cold

Heard: cool

by Your place stopped A possible difference cold A difference found so far
The mark shows your place. Different underline shapes show possible and confirmed differences, and you can see what was heard.

Follow your reading as you record

During continuous recording, the Studio follows your reading and marks possible differences before you stop. It checks against the words you are meant to say, including alternative readings.

Get a useful first answer soon after Stop

The realtime model now checks every finished recording first, even when this device uses Sync after stop. It quickly marks which sentences were covered and which words may differ while the final model continues in the background. Its preliminary marks are then replaced by the final answer rather than being saved as though they were final.

Hear the whole sentence when you play it back

Playback now starts where the sound begins, not where the word was marked. Speech recognition marks the moment a word was recognised, which on the reference recording was between 0.3 and 1.5 seconds after the narrator started speaking. Playing from that mark cut the opening off every sentence, and sometimes the end too. The Studio now listens to how quiet the room is in the recording and widens the playback window outward until the sound drops back to that room level, stopping halfway to the neighbouring sentence so that two sentences read in one breath cannot replay each other. Nothing is stored or sent. The word times and the delivered audio are unchanged. Trim is once again a tool for removing silence, not for getting your own words back.

Trimming and recording again

Trimming and punch-in recording produce the same result whether the take is still on this device, is being uploaded or already exists on the server.

If a trim cannot be saved, the message distinguishes a recording that is being filed, one that has left the device and an edit that did not reach the server. The explanation says whether to wait or open the sentence and save the trim again.

Small changes while working through a chapter

Final sync and the recording queue

A recording after Stop

Final sync queued The first sentence is waiting.

Final transcription running The next sentence has been reached.

Transcribed on this device · upload still waiting

A sentence shows both parts of the work: what the speech model has heard and whether the recording has reached the server.

Continue while the accurate transcript is prepared

Stopping a recording no longer makes you wait for the most accurate speech model. SpokenWords saves the audio on your device, adds a Final sync job and lets you continue. Several recordings can wait in the queue and are handled one at a time.

See progress on each sentence

Every sentence covered by unfinished work shows whether final sync is queued, running, paused or failed. The background tint remains visible in Clean, Word marks and Status dots views. Running work moves across the sentence; people who prefer reduced motion see a still pattern instead.

A double underline means this device already holds the transcript for that sentence while the upload is still waiting. The underline disappears after the take is filed. Screen readers receive the same distinction in the sentence's status name.

The status bar reports what the queue is doing

Transcribing one of 3 recordings — about 40 seconds left

A finished take being read on this computer, with a progress bar.

2 recordings were interrupted on their way to the server

A tab was closed mid-upload. Try again restarts them.

Two of the status bar's lines. The queue says what it is doing, how long it expects to take, and what happened to a recording that was left behind.

The status bar at the bottom of the Studio summarises the recordings this device still holds. It used to remember that a sync had started and could keep spinning after the work had stopped, with no control beside it and nothing but a reload to correct it. It now asks the queue what is running every time it repaints.

Try again always answers

Pressing Try again on a queue where nothing had failed used to do nothing, not even a flicker. Every press now produces a short message saying what it found:

A press reaches every recording it is about, including one waiting out a pause between attempts or parked behind the lock a recording holds, and a page you return to with the browser's Back button runs the same recovery as a fresh load.

Recordings keep all their edits

Work through a lost connection

On this device

Audio, the words heard and word timings are available now.

Waiting for the server

The recording uploads when the connection returns.

Recognition and upload are separate. Losing the connection delays the server copy without taking away the local recording or transcript.

Speech recognition

Downloading speech recognition for this device

300.0 MB of 500.0 MB, about 2 minutes left

Reading settings
Speech recognition

On this device

Or, when a network service is used:

Sent to a network service
See the download's progress and where your speech is recognised. The amounts and time shown here are examples.

Recognition can run on your device

SpokenWords can now recognise speech on your device when your team has that service available. It prepares the required download after you sign in, even if you have not opened the Studio yet. This happens for anyone who could end up reading: a Narrator, a Producer, a Lead or an Administrator. An Editor or a Reviewer downloads nothing.

Reading settings shows whether speech recognition runs On this device or your recording is Sent to a network service. This describes where speech is recognised; recordings still need to be saved to the production.

Checking recordings also does more to reject words that the audio does not support, including words invented where nothing was said, and a word the recogniser repeated over and over instead of following your reading.

Two models, two names

Your computer runs two speech models. The realtime model follows your reading while you record and marks where you have got to. The final model reads the finished take after Stop and produces the transcript of record, replacing whatever the realtime model heard. Which published model fills each role is decided by measuring the machine, so the Speech recognition page names the model itself where its size matters, such as KB-Whisper tiny or KB-Whisper medium. Everything a reader sees uses these names, including the Final model load row on that page.

The machine decides, and prices before it downloads

The model check shows how far it has got

Testing speech models on this device — 1 of 2 tested, about 20 seconds left

The install times each model on real speech and shows how far it has got.

This computer could not install speech recognition.

Your team has a network service, so recording still works here and finished takes are transcribed there instead.

The install reports its progress, and a computer that cannot finish it says so instead of letting the progress bar quietly disappear.

After the download, the install times each model on this machine. The status bar entry reads Testing speech models on this device — 1 of 2 tested, about 20 seconds left, with a progress bar. The bar shows the share of the work behind you rather than a count, because one model takes most of the time and a count would stand still and then jump. The cheapest model is checked first, so the first reading lands within about a second and prices the rest. The time estimate appears only after that first reading, and if a check runs past its estimate the estimate is withdrawn rather than counted past zero. Download, model check and transcription all word their remaining time the same way: whole minutes for a long wait, five-second steps for a short one.

A computer that cannot run speech recognition says so

A machine that could not build its speech model used to tell nobody. The download bar simply vanished, exactly as it does after a successful install, and the Record button then sent a narrator whose team had a perfectly good network service to ask an administrator for one. Three sentences replace that silence:

See how much feedback your device can give

Following a reading word by word asks more of a computer than some can manage, so the app measures this device and sets Speech feedback in Reading settings to what it found:

The installation only estimates the level, from the check it runs while preparing. Your own reading measures it properly: once you have recorded for a while, the app times its work on your speech and sets the level from that, so the level shown — and which levels are switched off — can change from what the installation estimated. You can choose a setting below the measured one, and the levels above it are switched off with the reason: This device was measured as too slow for this level. The choice applies to your next recording. It changes nothing about what was installed, because the installation follows the measurement.

Beside it, Cloud sync decides where the finished recording is checked: off, on this device; on, at your team's network service. It appears only when your team has such a service, and it is switched on and fixed there when speech recognition cannot run on this device.

Saving, connection problems and signing in

A walk by the river: no speech was recorded, so nothing was saved.

The article was imported again — this recording cannot be used

The recording was kept on the server, with a copy on this device. Record the passage again.

Stopping without speaking is reported as ordinary news, and asks nothing of you. Audio kept after an article was imported again is a warning, because you have to record the passage over.

Know whether a recording produced a saved take

The app no longer reports that everything was saved when a recording produced no take. It distinguishes a recording with no speech from one that could not be matched to the manuscript. When needed, it asks you to read the passage again.

A sentence no longer stays marked as saving for the rest of the visit after such a recording. If the app cannot check whether recordings are waiting to sync, it says so instead of claiming that syncing has finished.

Keep a recording with the version you read

If someone imports an article again while a recording is waiting to be saved, the app does not attach that recording to the new sentences. It can keep the audio on the server, with a copy on your device, and tells you to record the passage again. This prevents an older recording from being mistaken for work on the new import.

Stay signed in

A connection failure no longer signs you out simply because the app could not check your sign-in. It keeps your sign-in information and tries again. Connection messages now appear on every page instead of only in the Studio, and tell you when the server cannot be reached and when the connection is back.

A tab left open in the background no longer signs you out of the tab you are working in. When such a tab's own session ran out, it used to discard the sign-in information for this browser, including the sign-in the tab you were using had just renewed, so your changes stopped being saved there. It now discards only the sign-in it was still using itself.

Pages open on the sign-in this device already holds. Every page used to spend a round trip to the server confirming your session before anything else could start. A session that is stored whole, matches and has more than a minute left is applied at once. Anything doubtful is still checked with the server, and the server keeps the last word: a session it has ended still sends you to sign in on the first real request.

When a request never reaches the server at all, the message now reads The server could not be reached. Check your connection and try again. It used to name what failed and then end with the browser's own error text, such as “Failed to fetch”. This applies wherever the app reports a failure, not only in the Studio.

Clearer messages when a change is refused

Saving messages are clearer about changes that were refused or could not be stored, including reading position and sentence edits. These arrive over a working connection: they are the server's answer when a change could not be written. The English and Swedish messages now describe these outcomes more consistently.

Productions and teams

Most of this section is for the people who run a production. If your part is recording, editing or reviewing, you will not see the import and export controls described here.

A walk by the river

This article uses the speech settings saved when it was imported. The product's settings have since changed. Existing recordings are unaffected; the saved settings apply to sentences read by text-to-speech. To use the current settings, choose Import again. This deletes its recordings, reviews and bookmarks.

The gear marker on the article's row is named Uses earlier speech settings. Resting on it explains what importing again would delete.

See why an issue cannot be imported

The issue list now shows when an upload is waiting, running or has failed. Issues with unfinished uploads cannot be selected for import. The list also explains when nothing has been uploaded or there is no article content to import.

If an import is refused, the message now names the real reason. A refusal used to say the issue was already being produced even when speech was still being generated. An issue whose upload has not finished is now refused with that reason as well; before, such an issue could be imported with only the articles that had arrived so far. The reason also stays in the notification list after the message disappears.

See when speech settings have changed

If you manage the production's articles, a small gear marker now appears on the article's row before you export it. Its name is Uses earlier speech settings, and resting on it explains what has changed. The same information is reported when an export uses the earlier settings.

This means the product's speech settings have changed since the article was imported. Existing recordings are unaffected. Sentences read by text-to-speech still use the settings saved at import.

To use the current settings, choose Import again. This deletes the affected article's recordings, reviews and bookmarks. The message now explains this consequence.

Recording controls match the assignment

Only a narrator or producer assignment can own a recording. The Studio no longer offers Record to someone whose other role cannot support the take, and its explanation says that recording needs a narrator or producer assignment.

When an eligible lead or supplier administrator reaches an unstaffed chapter, the status bar can offer Assign me as producer. A producer can record and can staff the chapter, so one assignment unlocks both jobs. An assignment added by somebody else reaches an already open Studio without requiring a reload, and device speech recognition starts when the new right to narrate arrives.

Supplier staffing follows the granted work

Production information explains itself

Article rows now explain states such as no work, recording, rejected and approved. Production stage symbols describe what Draft, Recording, Review and Done mean rather than repeating the short label.

Content type is now a read-only field in Production details because it comes from the source product. The help text names that product when it can and directs a manager to correct the product if the value is wrong. The ambiguous clock date has left production cards; Last updated is named in the details view instead. A declined assignment is also visible on the production card without opening it.

Resolve problems without searching for the right screen

Production messages use the publication's familiar terms more consistently: chapters for a book, sections for a document, features for a magazine and articles for a newspaper.

Keyboard controls and messages

Advanced and help

Check speech recognition

The account menu's Advanced choices now include Speech recognition. It shows download and checking progress, what is available on this device, and why speech recognition may be unavailable.

The page opens with the answer: a few tinted sentences saying what this machine will do, such as Marking the reading: KB-Whisper tiny, on the processor and This machine has no graphics adapter, followed by Expected wait after Stop with the estimate for each candidate. The expected wait itself was corrected from about 3.4 times the recording to about 1.2 times, after the two figures behind it were measured on the same recording. The tables sit in sections you can open: Models, Engines, and what they report, Deployment checks, What this device remembers and Try the realtime recogniser. Server checks that passed are folded away with the browser details, so the ones that failed are easier to find, and a failed deployment check opens its section by itself, so it is never hidden behind a closed heading.

Once the installation has finished, choose Init model to start speech recognition on this device. The recording controls then appear: choose Record, or Play a file, set the Language — it starts on the language the app is shown in — or choose Detect, and watch the words appear. After stopping, Decode the take shows what the finished recording was heard as, and Save the take downloads the recording to your computer as an audio file. These controls let you try speech recognition without recording into a chapter.

If installation checks failed, Retry verification lets you try again after the cause has been corrected. It reloads the page and clears the saved failures. It keeps the downloaded files and successful checks.

See files kept on the device

Advanced also includes Private file system. It lists stored files, their sizes and whether the expected download is complete. You can refresh the list, delete a file or choose Delete everything, with a confirmation before deletion.

Deleting speech recognition files means they will need to be downloaded again. This page explains that signing out or clearing the usual browser cache does not remove these files.

Reproduce queue states on demand

Administrators and testers get two rows in the account menu under Advanced. They exist to reproduce the queue states described under Final sync and the recording queue on demand, so a problem report can be checked against a known state. They are troubleshooting instruments, not part of the recording workflow.

Clearer guides

The guides have been corrected against the controls this release moved or renamed, and use simpler wording. They explain the production card's Production roles and Production details choices, the panel rail that now opens Reading settings, the floating status bar, the current recording controls, the read-only content type, and which editing instructions return after reopening a chapter.

Read the Overview, Editor's guide, Narrator's guide, Reviewer's guide, Lead's guide, What a producer manages or Administrator's guide for help with your role.

Behind the scenes

You do not need to learn new controls for these changes.