Every library has a file with no subtitle track: an old rip, a home recording, a release that shipped without one. When somebody in the room needs subtitles for it, what do you do? Usually you search a couple of subtitle sites and end up with a file that’s out of sync by a second or two.
Quven can write one for you instead, from the dialogue, on your own machine, and it should take a few minutes. It isn’t the best subtitle you can get; an authored track written by a person will almost always read better. It is a real, timed, reusable subtitle for a file that otherwise has nothing.
This needs Quven 1.1.0 or later. On an earlier release the option described below isn’t in the player, and no setting will bring it forward.
Install the speech model once
Recognition runs on the server, so the model has to live there too. It isn’t shipped inside the download: it’s close to half a gigabyte and most libraries will never need it, so Quven fetches it on request instead of making every installation carry it.
On the desktop client and in the hosted web client a server administrator installs it explicitly. Open the subtitle controls during playback and the install action will be there when the model is missing; the panel reports its own percentage and megabytes while it downloads, verifies what it fetched, and discards the file if the verification fails. On a normal connection it’s a matter of seconds, not an afternoon. An ordinary viewer is told that an administrator has to install it, so they know who to ask.
The phone and tablet clients, and the television ones, carry nothing to do with the model at all. That’s by design, because those screens are for watching and the server is managed elsewhere. They offer translation of a subtitle that already exists, and of the speech model they show neither status nor prompt nor install path. Only an administrator can ask for a transcription, from the desktop or the hosted web client; while one runs, a phone or a television sees the job and uses what it produces like any other track.
Ask for subtitles from the player
Start the title, open the subtitle controls and choose Generate or translate subtitles; you’ll need to be an administrator for the generate half. Quven inspects the file first and shows you what it found.
If the file has any usable subtitle source, that’s what you’ll see offered under Subtitle source, with each entry labelled by language and kind: plain text, an image track that needs recognition, a sidecar file, or something Quven generated earlier. Leave it on Automatic (best available) unless you have a specific reason not to. Forced tracks and signs-only tracks are deliberately kept out of that list, because they aren’t a script.
When there genuinely is nothing usable, the panel says so and offers the audio instead. You should pick the dialogue audio stream carefully. A film can carry a dub, a commentary, an audio description and the original language, and the subtitle you get will faithfully match whichever one you chose. The spoken language can be left on automatic detection; setting it explicitly can be worth doing when the track is short, noisy, or bilingual.
Leave the player open while it works
The job belongs to the player. Quven tells you what stage it’s in and roughly how long it expects to take: around five minutes when it can start from a text track, ten when images have to be recognised first, twenty when it’s working from audio. Those are budgets. They aren’t promises, and a long film on a slow machine will use all of them.
The job belongs to the open player, on every client. Leaving will cancel the work, and Quven warns you and asks you to confirm before it lets you navigate away. This is deliberate. An unattended job that finishes into an empty room tends to be a job nobody notices has failed. The subtitle menu is the only place the request starts, and the progress dialog stays inside the player and won’t follow you around the app.
Underneath, the work is durable on the server. It survives a server restart, two people asking for the same thing share one job, and a retry gets its own identity instead of resurrecting the old one. So don’t queue the same request three times because the first one looked slow: you’ll simply be waiting on the same job.
On a machine with a supported graphics card the recognition step runs there instead of on the processor, which is both much faster and much less rude to whatever else you were doing. Without one it falls back to the processor and still works, just more slowly, and it deliberately leaves you a core so the machine stays responsive.
What you get, and what to do with it
The result is a normal subtitle asset, not an overlay drawn over the current frame. It’s stored against that specific physical file, it appears in the subtitle menu, and it will still be there when you open the title again next week. If the job fails or you cancel it, nothing half-finished is left attached.
A word about quality. Recognition sets the ceiling: what comes back is usually well timed and readable, and its mistakes are the awkward kind: a line that’s perfectly grammatical and simply not what the actor said. For a film that has a proper subtitle track available somewhere, that track is still the better choice. For the file that has nothing, this is the difference between watching it and not.
Generation runs entirely on your server and stays there. The audio, the decoded samples and the recogniser’s working files never leave the machine, and the temporary ones are cleaned up whether the job succeeds, fails or is cancelled. That’s also why it costs nothing to use: it’s your hardware doing the work.
Translating that subtitle into another language is a separate step with a different shape, because the text does leave the machine. That one is covered in the guide on translating subtitles with AI. If subtitles show up on a particular device but render oddly, the native playback guide is the better place to start.
Frequently asked questions
Does my audio leave the machine?
No. The server transcribes the dialogue on your own hardware and saves the result as an ordinary, reusable track. The audio itself is never sent anywhere.
Which release do I need?
Quven 1.1.0 or later. On an earlier release the option is simply not there, and no library setting brings it back.
Why can’t I install the speech model from my phone?
By design. Recognition runs on the server, so the model lives there and a desktop or hosted-web administrator installs it explicitly. Phone, tablet and television clients carry nothing about the model at all, though they will happily play a track generated elsewhere.