Skip to content

The Field Guide · product · May 11, 2026

How we chase the last few points of transcription accuracy

A word-error rate you'd forgive in a voice memo is fatal in research, where the difference between 'can' and 'can't' is the roadmap. Inside our accuracy pipeline.


In most transcription products, a 4% word-error rate is a footnote. In research it’s a trap, because the errors aren’t random — they cluster exactly where the signal is.

Interviews are full of product names, competitor names, half-finished sentences, and the single syllable that separates “we can work around it” from “we can’t work around it.” Generic speech models were trained on none of that. So a transcript can read fluently and still be wrong in the five places your report will quote.

Here’s what we do about it.

Vocabulary is workspace-local

Every Fieldnote workspace builds its own vocabulary from the material you bring: product names, feature names, the acronyms your customers actually use. The model is biased toward that vocabulary at decode time, which is why your beta feature’s made-up name survives transcription in week one.

Negations get special treatment

We run a second pass over high-stakes constructions — negations, hedges, comparatives — the places where a one-word error flips the meaning. When confidence is low on a “can/can’t”, we flag it in the transcript rather than guessing silently. A visible uncertainty beats an invisible error every time a human quotes the line.

Speakers are separated before words are finalized

Cross-talk is where transcripts quietly fall apart: two people, one channel, and suddenly the customer is credited with your interviewer’s opinion. We diarize first and let speaker boundaries inform the language model, which keeps “that would be great” attached to the person who actually said it.

Corrections teach the workspace

When you fix a word in Fieldnote, that correction joins the workspace vocabulary immediately. Teams tell us transcripts feel noticeably sharper after the first dozen interviews — that isn’t the base model improving, it’s your workspace learning your world.


The honest summary: we can’t promise perfection, and anyone who does is selling you a voice memo app. What we promise is that errors get rarer where they matter most, visible where we’re unsure, and easier to fix every week you use it — because a research tool is only as good as the quotes it lets you stand behind.