Back to blog
Guides

Getting the Most out of Whistle

A practical guide to using Whistle.

JM

Jakub Mroz

||4 min read

Whistle is 16.9 MB of speech to text on the same engine as Needle. If you are interested in the model and its benchmarks, visit the launch post.

Input

Whistle takes 16 kHz mono float samples in [-1, 1], or a WAV path, up to thirty seconds in one call. Longer audio is refused, so split it or use the streaming API below.

Output

It returns the transcript, the detected language, the time to first token and the decode speed.

import needle

result = needle.transcribe("clip.wav")
print(result["text"], result["language"], result["ttft_ms"], result["decode_tps"])

Word timestamps

word_timestamps=True adds a words list beside the transcript, one entry per word.

{"word": "bedroom", "start": 0.48, "end": 0.80, "probability": 0.997}

The times come from the decoder's attention over the encoder frames. Their resolution is 80 ms.

Live transcription

needle.stream takes chunks and transcribes as they arrive. Every chunk re-transcribes the buffered audio and commits the words two consecutive passes agree on, so committed text never changes under you.

for step in needle.stream(chunks):
    print(step["text"], step["pending"], step["pass_ms"])

Each step carries the words it committed with times from the start of the stream, the unconfirmed tail, and the milliseconds that pass took.

Language

Whistle supports seven languages: English, German, French, Spanish, Italian, Dutch and Polish. The language is one token, chosen before the transcript starts, and the rest of the transcript decodes based on it.

needle.transcribe("clip.wav", language="de")

Leave detection on when the input is genuinely mixed.

Keywords

Keywords are words and phrases you pass with the clip, and Whistle lifts their probability while it decodes, so a name it has never heard can compete with the ordinary words that would otherwise win.

Words on the list
B-WER, percent
no list18.43
100 keywords4.46
1,000 keywords5.04
Every other word
U-WER, percent
no list3.05
100 keywords2.74
1,000 keywords2.84
All of them
WER, percent
no list4.73
100 keywords2.93
1,000 keywords3.08
LibriSpeech test-clean, 2,611 utterances, with the per-utterance biasing lists and the scorer from Le et al. (Interspeech 2021), the paper that defines B-WER and U-WER.

A hundred keywords cut the error on the words in the list by three quarters and left every other word alone.

List the words your users say that most people do not: contact names, rooms, devices, stations, product codes, menu items. Common words gain nothing.

contacts = [person.name for person in address_book]
transcription = needle.transcribe("clip.wav", keywords=contacts)["text"]

A phrase counts as one entry, so living room is matched whole. Each entry is biased in three casings, so GPS also covers gps and Gps.

Depth

The decoder is an intelligence ladder: every depth from two layers up was trained as a model of its own, they all live in the same file.

Pick a depth for the budgetevery depth from 2 to 8 decoder layers is a trained model inside the one filedepthword error ratedecode8L4.751,182/s7L4.991,307/s6L5.071,430/s5L5.301,597/s4L5.591,799/s3L6.522,033/s2L7.662,557/s
Word error rate over all 2,611 LibriSpeech test-clean clips. Speed is the median of 100 runs on a single 10s clip that decodes to 100 tokens. All on an Apple M4 Pro. Depth 8 is the full decoder.

Whistle and Needle

Running both

First transcribe the clip with Whistle. Then pass the text to Needle to get the calls.

transcription = needle.transcribe("clip.wav")["text"]
action = needle.Needle(tools=tools).complete(transcription)

A mis-heard word is a missing call

Needle takes argument values straight from the transcript. If a word comes out wrong, there is nothing left to match the argument against, and no call comes out.

Use your schemas as keywords

Your tools already contain most of the words people will say. Use them as keywords. Take the enum values, the tool and property names, and the words in the descriptions. Write each one the way people say it, so living_room becomes living room.

Pick keywords from your schemawords in the description{ "name": "set_lights", "description": "Turn the lights in a room on or off.", "parameters": { "properties": { "room": { "enum": ["kitchen", "living room", "bedroom"] }, "state": { "enum": ["on", "off"] } } } }keywordskitchen · living room · bedroom · off · set · lights · roomstate · Turn · the
Take the values, names and descriptions from your tools. Write them the way people say them. Pass the list with every clip.
keywords = [
    "kitchen", "living room", "bedroom", "off",  # enum values
    "set", "lights", "room", "state",            # names
    "Turn", "the",                               # description
]
transcription = needle.transcribe("clip.wav", keywords=keywords)["text"]
action = needle.Needle(tools=tools).complete(transcription)