Blog
Documenting on-device science and engineering.
Getting the Most out of Whistle
A practical guide to using Whistle.
Jakub Mroz · · 4 min read
Whistle: Speech to Text in 16.9 MB
An open speech recognition model that runs on the same CPU engine as Needle. It transcribes seven languages, reaches the first token in 11 ms, and loads beside Needle so one binary turns a clip straight into tool calls.
Jakub Mroz, Henry Ndubuaku · · 6 min read
Cactus x Raspberry Pi
Raspberry Pi wrote about running Needle on a Pi 5: a 14 MB function-calling model that turns plain English into local Python actions, entirely offline on the CPU.
Raspberry Pi · · 6 min read · raspberrypi.com
Where an Attention Head Spends Its Bytes
We shrank the query-key path and the value path of Needle's attention heads to the same file sizes and measured the loss. Bytes taken from the values cost up to three times more than bytes taken from the routing, and widening the values alone is the cheapest capacity we have found.
Karen Mosoyan · · 8 min read
Can We Trade MLPs for Engrams?
Needle keeps its facts in hashed n-gram tables instead of feed-forward layers. 70.8M of its 121M parameters live there, and a token reads 30 rows of them and multiplies none. This post is the lookup, the gate, where the sites sit, and what the tables buy.
Henry Ndubuaku · · 12 min read
The .cact Format
Needle ships as one file that the engine maps into memory and reads in place: a 196-byte header with the whole architecture, a nameless tensor directory, and Cactus-Quantised blobs at 2.125 bits per weight. What is in it, how the quantisation works, how the kernel reads it without ever unpacking, and how to parse it yourself in twenty lines.
Parkirat Sandhu · · 9 min read
How to Design Tools for Needle 3
Needle reads a schema literally and copies every argument from the request, so the toolset is the product. The rules we learned shipping Needle 3: one tool per action, names users would say, formats in descriptions, constraints in the grammar, triggers for the phrasings a description cannot list.
Satyajit Kumar · · 10 min read
Fine-tuning Needle
Two ways to fine-tune Needle 3 from one package: LoRA on your machine, or the full model on the Cactus Platform with every depth trained and scored. The data format, the commands, how to read the loss, how much data a task needs, and what each path changes.
Jakub Mroz, Roman Shemet · · 9 min read
Leveraging Needle's Confidence
Every Needle response carries a calibrated confidence score. What it measures, what the engine already does with it, and how to turn it into a product decision: act, confirm, or refuse.
Karen Mosoyan · · 7 min read
Needle Python Docs
The reference for the cactus-needle package: the Needle class, tools three ways, the response shape, the behaviour contract, system facts, tool retrieval, tuned weights, offline devices, environments, the CLI, and the mistakes to avoid.
Roman Shemet · · 12 min read
What Devices Are Supported on Needle
Thirteen platform folders, one engine under 1 MB each, one set of weights at any depth from 2 to 20 layers. Which folder is your device, what ships in it, how to run it from the command line, C, a browser or a WASI host, and how to set it up with no network at all.
Justin H. Lee · · 8 min read
Porting Needle 3
Notes for anyone writing their own Needle 3 runtime, from the first community port: which oracle to test against and how to read the numbers, the tensor order that the .cact container promises, what the prompt looks like on the wire, how the ladder picks blocks, and how retrieval works while the release ships without a contrastive head.
Henry Ndubuaku · · 10 min read
Structured JSON Extraction with Needle
Extraction in Needle is tool calling with one tool: declare the record, pass the text, get a typed object whose every field is a span of the passage. How the grammar guarantees the shape, how grounding keeps values honest, and how the same trick does classification.
Noah Cylich · · 8 min read
The Hadamard MLP for Channel Mixing for Almost No Parameters
Needle replaces the transformer's feed-forward layer with three Kronecker-factored mixes that start as the Walsh-Hadamard transform and learn from there. 25.6K parameters per layer instead of 4.7M, and a fifth of the compute per token.
Henry Ndubuaku · · 13 min read
Intelligence Ladders: One Set of Weights, Every Depth a Model
Needle 3 is trained so that any depth from 2 to 20 layers is a deployable model. This is how the subnetworks are chosen, how they are trained without training five models, and what each one scores.
Henry Ndubuaku · · 12 min read
What a Transformer Loses Without Feed-Forward Layers
We deleted the feed-forward network from a transformer and measured what was actually lost, under controls for parameters, compute and depth. At matched parameters the answer is 0.006 nats, and all of it lives on tokens with nothing to look up.
Henry Ndubuaku · · 12 min read
How to Design Tool Environments for Needle 2
A practical guide to tool names, descriptions, enums, and execution rules for reliable function calling with Needle 2.
Roman Shemet · · 8 min read
Needle 2: The 14 MB Agentic LLM for Tiny Devices
An open 45M-parameter model for tool calling, device use, and structured extraction. Needle 2 runs as a 14 MB binary in 28 MB of session RAM.
Henry Ndubuaku · · 14 min read
We Distilled Gemini Tool Calling into a 26M Model
An open-source 26M parameter function-calling model that runs at 6000 tok/s prefill and 1200 tok/s decode on consumer devices.
Henry Ndubuaku · · 3 min read
