Cactus Hybrid
Cloud accuracy without the cloud cost

Post-trained models that know when they're wrong and request help from the cloud.

Cactus routes agent commands based on complexity: on-device for simple tasks, cloud for complex operations.

Command

Set the thermostat to 72 degrees

Cactus Hybrid Router
On-Device
Cloud
Complexity
Output
Waiting for command...
Intelligent routing for function calls
Try the demo
$brew install cactus-compute/cactus/cactus
$cactus run
Gemma 4 E2B Hybrid

The smallest Gemma 4 model matches Gemini 3.1 Flash-Lite on a given task by handing off the following share of queries:

BenchmarkFP164-bit
ChartQA vision15–20%25–30%
MMBench vision30–35%40–45%
LibriSpeech audio25–30%35–40%
GigaSpeech audio30–35%40–45%
MMAU audio30–35%35–40%
MMLU-Pro text45–55%~90%
Lower score (%) means fewer cloud calls.Learn more

Router in the Weights

We post-train a router inside the model weights, which reads internal activations and calculates a confidence score for each task.

Runs on Your Existing Stack

Run natively with the Cactus Engine, as well as llama.cpp, MLX, or Transformers.

Beyond Gemma

The post-training recipe extends to other models — as well as to your proprietary weights. Talk to us.

No compromise

Get the best of both on-device and cloud.

Traditional Cloud AI
Cactus On-Device
Cactus Hybrid
Sub 150ms Latency
Handles Noisy Audio
Works Offline
Data Privacy
Cost Efficient
Smart Routing

Ready to get started?

Add hybrid inference to your app in minutes. Free to start, scales with you.