CactusCactus
NeedlePlatformResearchBlog
4.2k+ starsTalk to us

[NEW]Needle 3: our 8-29 MB foundation model for tiny devices

Research

The papers behind Needle and the Cactus engine: small attention-only models, quantisation, sparsity, tokenisation and training.

  • A Controlled Study of Attention-Only Transformers

    Henry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee

    arXiv · 2026

  • Parameter-Efficient Transformer Embeddings via Functional Factorization

    Henry Ndubuaku

    ICML · 2026

  • Depth Over Specialization in Small Multimodal Transformers

    Jakub Mroz, Henry Ndubuaku

    ICLR · 2026

  • Calibration-Aware Activation Sparsity for Instruction-Tuned LLMs

    Noah Cylich, Karen Mosoyan, Henry Ndubuaku

    ICML · 2026

  • A Blazing Fast LM-Head Replacement

    Satyajit Kumar, Henry Ndubuaku

    ICLR · 2026

  • Token-Aware Chunked Encoding

    Parkirat Sandhu, Henry Ndubuaku, Justin Lee, Satyajit Kumar, Noah Cylich, Karen Mosoyan, Jakub Mróz

    ICLR · 2026

  • Just Enough Learning: GRPO-Guided Controllers for Hyperparameter Sweeps

    Justin H Lee, Henry Ndubuaku

    ICLR · 2026

CactusCactus
Hugging FaceGitHubCactus Platform
TermsPrivacy

© 2026 Cactus Compute. All rights reserved.