Monday, August 17, 2026

Needle

★6,815Must watch
Demo

A 14MB foundation model that does tool calling and structured extraction on phones, wearables, and robots.

Why it matters

Needle 2 is an open 45M-parameter Simple Attention Network from Cactus Compute, compressed to about 14MB with Cactus Quants and baked into its own engine. A full session runs near 28MB of RAM. It is built for tool calling, device control, and structured extraction, with a learned confidence head and a sliding window that keeps memory bounded no matter how long the conversation runs.

Most on-device models still chat. Founders shipping voice, home, and robot products need tiny models that call tools with JSON they can trust, stay offline, and fit in tens of megabytes of RAM. Needle turns that into a single pip package and a single binary engine.

How it works

pip install cactus-needle, decorate Python functions as tools or pass schemas, then run Needle so it picks calls, executes them, and returns structured results. For extraction, declare a Pydantic model or JSON schema and call extract(). Optional LoRA fine-tune plus export still lands as one .cact binary. A local playground UI is one command away for trying tools and prompts.

This is not another giant local chat model or streaming MoE trick. It is a purpose-built edge foundation model for tools and structured IO, grammar-constrained decode from your schemas, confidence gating, and tool retrieval over large catalogues while staying under a phone-class memory budget.

Capabilities

Demo
  • Public demo available

Similar tools

on-device-aitool-callingedgelocal-llmstructured-output
Source ↗

Via github

X