Cactus Releases Needle 3, Tiny On-Device AI Models It Says Can Beat DeepSeek V4 Flash on a Narrow Task After Fine-Tuning
The 8–29 MB models turn plain-English requests into app function calls on phones, robots and wearables. Cactus's DeepSeek comparison is for versions fine-tuned on a phone-command benchmark, and early testers reported mixed results, with indirect requests often failing.

The takeaway: Cactus has released Needle 3, a family of very small AI models, which it describes as 8 to 29 MB binaries (another part of its page says 9 to 29 MB), built to turn plain-English requests into app function calls and structured data on phones, wearables, robots and microcontrollers. Cactus claims that, once fine-tuned on a specific task, versions as small as its 4-layer, 29M-parameter subnetwork pass DeepSeek V4 Flash on the DroidCall phone-command benchmark. Early testers on Hacker News reported mixed results with the demo; Cactus agreed in the thread that the model works best with direct language.
What Cactus says Needle 3 does
According to Cactus's launch page, Needle 3 trades general chat ability for three jobs:
- Tool calls: given the functions an app exposes, pick the right ones and fill in the arguments. Cactus says a request that no tool covers should return an empty list, "not a guess".
- Structured extraction: turn messy text, such as an invoice or a notification, into typed fields.
- Text embedding: return a vector for a sentence, for on-device search and matching.
The model is built as what Cactus calls an "intelligence ladder": one set of weights in which every depth from 2 to 20 layers works as a model of its own, so a developer can pick a smaller slice for a smaller device. Cactus lists the full model at 121M parameters, trained on 360B tokens of what it calls a "proprietary structured dataset". It quotes decode speeds of 400 to 4,000 tokens per second on a Raspberry Pi 5. The code on GitHub is under the Apache 2.0 licence.
The DeepSeek comparison, and what it actually covers
The Show HN post's title says the models "can match DeepSeek V4 Flash"; its body says this means fine-tuned performance "on a narrow task with just 4L, stress on "narrow task"". Cactus's page gives the detail. It says the shipped model, before fine-tuning, beats models 10 times its size on mobile tool calls and matches models two to three times its size on extraction; the Show HN post adds, of those results, "we do not win everywhere ofc". The DeepSeek comparison applies after fine-tuning: Cactus says fine-tuning on the DroidCall benchmark "lifts every subnetwork by 18 to 36 points", and that "from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash". Its chart also shows a second phone-command benchmark, Mobile Actions, and a separate line on the page says a 4-layer model "can match DeepSeek V4 Flash when tuned on downstream tasks for one epoch".
In our reading, that is a claim about narrow, task-specific accuracy, not general ability, and Cactus frames it the same way: "Constraining the capacity to a narrow, well-defined task is what lets it reach frontier-level accuracy there." The benchmarks are Cactus's own runs, and we haven't reproduced them.
What early testers found
The Hacker News thread for the launch had mixed reports, mostly from people trying the demo with smart-home style commands:
- One user wrote: "None of the queries I asked worked", listing requests such as "more light" and "both doors should be locked".
- Another said "turn all the lights on/off" and "it's too dark in the bathroom" worked, "but anything less direct didn't". They said confidence on the bad responses was "pretty low", though one wrong thermostat change came with high confidence.
- A third reported, "from my limited tests", that "it can work with up to 10 tools/definitions. over that and it gets confused".
- Others had better luck: one said "Illuminate (roomname), de-illuminate (roomname)" works well; another got "a beautiful JSON doc" from a two-part request.
- One tester who fine-tuned Needle 3 on their own tool-calling task reported 20.4% exact-argument accuracy with Needle 3 quantised to 4-bit weights, against 85.2% for a fine-tuned FunctionGemma at full BF16 precision, while calling it "a solid improvement over Needle 2". Cactus asked for details of the fine-tuning setup.
- A broader critique: "the growing number of dubious claims that a tiny model beats LLMs will make any useful innovation be overlooked", with a request for Cactus to spell out what the model can't do.
The Show HN submitter, writing for Cactus, replied that the demo is "just a "get started" preset" whose tools can be edited, updated several tool definitions during the thread, and agreed that "the cleanest use cases involve direct language". He also explained the size range: "The first 4 layers alone are 8 MB, all 20 are 29 MB."
These are informal tests, not benchmarks, but they match Cactus's own advice that describing tools well "is the whole game" and that developers should keep the toolset per turn small.
Why it matters
Running small models entirely on a device means commands can work offline, with low latency and without sending data to a server. In our reading, Needle 3 is a bet that most of those commands don't need a general-purpose model at all, just a small one tuned hard on one product's tools. Cactus's own fine-tuned results, which we haven't reproduced, support that for narrow tasks. The early hands-on reports suggest the demo still struggles with indirect phrasing, something Cactus acknowledged in the thread.
Ask Relay — he reads every question himself and replies personally by email.
