The biggest model is rarely the best tool for the job. A smaller model that is shaped to your work is usually faster, cheaper, and just as productive.
You would not hire a Nobel laureate to file invoices. They could do it, but they would charge a fortune, while someone specifically trained for the job would be faster.
Frontier models answer any topic from anyone. That breadth of data costs a lot, and most of it is irrelevant to the work a business repeats daily: reading claims, checking contracts, sorting tickets. We can teach a smaller model that one job it will match the giant for a fraction of the cost.
of parameters: trained to do everything for everyone. They are chasing artificial general intelligence, so they want to take in every scrap of information we expose.
They runs on thousands of racks in data-center accelerators. A single dense rack can draw as much power as an 80-house neighborhood.
of parameters that are trained to do your job.
This runs on an NetApp® AIPod Mini in your data center or your own office. Small enough to run on a handheld too.
Circles suggest the gap, not to scale. Roughly 60–300× lower capital cost for the same outcome.
Small enough to fit in a hand. Iterate runs retrieval — searching a company's own documents — on handhelds carried by mine workers underground and waitstaff on the floor. No mains power. No internet. The model is on the device.
Nobody builds a small model by starting small. They start with a big one and shrink it. There are five core techniques do that. Each method brings significant and unique advantages:
A big model teaches a small one. The small one watches how the big one answers thousands of questions. Then it learns to answer the same way.
Like: an apprentice who trains beside a master, then works alone.
Cutting out the parts that never do anything. Most of a big model sits idle on any one job. Those parts are removed.
Like: trimming dead branches. The tree stands, and it is lighter.
Storing each number less precisely. Use 3.14, not 3.14159265. The answers barely change. The model gets much smaller and quicker.
Like: rounding prices to the nearest dollar. The total still works.
Training a general model further on your own records. It picks up your words, your formats, and the way your business decides things.
Like: an experienced hire spending a month learning how you do it here.
Joining trained pieces of two or more models into one. You keep the strengths of each, without training something new from scratch.
Like: building one team out of the best people from two departments.
The models running on an AIPod Mini are right-sized ones, shrunk by exactly these techniques. That is why a box that fits in your office does work that used to need a data centre — and why the model, and everything it learns, stays yours.
Ask what the model has to do. Then buy the smallest one that does it well. Bigger is a cost, not a feature. And the small one is the one you can afford to own. A small model that is fine-tuned specific to the use case is often much more accurate.
Full paper, 12 pages — iterate.ai/partners/netapp/papers. On the wider cost argument, see The End of “Per Token” Costs; on ownership, What Private AI Actually Means.
A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.
