Research Notes

What Our First Sparse Model Taught Us About Local AI

Our latest CFI experiments show why using fewer computations is not the same as delivering faster, more capable AI on consumer hardware.

Bryce CourtneyLocal AI / Consumer Frontier Intelligence / Sparse Models / Research
close up circuit board

Artificial intelligence does not become practical on consumer hardware simply because an architecture performs fewer calculations on paper.

The calculations have to preserve useful behavior. The system has to run faster or use fewer resources when measured. And any apparent advantage has to survive comparison with a simpler alternative.

Those distinctions have shaped our recent work on Consumer Frontier Intelligence, or CFI.

We began by studying where computation might be reduced in an existing model. More recently, we built and trained a small, independent CFI model to test sparse expert routing directly. Neither track has produced a frontier-capable local model. Both have made the engineering problem clearer.

A promising local signal can disappear during generation

In our earlier experiments, we examined the later layers of a frozen Qwen3-4B-Base model. Some interventions caused relatively small changes when we evaluated individual positions against a fixed reference. That made parts of the network worth investigating as possible targets for reduced computation.

But generation is not a sequence of independent tests. Each new token becomes part of the context for the next one. A small difference can change the path the model follows.

One experiment made that distinction especially clear. We kept the model’s layer-36 attention and its key-value cache intact, while replacing only that layer’s MLP output with a compact auxiliary predictor. On held-out, teacher-forced traces, the predictor moved the output distribution closer to the full model than simply zeroing the MLP did. During free-running generation, however, the replacement diverged from the full model’s token sequence early. The measured runtimes did not establish a dependable speedup. The cache-safe reconstruction experiment records the setup, failures, and measurements.

A later scorable evaluation added an important nuance: on three selected questions, the full model, the predictor replacement, and the zero-MLP control all reached the same final numeric answers, despite taking different token paths. That result does not prove the predictor preserves general behavior. Nor does a changed token path, by itself, prove those particular answers were lost. Three questions cannot settle either claim.

The lesson is that a useful intermediate metric is a reason to investigate further—not permission to announce an acceleration.

Our first independent model exposed a different problem

The Qwen experiments trained only auxiliary components; they did not train or fine-tune Qwen itself. To test a CFI-defined architecture rather than a replacement inside another model, we then trained a small causal model from scratch with sparse expert dispatch and a paired dense control.

That first independent-model run was an engineering smoke test, not a useful language model. The model trained, saved, reloaded, and generated without Qwen weights. It also showed a clear failure: the router sent every held-out position to one expert. The sparse version performed worse than its dense control and was slower in the short timing test.

We followed that with a more controlled comparison across three seeds. Stronger load balancing prevented the observed route collapse: all eight experts were used. Yet distributing work across experts was not the same as gaining useful capacity. A dense model with approximately the same number of total parameters achieved better held-out loss. The balanced sparse model was also slower than the same-active-width dense control in every seed; its median measured throughput was about 161 versus 245 generated bytes per second.

All expert weights remained in GPU memory. Selecting one expert per token reduced some computation, but it did not create a measured memory saving or a practical efficiency advantage in this setup. The generated samples were not useful long-form language.

What we take forward

These results do not rule out sparse models or conditional computation. They rule out treating their appealing properties as results before the measurements support them.

We can make a router use every expert. We have not shown that those experts provide a worthwhile quality benefit. We can omit or replace a piece of computation. We have not shown that the change reliably preserves behavior while making inference faster. And we can train an original model on consumer hardware. That does not mean it is ready to scale into a capable assistant.

CFI’s aim remains ambitious: more useful intelligence within hardware people can own and control. The path toward it now has sharper tests. A proposed improvement must earn its place against dense controls, held-out behavior, actual runtime, and memory use—not just parameter counts or promising intermediate scores.

We will keep publishing the experiments that work and the ones that do not. The negative results are part of how we find out what is worth building next. For the broader motivation, read why we’re building for the edge.

Related research: Consumer Frontier Intelligence