Mostik Links Frontier and Small AI Models Through Latent-Space Reasoning

| September 24 | The Ascendants
Mostik, Mostik AI, Sasha Malysheva, Stanislav Smirnov, General Catalyst, AI Models, Latent Space, Frontier Models, Open Weight Models, Artificial Intelligence, ARC-AGI, GLM-5.2, Qwen-3.5, AI Reasoning, AI Infrastructure, Model Interoperability, AI Startups

Share

For years, one of the biggest conversations in artificial intelligence has revolved around a familiar question: can smaller, open models eventually catch up with the enormous frontier systems built at vast cost? Mostik is approaching the problem from a very different direction.

The startup is not trying to make a 4-billion-parameter model become a 753-billion-parameter model. Instead, it is asking whether the smaller model needs to reproduce all of that intelligence in the first place.

What if the larger model could do the hard thinking, then pass what it has understood directly to the smaller model without translating that reasoning into words?

That question sits at the centre of Mostik’s work.

The company says its protocol allows AI models to communicate in latent space, passing hidden internal representations from one model to another without using text as the intermediary. Crucially, Mostik says the models involved do not need to be fine-tuned and can come from different model families.

In simple terms, the frontier model can reason about the problem, while a much smaller model running closer to the user can produce the answer.

It is an unusual attempt to separate where the reasoning happens from where the response is generated.

A bridge between two very different models

Mostik’s name itself means “bridge” in Russian, an unusually literal fit for what the company is trying to build.

Rather than having one model generate text and forcing another model to read and interpret it, Mostik creates a mathematical connection between their internal representations. One description of the system says a lightweight bridge matrix projects the high-dimensional hidden state of the larger model into the latent input space of the smaller one, without altering the original models’ weights.

That distinction matters.

AI systems working together commonly communicate through generated tokens. One model produces an answer, explanation or intermediate output, and another model consumes it. Mostik’s approach attempts to remove that textual handoff altogether.

Its bet is that models do not necessarily have to “speak” to one another in human-readable language to exchange useful information.

That could make collaboration between models less dependent on repeatedly generating and processing long streams of tokens.

A 753B model thinks. A 4B model answers.

The company’s most striking demonstration brings the idea into focus.

Mostik connected a 753-billion-parameter GLM-5.2 model with a 4-billion-parameter Qwen-3.5 model, with the smaller model positioned as an edge-class system capable of operating much closer to the end user.

The startup says the larger model can read and reason about the problem before its hidden state is transferred to the smaller model, which then produces the answer.

Mostik says the resulting setup achieved roughly 80% of the accuracy of the larger reference model. The company’s own account describes the system as delivering that level of performance while running about 20 times faster. Separately, published descriptions of the experiment report inference cost at roughly one-twentieth of running the full 753B model.

The important part is not simply that a small model produced a useful answer.

It is that Mostik says it did so after receiving internal reasoning information from a vastly larger model, rather than receiving the larger model’s finished text.

For companies trying to run AI on their own infrastructure, devices or specialised environments, that is the larger proposition behind the technology.

The frontier model does not necessarily have to generate every final token itself.

Built by mathematicians, not around a bigger-model race

Mostik says the work came together remarkably quickly.

The team describes itself as 15 people, including 12 PhDs and a Fields Medalist, working on the problem over four months.

At the centre of the company is CEO and co-founder Sasha Malysheva, who has been identified as the main developer of the approach. Mostik’s chief scientist is Stanislav Smirnov, a University of Geneva professor and winner of the 2010 Fields Medal.

mostik team

That mathematical background helps explain why Mostik’s pitch feels different from the usual AI infrastructure story.

It is not primarily about training an even larger model, building a new chatbot interface or introducing another orchestration layer. The company’s focus is the mathematical relationship between the internal representations of existing models.

The underlying idea is that models trained separately may still contain compatible semantic structures that can be connected.

If that premise continues to hold across model families, Mostik could potentially make it easier to combine a highly capable general model with smaller systems chosen for deployment constraints or specialised tasks.

The ARC-AGI result the company is not fully discussing yet

Mostik is also pointing to an early competitive result.

The company says its technology reached first place on the ARC-AGI leaderboard, although it has deliberately kept details limited while the competition is still running.

Other coverage describes the system as having reached the top of ARC-AGI 3, a benchmark intended to test AI reasoning.

For now, the result is best treated for what it is: an early indicator attached to a technology whose broader capabilities are still being established.

The more consequential question is whether the same model-to-model communication can work reliably across many architectures, providers and real-world workloads.

That is where Mostik’s ambitions extend beyond a benchmark.

A challenge to frontier-model lock-in

Mostik is also framing the technology as an infrastructure play.

The startup says it wants to reduce dependence on any single frontier-model provider and is already working with inference providers to support greater adoption of open-weight models.

That vision is straightforward.

A company could potentially use a powerful frontier system for the part of a task where deep reasoning matters, while relying on a much smaller model on its own infrastructure for the final interaction or execution.

Instead of choosing between frontier intelligence and local deployment, Mostik wants to connect the two.

The company says neither model needs to be fine-tuned for the bridge to work, which, if the approach proves broadly transferable, could make model combinations easier to deploy without rebuilding the underlying systems.

Mostik is backed by General Catalyst, according to the company.

A different question for the AI industry

There is still much that will have to be demonstrated.

Connecting internal representations between independently developed models is technically difficult. Different architectures, training processes and internal structures can make their latent spaces fundamentally different, and outside observers have noted that the bigger test will be whether Mostik’s approach holds up across many unrelated model families and production environments.

But the company’s most interesting contribution may already be the question it has placed on the table.

AI development has largely treated the model generating the final answer as the model that must contain the intelligence required to produce it.

Mostik is challenging that assumption.

If reasoning can be transferred directly between models, the future may not be defined solely by which company builds the biggest single system.

It may also depend on how intelligently different models can work together.

And for a startup built around the idea of creating a bridge between them, that is exactly the point.

Also Read: Meta VR Glasses: All About Price, Weight, Features, Launch Date and Muse Charm

Leave the first comment