We Didn’t See It Coming: Murati’s Inkling-Small and the Silent Cost Curve
Culture
|
CryptoStack
|
We didn’t.
We didn’t need another open-weight model, not really. Every week another lab releases something that claims to crack the code, and every week the hugging face marketplace gets a little noisier. But when the release comes with the name of the former OpenAI CTO behind it, the noise changes pitch.
Mira Murati’s Thinking Machines Lab dropped Inkling-Small this week. It’s a Mixture-of-Experts multimodal reasoning model with 276B total parameters and 12B active. The larger sibling, Inkling, sits at 975B total, 41B active. On Artificial Analysis’s Intelligence Index, Inkling-Small scores 40, just one point behind Inkling’s 41. On SWE-bench Verified and HLE, the smaller model actually flips the script and beats the bigger one. The weights are Apache 2.0. The API price is $1.20 per million output tokens, about 70% cheaper than Inkling. The quantized weights are 171GB. The press release doesn't mention the training compute, the data mix, or the alignment methods.
That last sentence is where the story starts.
Let me be honest about my own bias. In 2018, I spent 40 hours reverse-engineering Raptor Protocol’s smart contracts, convinced I had found the yield strategy that would define a cycle. I published a 3,000-word thesis right before a reentrancy bug drained $2 million. The narrative was right until it wasn’t. I learned that the ledger’s silence matters more than the numbers it shouts. In this case, the silence is deafening.
MoE is not new. The architecture is a routing layer that sends each token to a small subset of expert networks, which gives you the capacity of a huge model at the cost of a small one. Inkling-Small is a textbook example. The ratio between the full model and the small model is nearly identical in total parameters (975B vs 276B, roughly 3.5x) and active parameters (41B vs 12B, roughly 3.4x). That’s the signature of the same breed sliced into two sizes, not a new species. The innovation is not architectural; it’s compositional. The real signal is that Inkling-Small outperforms Inkling on coding and exam-hard reasoning benchmarks. That doesn’t happen by accident. You don’t shrink a model and get better results by simple distillation. You get better results because you rebalanced the data, weighted the curriculum, and did targeted post-training on code and math. In other words, the small model isn’t a scaled-down Inkling; it’s a purpose-built specialist wearing the same brand.
This is where my crypto brain starts firing. The parameter-to-performance relationship here is the same as the settlement-to-execution split that every Layer 2 has been selling for years. A fat, expensive training run is the settlement layer; the 12B-active inference path is the execution shard. The catch is the same one I’ve been shouting about since 2021: the sequencer is still a single point of control. Here, the sequencer is the H100 cluster, the data pipeline, and the undisclosed compute supplier. “Training heavy, inference light” is a beautiful structural line, but it means the real trust lives in the part of the pipeline nobody wrote a blog post about.
Every bull run is a myth waiting to be debunked, and the AI bull run is no exception. Debunk the most seductive part of the Inkling-Small myth: that open weights mean democratization. Apache 2.0 is the most permissive license in the game. But 171GB of quantized weights is not a gift to a solo developer with a laptop. That’s a threshold that assumes someone has a server rack and an ops team. This is B2B open-source, or what I like to call open-washing: the legal friction disappears while the infrastructure friction stays exactly where it is. The model is open enough to let enterprises say they own their AI stack, but closed enough that the average user still has to rent it. At $1.20 per million output tokens, with 12B active parameters, the gross margin per token is far healthier than any comparably priced competitor. That is not a discount; it is a land grab. Yield is the bait, liquidity is the trap: the developer mindshare, the data flywheel, the routing telemetry.
But let’s push the contrarian knife in deeper. The “one point behind Inkling” narrative is designed to feel close. In absolute terms, a 40 on the Intelligence Index is not a top-of-the-food-chain score. The state of the art is in the 50s. So what we have here is a comparison between two models in the same family, on a modest step of the staircase, and the marketing team is asking us to measure the centimeter between the fourth and fifth steps while pretending the staircase doesn’t keep going up. The fact that Thinking Machines Lab chose not to release comparisons against GPT-5, Claude Opus, or Gemini Ultra is itself a data point. When you have a native advantage, you put it on the front page. When you have a relative advantage, you put it in a footnote.
There is also the missing section on safety. Murati comes from OpenAI, where safety culture was the public religion. Yet the Inkling-Small release omits the red-team results, the model card transparency, and the alignment methodology. Apache 2.0 weights with multimodal audio input means the downstream control is zero once the weights are on the mirror. Any third party can fine-tune, de-align, or embed hidden instructions in audio spectrograms. In the ledger’s silence, the true story whispers: either the safety work isn’t done, or it isn’t finished, or it doesn’t make the model look as good.
Code is law, but humans write the bugs. The launch may be technically impressive, but the governance of open reasoning models is still the same bug — the one where you let the code run and hope no one audits the backdoor. If Inkling-Small becomes the backbone of a developer toolchain, the exploit surface doesn’t disappear; it moves into the supply chain. A poisoned fine-tune, a malicious LoRA adapter, a hidden audio trigger — the surface area of the open-weights world is vastly larger than any closed API.
Let’s zoom out. The real industrial impact of Inkling-Small is not the benchmark points. It’s the price anchor. By pricing the API at $1.20 per million output tokens, Thinking Machines Lab is telling the market that near-frontier coding intelligence can be had cheaply. That puts downward pressure on every closed API vendor in the category. And because the weights are open, any enterprise that wants to avoid sending code to a third party can now self-host. The “don’t exfiltrate our source code” objection melts away. That is a wedge into the code-agent market, which is exactly where the next wave of AI revenue will live.
The deeper irony is that this release is not a crypto story, yet it runs on crypto’s oldest logic. Open source as a trust anchor. Cost-per-smartness as a unit of account. A heavy settlement layer paying for a lean execution path. The same modularity debates that dominate Layer 2 discourse are playing out in model design. Same blind spots: the assumption that transparency of the ledger equals transparency of intent. Strip away the model weights and what do you see? A classic DeFi launch: clever tokenomics, a trust flaw hidden in the infrastructure layer, and a narrative built on a relative metric.
Based on my audit experience, I can tell you what I would ask before deploying Inkling-Small in production. Where is the data bill of materials? What is the routing loss under adversarial load? Which quantization method was used for the 171GB release, and what is the degradation curve? What is the actual context window? Where is the independent third-party evaluation? If the answers don’t exist yet, then this is not a deployment, it’s an experiment with someone else’s capital.
Sentiment is a shifting tide, not a solid ground. Right now the tide says “Mira Murati, OpenAI pedigree, Apache 2.0, cheap tokens, good code.” The tide will shift when someone measures the model in a realistic workflow and finds the gap. Not if. When.
What comes next? The battle won’t be decided by weights. It will be decided by the infrastructure that routes, pays for, and governs these models. Agents need to move value. Models need to pay for compute, attest to outputs, and prove data provenance. That is a ledger problem. And if the open-weights crowd can’t solve it, the closed-API crowd will solve it for them — and charge a toll on every thought the machines have.
We didn’t need another open-weight model. But we might need a ledger for the ones we built. And no one has written that contract yet.