Inkling-Small’s Real Benchmark Is the Price Chart, Not the Leaderboard

CryptoTiger Guide
Mira Murati’s new lab just fired a shot that most AI commentary will misread as a model release. It is not a model release. It is a pricing attack on the entire reasoning-token value chain. The proof is not the benchmark score. The proof is $1.20 per million output tokens on a 276B-parameter Mixture-of-Experts architecture that still manages to beat its bigger sibling on code and hard reasoning tasks. In a market where cost-per-intelligence is becoming the only religion, that number matters more than the next Intelligence Index tick. Context: Thinking Machines Lab, the company founded by former OpenAI CTO Mira Murati, published Inkling-Small under Apache 2.0. Alongside it sits Inkling, a much larger model: 975B total parameters, 41B active. Inkling-Small is 276B total, 12B active. That ratio matters. 276 divided by 975 is roughly 0.28. 12 divided by 41 is roughly 0.29. Same MoE blueprint, smaller body. Yet Inkling-Small scores 40 on Artificial Analysis’s Intelligence Index while Inkling scores 41. It flips the larger model on SWE-bench Verified and HLE. It accepts text, image, and audio input, and it does not generate images or audio. Quantized weights weigh 171GB. The parameter math is the first tell. A 12B-active-parameter model is not a “small model” in the sense of something a laptop can run. It is a deployment-tier choice. At 12B active parameters, the per-token compute cost lands in the mid-size range, while the total parameter count carries a larger share of knowledge capacity. That split is the physical foundation of the $1.20 price. A dense 276B model could not sell tokens at that price without bleeding. MoE can, because only a slice of the network wakes up for each token. The second tell is the benchmark inversion. Distillation shrinks a model and usually gives back a little knowledge loss across the board. Inkling-Small is not simply a distilled Inkling. It is better than Inkling on SWE-bench Verified and HLE. You do not get that by shrinking. You get it by reweighting the training diet. The most reasonable read: Inkling-Small was trained or post-trained with a higher share of code, math, and formal-reasoning data, or with a curriculum that pushes those domains forward. That is a deliberate specialization, not an accident. This also creates a weakness. The same release notes say Inkling is better on knowledge coverage and factual accuracy. Translate that honestly: Inkling-Small is a specialist. It is built for toolchains, code agents, and high-stakes reasoning pipelines. It is not a general-purpose assistant that will answer trivia better than its bigger sibling. The multimodal input side reinforces the narrowness: text, image, and audio in, no generation out. That keeps the ceiling low and the inference cost lower. It also means Thinking Machines is not trying to out-Gemini Google. It is trying to own the slice of the market where “produce the correct output” matters more than “produce a beautiful image.” Now do the cost arithmetic. I spent weeks reverse-engineering Compound’s cToken contracts before deploying a dollar, and that habit stuck: distrust any chart that cannot be traced back to a mechanism. The mechanism here is brutal. Training a 975B-parameter model, even with MoE, requires at least the order of 10^25 FLOPs. On an H100 cluster with realistic utilization, that is thousands of GPUs for months. No lab burns that compute and then sets API prices to maximize near-term profit. The $1.20 price is penetration pricing. It is designed to pull developer volume, collect usage data, and establish the model as the default reasoning engine inside code workflows before raising prices or selling higher-margin services like fine-tuning, managed hosting, and enterprise SLAs. “Cheaper than the flagship” is a marketing anchor. It is also a migration subsidy. The investment logic follows the same pattern. Mira Murati’s name is an asset that discount rates respect. Founder premium is real. But the burn rate is also real. A lab that trains a 975B-parameter model and then prices its small model below commodity rates is not trying to earn profitability in year one. It is trying to own a position in the developer stack before the next funding round. Apache 2.0 accelerates that play by removing procurement friction. It also makes the company an attractive acquisition target, because a future acquirer can integrate the weights without license chains. Numbers do not lie, but they do hide. The published benchmark table hides the balance sheet. The contrarian read is not that Inkling-Small is overhyped. It is that the open-source framing has been oversold. Apache 2.0 gives an enterprise the right to own, modify, and deploy the weights. It does not give a solo developer the hardware to use them. 171GB of quantized weights is not a consumer artifact. It is a B2B procurement line item. This is open-source as an enterprise trust trick, not as grassroots democratization. It removes legal friction while preserving infrastructure gravity. The teams that can actually run Inkling-Small are exactly the teams that would otherwise be paying a closed API a per-token tax. Apache 2.0 is a hedge, not a charity. The second contrarian read is the benchmark band. “Only one point behind Inkling” sounds like a near miss. It is not. Artificial Analysis’s Intelligence Index is an aggregate, and global frontier models have historically sat far above the 40-level range. If the SOTA band is 50-plus, then Inkling-Small and Inkling are both below the front line. The “one point” narrative is an internal comparison inside a low band. The release conveniently avoids direct benchmarks against GPT, Claude, Gemini, DeepSeek, or Llama. Omission is a message. If the numbers were flattering, the team would publish them. This is exactly how order books work: the chart shows fear; the order book shows intent. The published chart is favorable. The missing order book is more informative. Then there is the security side. Apache 2.0 full weights on a multimodal reasoning model create an uncomfortable problem. Once weights are released, the lab’s control over model behavior drops to zero. Anyone can fine-tune, remove alignment, or embed malicious instructions. Audio input adds another attack surface. It is possible to imagine injected audio instructions or deepfake-analysis pipelines being gamed. The release does not disclose a detailed red-team analysis, model card, or alignment recipe. For a lab run by a former OpenAI CTO, the silence is loud. In crypto, I say security is a feature, not a marketing slide. The same applies here. If safety infrastructure were mature, the release package would say so. It does not. Survival precedes profit in the unregulated wild. The AI market currently rewards speed and distribution over caution, but the bill eventually arrives. A model with full weights, no public safety eval, and a very low price is a product designed for adoption before scrutiny. That is not necessarily malicious. It is simply the pattern of a startup that needs momentum. The danger is that momentum becomes the only risk model, and risk models that ignore tail events end up repriced in a single afternoon. Patience is a tactical advantage, not a virtue. The next six months will separate the real ecosystem from the narrative ecosystem. Watch three signals. First, integrations: does Inkling-Small show up in mainstream IDEs, CI/CD tools, and agent frameworks? Second, third-party adapters: GGUF and AWQ quantizations, vLLM and SGLang support. Third, API volume and paid enterprise accounts. If those appear, this is the opening of a new liquidity pool in the AI stack. If they remain absent, the open weights are a tombstone for a well-funded thesis. Code does not negotiate. It executes or it fails. So will this release.

Inkling-Small’s Real Benchmark Is the Price Chart, Not the Leaderboard

Inkling-Small’s Real Benchmark Is the Price Chart, Not the Leaderboard

Inkling-Small’s Real Benchmark Is the Price Chart, Not the Leaderboard

Market Prices

BTC Bitcoin
$63,182.1 +0.13%
ETH Ethereum
$1,858.94 -0.46%
SOL Solana
$73.13 +0.26%
BNB BNB Chain
$582.1 +0.47%
XRP XRP Ledger
$1.08 +1.41%
DOGE Dogecoin
$0.0700 +0.34%
ADA Cardano
$0.1887 +8.95%
AVAX Avalanche
$6.58 +3.48%
DOT Polkadot
$0.7950 +3.37%
LINK Chainlink
$8.3 +2.37%

Fear & Greed

27

Fear

Market Sentiment

Event Calendar

{{年份}}
30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

12
05
halving BCH Halving

Block reward halving event

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

18
03
unlock Sui Token Unlock

Team and early investor shares released

28
03
unlock Arbitrum Token Unlock

92 million ARB released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

Market Cap

All →
1
Bitcoin
BTC
$63,182.1
1
Ethereum
ETH
$1,858.94
1
Solana
SOL
$73.13
1
BNB Chain
BNB
$582.1
1
XRP Ledger
XRP
$1.08
1
Dogecoin
DOGE
$0.0700
1
Cardano
ADA
$0.1887
1
Avalanche
AVAX
$6.58
1
Polkadot
DOT
$0.7950
1
Chainlink
LINK
$8.3

Tools

All →

Altseason Index

44

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

🐋 Whale Tracker

🟢
0xdbad...dd3f
1d ago
In
7,256,797 DOGE
🔴
0x5e49...ab31
30m ago
Out
12,020 BNB
🔵
0xfc03...8392
5m ago
Stake
4,767.27 BTC

💡 Smart Money

0xb020...a031
Early Investor
+$1.4M
92%
0x9233...b1c8
Institutional Custody
-$0.9M
74%
0x8e79...4ed6
Market Maker
+$0.7M
74%