SOLTEN & CO. RESEARCH | FLAGSHIP REPORT | 31 AUGUST 2026 | VERSION 1.0
AI Models / Infrastructure / Cloud Economics / Capital Allocation
| CENTRAL THESIS: Chinese laboratories are not giving intelligence away. They are reducing the price of model access to accelerate adoption, shape technical standards, and redirect demand toward paid inference, cloud infrastructure, enterprise products, and applications. The strategy is already economically visible at Alibaba, but it is not proven across the private labs—and it can be disrupted by serving cost, regulation, security concerns, or government restrictions. |
Research snapshot
| Field | Detail |
|---|---|
| Publication date | 31 August 2026 |
| Evidence cut-off | 21 August 2026 |
| Research type | Flagship thematic and market-structure report |
| Primary audience | Family offices, venture/growth investors, lean investment teams and strategy leaders |
| Estimated reading time | 30–35 minutes |
| Venture Deal Protocol | Not applicable: this report does not underwrite a financing or acquisition |
| Core evidence boundary | Public data measures releases, prices, downloads and one listed company’s segment economics; it does not reveal private-lab gross margins or production workload share |
Executive summary
China’s leading AI laboratories are releasing model weights at a scale that would once have been treated as proprietary crown jewels. In August, Alibaba released weights for the 2.4-trillion-parameter Qwen3.8 flagship; Moonshot’s 2.8-trillion-parameter Kimi K3, DeepSeek V4, MiniMax M3 and Z.ai’s GLM family had already established a pattern. The immediate interpretation—China is giving away its best models for free—is memorable and incomplete. What is free is usually a licence to download weights. Training data, full reproducibility, reliable serving, compute, security, support and enterprise integration remain scarce or paid. Some licences also impose commercial conditions.[1]–[13]
The strategic logic is a substitution: sacrifice model-layer scarcity to acquire distribution. Open weights lower switching and experimentation costs, invite community optimisation, create derivatives and make a model family available through many clouds, gateways and on-premise stacks. That reach can stimulate demand for hosted APIs, cloud compute, accelerators, inference optimisation, security, evaluation and model-enabled applications. It can also influence which architectures and tooling conventions become defaults. Hugging Face counted 151,448 Qwen derivatives and roughly 180–210 new Qwen-based repositories per day during the first seven months of 2026. It simultaneously warned that Hub metrics are not commercial market share.[14]
Alibaba provides the clearest public evidence that the model can work. For the June 2026 quarter, its AI Cloud and Compute Services revenue rose 45% year over year to RMB48.4 billion and adjusted EBITA rose 133% to RMB5.6 billion. Alibaba also reported more than three billion Qwen downloads and more than 300,000 derivatives, while saying AI-related product revenue had delivered a twelfth consecutive quarter of triple-digit growth. Capital expenditure rose 75% to RMB67.7 billion. This is consistent with open models feeding a full-stack cloud business; it does not isolate Qwen’s causal contribution, and it shows that the strategy is capital intensive.[17], [18]
The strategy is not uniform. Qwen releases models across a broad size range, inviting developers to standardise on one family from local deployment to frontier workloads. Moonshot, Z.ai and MiniMax lead with very large models that few developers can serve directly; open weights create attention and distribution, but their paid APIs and subscriptions remain the practical product for many users. DeepSeek combines permissive MIT releases with official API prices far below U.S. frontier services. These are different business models, not evidence of a single centrally directed commercial plan.[4]–[14]
For investors, the consequence is a value migration rather than value destruction. Standalone model API pricing is under pressure. Compute demand, inference throughput and optimisation can expand. Enterprise control points—security, evaluation, observability, orchestration, proprietary data and workflow integration—become more important as model choice commoditises. Application companies can gain from lower variable cost, but only if competition does not pass the entire saving to customers. The most exposed companies are model vendors whose differentiation rests on generic capability and whose monetisation does not extend into cloud, distribution, proprietary data or workflow ownership.
The counter-thesis is substantial. The largest open models are expensive to serve: raw BF16 weight memory is approximately 4.8 terabytes for Qwen3.8’s 2.4 trillion parameters and 5.6 terabytes for Kimi K3’s 2.8 trillion, before runtime overhead. Download counts can reflect curiosity, mirrors or automated pipelines rather than durable production. Security review, censorship behaviour, licensing ambiguity and geopolitical restrictions can block enterprise adoption. Reuters reported that Chinese authorities had discussed possible limits on overseas access to advanced models; no policy decision had been confirmed at the evidence cut-off.[2], [4], [14], [26]
Our conclusion is conditional but decisive: the topic merits fundamental coverage because open-weight competition changes the price of intelligence, the geography of standards and the profit pools around AI. The investable signal is not another benchmark victory. It is whether open-model distribution converts into paid inference, cloud utilisation, enterprise deployments and application gross profit without destroying the margins needed to fund the next generation.
Key findings
- ‘Free’ is a distribution choice, not a cost structure. Most releases provide weights; they do not provide training data, full reproducibility, serving or enterprise assurance.
- China is not a single actor. Qwen’s full-spectrum strategy aims at developer standardisation; frontier-first labs use open releases as demand generation, credibility and ecosystem leverage.
- The strongest monetisation evidence is Alibaba’s cloud segment. It supports the mechanism but does not prove that every private lab can reproduce it.
- Open-model adoption is real but poorly measured. Derivatives and downloads indicate ecosystem activity; they do not reveal production workloads, routed dollars, retention or gross margin.
- Licences are becoming a strategic variable. DeepSeek and GLM use MIT for major releases, while Qwen3.8 and MiniMax M3 use custom terms that can limit or condition commercial use.
- Frontier API prices span orders of magnitude. DeepSeek V4 Pro lists $0.435 per million uncached input tokens and $0.87 output; OpenAI GPT-5.6 Sol lists $5 and $30; Anthropic Claude Fable 5 lists $10 and $50. Capability, latency, support and safety are not normalised.
- Open weights benefit cloud and hardware only if lower prices expand workload volume faster than unit margins compress. Capex and power remain unavoidable.
- The model layer can retain strategic value even if direct licensing value falls: model defaults influence toolchains, hardware optimisation, application design and standards.
- The largest threats are regulatory bifurcation, security concerns, serving economics, custom licence friction and a renewed capability gap in favour of closed models.
- The decisive 12–24 month indicators are production spend, enterprise deployments, paid API mix, cloud segment economics, independent capability-per-dollar, and any export or access restrictions.
Contents
| Sections 1–8 | Sections 9–15 |
|---|---|
| 1. What is actually being given away | 9. Value migration across the AI stack |
| 2. The August 2026 release wave | 10. Industrial-policy and geopolitical logic |
| 3. Two open-model strategies | 11. Counter-thesis and falsification |
| 4. The economics of distribution | 12. Scenarios for 2027–2029 |
| 5. Evidence that monetisation is occurring | 13. Investor implications |
| 6. Pricing pressure and deployment reality | 14. Open questions and monitoring triggers |
| 7. Adoption: strong signal, weak instrument | 15. Conclusion |
| 8. Licences, openness and control | Methodology, limitations and sources |
1. What is actually being given away
The phrase ‘open source’ conceals several different bundles. A fully reproducible system would disclose weights, architecture, inference code, training code, data provenance and a licence permitting modification and redistribution. Most frontier releases provide weights, architecture descriptions and enough code to run inference. They do not disclose the complete training corpus or a recipe that lets a third party recreate the model. The more precise term is open weight.[24], [25]
That distinction is economically important. A developer may avoid a per-token licence toll by downloading the weights, but still needs storage, accelerators, high-bandwidth memory, networking, serving software, optimisation, monitoring, security and people. For very large mixture-of-experts models, only part of the parameter set is active for each token, reducing compute relative to a dense model. The full weight set must nevertheless be stored and made available to the serving system. Quantisation can reduce memory; runtime state, redundancy and the key-value cache add it back.
Exhibit 1. The cost stack behind a free model

Source: Solten & Co. framework. A zero licence price does not imply zero total cost of ownership.
| Model | Total / active parameters | Approx. raw BF16 weight memory | Release / licence boundary |
|---|---|---|---|
| Qwen3.8-2.4T-A95B | 2.4T / 95B | 4.8 TB | Weights released; custom Qwen3.8-Max licence |
| Kimi K3 | 2.8T / vendor-disclosed MoE | 5.6 TB | Weights released; custom terms |
| DeepSeek V4 Pro | 1.6T / 49B | 3.2 TB | Official open weights; MIT |
| GLM-5.2 | 744B MoE | 1.49 TB | Weights released; MIT |
| MiniMax M3 | ~428B / ~23B | 0.86 TB | Weights released; custom community licence |
Memory figures are arithmetic: parameter count multiplied by two bytes, expressed in decimal terabytes. They exclude quantisation, runtime overhead, cache, redundancy and multimodal components. They are deployment-scale illustrations, not hardware bills.[2], [4], [6], [9], [12]
2. The August 2026 release wave
Alibaba’s August release brought the argument into focus. Qwen3.8-Max launched first as a managed multimodal model. The company then released a 2.4-trillion-parameter, 95-billion-active text checkpoint on 12 August and a 27-billion-parameter model on 14 August. It was the first open release at Alibaba’s Qwen-Max scale. The open checkpoint and managed service are related but not identical: the hosted product adds vision, longer default context and integrated tools.[1]–[3]
The surrounding market makes the release structural rather than episodic. Moonshot introduced Kimi K3 in July with 2.8 trillion parameters and a one-million-token context window, distributed through weights, a consumer product, Kimi Code, Kimi Work, enterprise subscriptions and a paid API. Z.ai released GLM-5.2 under MIT in June and announced GLM-5.3 in August, with weights scheduled after additional safety work. DeepSeek’s V4 family combines a 284-billion-class Flash model and a 1.6-trillion-parameter Pro model under MIT. MiniMax M3 provides a 428-billion-parameter multimodal model under a custom community licence.[4]–[13]
| Family | Latest relevant release | Open-weight status | Paid surface | Investor-relevant caveat |
|---|---|---|---|---|
| Qwen | Qwen3.8 Max / 2.4T-A95B | Released; custom licence | Alibaba Cloud, QwenWork, apps | Strongest public cloud monetisation anchor |
| Kimi | K3, 2.8T | Released; custom terms | API, subscriptions, enterprise | Serving scale makes hosted access practical |
| DeepSeek | V4 Flash / Pro | Released; MIT | Official API | Extremely low list price; private economics undisclosed |
| Z.ai | GLM-5.2 / 5.3 | 5.2 MIT; 5.3 promised | API, coding plans | Release timing and product version can diverge |
| MiniMax | M3, ~428B | Released; custom licence | API, agent and subscriptions | Commercial authorisation threshold in licence |
Vendor evaluations show these models close to the frontier on selected coding, agentic and long-context tasks, but no model is uniformly best. Harnesses, reasoning budgets, hardware, prompts and task mix differ. We use benchmarks to establish that the releases are strategically credible, not to crown a global winner.[2]–[7], [30]
3. Two open-model strategies
The release portfolios reveal two strategies. Qwen and parts of Tencent cover the full range from small local models to frontier systems. That lets a developer learn one family, fine-tune it, deploy a small version on-device and move to hosted frontier inference without changing the conceptual stack. The commercial prize is standardisation. Hugging Face reports that Qwen’s 2026 portfolio generated roughly 2.05 billion downloads across repositories with declared parameter counts—about 55 times Moonshot’s frontier-only portfolio—and 151,448 derivatives.[14]
Moonshot, Z.ai and MiniMax publish far less below 70 billion parameters. Their largest models are too expensive for most developers to serve directly, so the open release works as proof, distribution and a community-optimisation seed. Quantised variants appear quickly; third-party inference providers add access; the laboratory sells an official API, coding plan, consumer subscription or enterprise service. DeepSeek sits between the categories: its permissive releases achieve enormous distribution, while its official API is priced aggressively enough to compete with self-hosting for many workloads.[4]–[14]
Exhibit 2. Two different routes from openness to economic value

Source: Hugging Face release-portfolio analysis; company product disclosures; Solten & Co. classification.
This diversity matters for underwriting. Alibaba can subsidise model R&D with e-commerce cash flow and monetise through a large cloud. A venture-backed laboratory must turn attention into paid inference, subscriptions, enterprise contracts or strategic financing before cash runs out. The same open-weight tactic therefore has different runway, pricing and bargaining implications across firms.
4. The economics of distribution
Opening weights changes the customer-acquisition equation. A closed API asks developers to trust one vendor’s economics, availability and roadmap. An open model can be downloaded, inspected, fine-tuned, mirrored, quantised and offered by competing hosts. Each new deployment becomes a distribution endpoint the originating lab did not have to finance. Derivatives add languages, domains, formats and hardware targets. The community absorbs part of the optimisation expense.
The return can arrive through five channels. First, a laboratory can sell an official managed API to users who prefer reliability to self-hosting. Second, a cloud provider can monetise storage, accelerators, networking and inference. Third, the model can pull users into an application or subscription. Fourth, broad adoption can influence libraries, evaluation conventions and hardware optimisation. Fifth, ecosystem position can improve financing and strategic bargaining power even before operating profit appears.[14], [19], [20]
The strategy resembles loss-leader economics but with an important difference: model weights are non-rival digital assets. Once released, the marginal distribution cost is low, but the strategic concession is irreversible. Competitors can study, modify and serve the model; the originator cannot later restore exclusivity. The bet is that adoption, iteration and complementary revenue exceed the lost option value of keeping the model closed.
| Concession | Expected return | Evidence available | Evidence still missing |
|---|---|---|---|
| Zero / low licence price | Faster adoption and derivatives | Downloads, repositories, provider listings | Production share and customer retention |
| Replicable serving | Broader distribution | Third-party hosts and quantisations | Originator’s paid share of demand |
| Community modification | Faster optimisation and localisation | Derivative models and ports | Value captured by the original lab |
| Benchmark visibility | API and subscription demand | Launch traffic and price pages | Cohort conversion and gross margin |
| Hardware compatibility | Cloud / chip utilisation | Vendor integrations | Incremental revenue attributable to model |
5. Evidence that monetisation is occurring
Alibaba is the observable case. The company’s June-quarter release reports RMB48.437 billion of AI Cloud and Compute Services revenue, up 45% year over year, and RMB5.628 billion of adjusted EBITA, up 133%. The implied adjusted EBITA margin was approximately 11.6%. AI-related product revenue reached RMB12.38 billion and continued triple-digit growth for a twelfth consecutive quarter. The company reported RMB67.7 billion of quarterly capital expenditure, up 75%.[17], [18]
Management explicitly describes the full-stack logic: open Qwen models create reach; training, development and inference run on Alibaba Cloud; the company also supplies chips, servers, storage, networking and applications. Qwen had more than three billion company-stated downloads and more than 300,000 derivative models by the earnings date. That is consistent with a distribution flywheel.[17], [19], [20]
| WHAT THE ALIBABA EVIDENCE PROVES—AND WHAT IT DOES NOT It proves that a company pursuing open-model distribution can simultaneously grow a large, increasingly profitable cloud segment. It does not prove that Qwen caused the growth, that every Qwen deployment runs on Alibaba Cloud, or that a standalone model laboratory can fund the same strategy. |
The private-lab evidence is thinner. Kimi sells token-based API access and tiered consumer subscriptions; Z.ai sells API access and coding plans; DeepSeek operates an aggressively priced API; MiniMax sells APIs and applications. These paid surfaces refute the claim that the companies reject monetisation. Public sources do not establish revenue mix, inference gross margin, retention, subsidy levels or the share of open-weight users that convert to paid services.[4], [5], [8], [11]–[13]
Capital formation is a second, weaker form of evidence. Ecosystem traction supports valuations and access to strategic partners. It is not evidence of durable unit economics. For investors, the distinction between financing validation and customer validation is essential.
6. Pricing pressure and deployment reality
Official list prices show how aggressively Chinese providers can attack the hosted layer. DeepSeek V4 Pro lists $0.435 per million uncached input tokens and $0.87 per million output tokens. Z.ai lists GLM-5 at $1 and $3.20. Kimi K3 lists $3 and $15. OpenAI lists GPT-5.6 Sol at $5 and $30, while Anthropic lists Claude Fable 5 at $10 and $50. Qwen3.8-Max’s international Model Studio price is CNY14.988 for input and CNY44.965 for output.[5], [8], [11], [21]–[23]
| Managed model | Uncached input / 1M | Output / 1M | Context | Boundary |
|---|---|---|---|---|
| DeepSeek V4 Pro | $0.435 | $0.87 | 1M | Official list; capability and service levels not normalised |
| Z.ai GLM-5 | $1.00 | $3.20 | 200K | Current developer price; GLM-5.3 price not yet listed |
| Kimi K3 | $3.00 | $15.00 | ~1M | Cache-hit input is $0.30 |
| OpenAI GPT-5.6 Sol | $5.00 | $30.00 | ~1.05M | Short-context standard list price |
| Anthropic Claude Fable 5 | $10.00 | $50.00 | Vendor service | Global API list price |
| Qwen3.8-Max intl. | CNY14.988 | CNY44.965 | 1M | Currency intentionally not converted |
On list price alone, DeepSeek V4 Pro is about 11.5 times cheaper than GPT-5.6 Sol on uncached input and 34.5 times cheaper on output. The comparison is illustrative, not a price-performance verdict. Models differ in quality, tokenisation, reasoning-token consumption, latency, uptime, safety, data handling, regional availability and support. Enterprise contracts can also diverge materially from list price.[11], [22]
Self-hosting is not automatically cheaper. The economic choice depends on utilisation. A fully provisioned cluster can be attractive at sustained volume, strict data-sovereignty requirements or specialised optimisation. At low or bursty volume, managed inference converts fixed capacity into variable cost and absorbs reliability engineering. Open weights expand the option set; they do not make every customer a cloud operator.
7. Adoption: strong signal, weak instrument
The evidence for ecosystem reach is unusually strong. Hugging Face reports Chinese frontier models exceeding U.S. open releases in scale in almost every month of 2026. Qwen-based repositories reached 151,448 derivatives, and Qwen added roughly 180–210 derivatives per day in the first seven months. The ATOM Project reports Qwen rising from about 1% of new fine-tunes and adaptations in January 2024 to 69% in February 2026. AP reported that the five most-used models on OpenRouter over a recent month were Chinese and that Kimi consumer downloads accelerated after K3.[14]–[16]
The same evidence contains its own warning. Hugging Face found that 85.6% of model repositories had fewer than 200 lifetime downloads and that 1.5% of repositories accounted for 99.2% of downloads. Only one repository appeared in both the year’s top-25 lists by downloads and likes. Small models below one billion parameters accounted for 83% of all-time downloads among repositories declaring size; models above 100 billion accounted for 1%. Attention, adoption and production are different variables.[14]
Downloads can be triggered by mirrors, automated pipelines and repeated environments. API usage omits private deployments; private deployments omit hosted APIs. Derivatives count experimentation and ecosystem labour, but not users or revenue. OpenRouter token shares are provider-specific and can shift with price promotions. A credible market-share measure would combine routed spend, production tokens, active deployments, renewal and workload criticality. No public dataset does this comprehensively.
| Metric | Useful for | Unsafe inference |
|---|---|---|
| Downloads | Distribution and pipeline activity | Commercial share or unique users |
| Likes / launch traffic | Attention and developer interest | Durable adoption |
| Derivatives | Community investment and standardisation | Originator revenue |
| Router token share | Hosted workload momentum | All deployment channels |
| API list price | Competitive posture | Realised price or gross margin |
| Named enterprise deployments | Production credibility | Portfolio-wide penetration |
8. Licences, openness and control
Licensing is becoming a monetisation surface rather than a footnote. Hugging Face found that 59% of 178 Chinese releases above 20 billion parameters in 2026 used Apache 2.0 and 22% used MIT. DeepSeek and Z.ai released major frontier models under MIT. The latest largest releases are less uniform: Qwen3.8 uses a custom licence, and MiniMax M3 requires attribution for commercial use and prior written authorisation above a $20 million annual-revenue threshold.[2], [6], [9], [13], [14]
This makes ‘free’ conditional. A company may be able to test, modify and deploy without paying the lab, yet still face attribution, usage, revenue or geographic terms. Licences can change between generations. The practical switching cost is therefore not only technical compatibility; it includes legal review and the risk that a future flagship has different conditions.
Control also survives outside the licence. The laboratory chooses release timing, safety tuning, documentation, tokenizer and architecture. The hosted version can include tools, multimodality, longer context, faster inference and support not present in the checkpoint. A lab can open yesterday’s weights while monetising today’s integrated product.
9. Value migration across the AI stack
Exhibit 3. Where value can move as model access gets cheaper

Source: Solten & Co. value-chain framework.
Model commoditisation is not binary. Frontier capability can retain a premium while adequate capability becomes abundant. Most enterprise tasks have a threshold: once accuracy, reliability and latency are sufficient, cost, data control and integration dominate. Open models accelerate competition below the absolute frontier and make model substitution easier. This pressures generic API margins and increases the value of routing, evaluation and proprietary workflow context.
Compute and cloud can benefit through an elasticity effect. Lower prices stimulate experiments, longer contexts, more agentic steps and more users. If token demand rises faster than unit prices fall, total inference revenue grows. Hardware vendors also use open models to demonstrate and optimise their systems. Hugging Face notes that NVIDIA and AMD were the most prolific U.S. model publishers in 2026 and interprets the releases as a route to chip demand.[14]
Tools and applications face a more selective outcome. Gateways, observability, security, evaluation and optimisation become valuable because model choice multiplies operational complexity. Application companies benefit if lower inference cost improves contribution margin or enables new usage. They lose the advantage if every competitor accesses the same models and passes savings to customers. Distribution, proprietary data and workflow ownership—not model access alone—determine who keeps the surplus.
| Layer | Likely first-order effect | What creates durable value | Principal risk |
|---|---|---|---|
| Frontier model APIs | Price pressure below the top capability tier | Unique capability, trust, distribution | Adequate open substitutes |
| Cloud / inference | Higher workload volume | Utilisation, cost curve, capacity | Capex and price competition |
| Chips / systems | More serving demand and optimisation | Performance per watt and ecosystem | Domestic substitution / export controls |
| Deployment tooling | More complexity and choice | Workflow data, governance, switching | Cloud bundling |
| Applications | Lower variable cost | Distribution and proprietary context | Savings competed away |
10. Industrial-policy and geopolitical logic
Open models also serve an industrial strategy. The U.S.-China Economic and Security Review Commission argues that low-cost, open-weight deployment can create two feedback loops: a digital loop of adoption and iteration and a physical loop in manufacturing, robotics and research that generates specialised real-world data. The paper treats China’s compute constraints, state support and industrial base as reinforcing conditions. It is a policy analysis, not neutral proof of commercial causality.[24]
Stanford HAI’s review is useful because it resists a monolithic account. Chinese firms vary in licence, architecture, target users and relationship to the state. Export controls may have encouraged efficiency and openness, but academic work finding association between U.S. policy and developer engagement does not establish a clean counterfactual. The observed ecosystem is the product of company strategy, capital constraints, cloud economics, policy support and developer demand.[25], [28]
Standards influence may be the largest long-duration prize. A widely fine-tuned model shapes toolchains, file formats, inference kernels, evaluation practice and developer skills. Hardware vendors optimise around it; universities teach it; applications inherit its behaviours. Those complements can persist even when another model wins the next benchmark.
Geopolitics can reverse the distribution advantage. Reuters reported in July that Chinese authorities had discussed restricting overseas access to advanced models, including open and closed systems; the scope was undecided and ministries and companies did not confirm the discussions. U.S. policymakers have also debated restrictions on Chinese models. A bifurcated market would reduce global scale but increase demand for sovereign hosting, regional clouds and compliance tooling.[26], [27]
11. Counter-thesis and falsification
A strong thesis must specify what would make it wrong. The first possibility is that open-model activity is mostly attention. Downloads and derivatives may fail to convert into production workloads, paid APIs or enterprise renewal. The second is that serving cost overwhelms the distribution benefit: very large checkpoints remain impractical outside specialist hosts, and official APIs price below sustainable economics. The third is that closed labs preserve a large capability or reliability gap and can continue charging a premium.
The fourth failure mode is institutional. Security teams may reject Chinese models because of data governance, supply-chain review, content behaviour or policy uncertainty. Custom licences may deter commercial use. China may limit overseas access, or the United States and allies may restrict procurement and distribution. The fifth is competitive: clouds and hardware vendors can adopt open models while capturing most of the economics, leaving the originating lab with little more than brand awareness.
| Thesis claim | Falsification test | Observable evidence |
|---|---|---|
| Open releases drive economic demand | Production spend and renewal do not follow ecosystem activity | API revenue, routed spend, named renewals |
| Value migrates to cloud / inference | Cloud AI growth or margins fail despite model adoption | Segment revenue, EBITA, utilisation, capex |
| Open models compress premium pricing | Closed APIs retain share and pricing at adequate workloads | Price changes, router mix, enterprise contracts |
| Standards create durable influence | Derivative and toolchain growth shifts away from Qwen | Repository creation, framework defaults, integrations |
| Global distribution compounds | Regulation or security blocks cross-border production use | Procurement bans, export limits, provider delistings |
| Self-hosting is a credible option | Total cost remains consistently above managed APIs | Independent TCO studies and utilisation data |
12. Scenarios for 2027–2029
Exhibit 4. Scenario map

Source: Solten & Co. scenario framework. Not a forecast.
In the open-substrate scenario, enterprises treat models as interchangeable components. Open families become default starting points; premium closed models are invoked only when incremental capability justifies cost. Model API prices compress, but inference volumes expand. Full-stack clouds, efficient hardware, deployment tooling and well-distributed applications capture the value.
In the bifurcated-stacks scenario, trust, regulation and supply chains separate markets. Chinese open models dominate parts of Asia, the Global South and private deployment; Western closed models retain regulated and high-trust workloads in the United States and allied markets. Sovereign hosting, localisation and compliance become larger profit pools. Standards split rather than converge.
In the constrained-diffusion scenario, large open models remain influential research and specialist assets but do not become the default production substrate. Serving cost, security reviews and access controls slow adoption. Managed inference reconcentrates demand, and closed frontier vendors preserve pricing power. The scenario is more likely if independent benchmarks show a widening capability gap or if official restrictions disrupt model availability.
13. Investor implications
For public-market investors, Alibaba is the cleanest live experiment. The relevant question is not whether Qwen wins benchmarks but whether AI-related cloud revenue, external customer growth and adjusted EBITA outpace capex and depreciation over a full cycle. The June quarter is encouraging: growth accelerated and segment profit expanded despite investment. One quarter cannot establish returns on invested capital.[17], [18]
For venture and growth investors, a model laboratory without cloud ownership needs a credible conversion surface. API revenue, enterprise subscriptions, consumer products, proprietary data or strategic distribution must fund training and serving. Downloads are a lead indicator at best. Diligence should prioritise realised price, inference contribution margin, paid conversion, retention, workload criticality, compute commitments and licence obligations.
For infrastructure investments, demand elasticity is decisive. Lower model prices can expand token volume, context length and agentic workloads, benefiting compute, memory, networking, power and cooling. It can also intensify price competition and strand inefficient capacity. Underwriting should be based on contracted utilisation, performance per watt, customer concentration and the ability to support multiple model families.
For software investors, lower inference cost is not automatically a moat. It can improve gross margin temporarily, then be competed away. Durable beneficiaries combine AI with distribution, proprietary workflow data, high switching costs or a control point such as security, evaluation or orchestration. A product whose only advantage is access to a strong model is increasingly fragile.
Diligence scorecard
| Question | Strong evidence | Weak evidence |
|---|---|---|
| Is adoption commercial? | Paid production spend, renewal, critical workloads | Downloads, likes, launch traffic |
| Is pricing sustainable? | Contribution margin after serving and support | Low list price without cost disclosure |
| Is the ecosystem defensible? | Derivatives plus tools, hardware and enterprise integrations | A single benchmark lead |
| Can the company fund the cycle? | Cash runway, cloud subsidy or contracted capacity | Strategic valuation alone |
| Can customers deploy safely? | Audits, governance, licence clarity, regional hosting | Open weights alone |
| Who captures lower cost? | Retained application margin or higher usage | Gross-cost decline with no pricing power |
14. Open questions and monitoring triggers
Open questions
- What share of Qwen, Kimi, DeepSeek, GLM and MiniMax usage is paid production rather than evaluation or community activity?
- What are realised API prices, inference gross margins and compute subsidies by laboratory?
- How much of Alibaba Cloud’s AI growth is attributable to Qwen-led workloads rather than general compute demand?
- Which open-model derivatives are used in regulated enterprise production outside China?
- How will custom licences evolve as the largest models become more expensive to train?
- Do Chinese open models retain capability-per-dollar leadership under independent, task-specific evaluation?
- What security, censorship and data-governance behaviours emerge across self-hosted and managed routes?
- Will Chinese or Western governments restrict model-weight distribution, procurement or cloud access?
- Does community optimisation create value for the originating lab or primarily for third-party clouds and hosts?
- At what utilisation and configuration does self-hosting beat managed inference on total cost and reliability?
Monitoring triggers
| Signal | Frequency | Thesis strengthens if | Thesis weakens if |
|---|---|---|---|
| Qwen derivatives and established downloads | Monthly | Growth persists beyond launch cohorts | Activity decays or shifts to another family |
| Router token and dollar share | Monthly | Chinese models gain paid production spend | Share is promotion-driven or transient |
| Alibaba AI Cloud revenue / EBITA / capex | Quarterly | Profit grows with utilisation | Capex rises without durable margin |
| Official API prices | Monthly | Volume expands without destructive repricing | Repeated cuts imply unsustainable competition |
| Independent capability-per-dollar | Per release | Open models remain adequate for enterprise tasks | Closed frontier gap widens materially |
| Licences and weight availability | Per release | Commercial terms remain usable | Restrictions and revenue conditions broaden |
| Named enterprise deployments | Quarterly | Renewed, regulated workloads appear | Pilots fail to reach production |
| China / U.S. policy | Event-driven | Cross-border access remains open | Export, procurement or hosting limits expand |
| Hardware and cloud integrations | Monthly | More optimised, multi-provider serving | Availability narrows to captive stacks |
15. Conclusion
China’s strongest open-weight releases are not acts of commercial surrender. They are bids to make model intelligence abundant enough that distribution, infrastructure and ecosystem position become the scarce assets. The model works most clearly for a full-stack company such as Alibaba, which can monetise compute, cloud services, chips and applications around Qwen. It is more speculative for independent laboratories that must finance frontier training while competing on API price.
The strategy matters beyond the laboratories themselves. It lowers the reservation price for adequate intelligence, weakens model exclusivity, accelerates multi-model deployment and raises the value of compute efficiency, governance, proprietary data and application distribution. It also exports architectural choices and developer habits. Those structural effects can persist even if a Chinese model never holds the absolute benchmark lead.
The correct investment stance is neither ‘open source wins’ nor ‘closed models win.’ It is to follow the conversion from reach to economics. The winning companies will be those that turn cheaper intelligence into paid throughput, durable workflows and attractive returns on capital. Downloads are evidence of movement. Revenue, retention, margin and standards are evidence of power.
Methodology and limitations
This report triangulates official model cards, licences, API price pages, listed-company disclosures, platform adoption data, policy research, academic work and high-quality reporting. Primary sources establish release facts, stated prices and company-reported metrics. Vendor benchmarks are not normalised and are not used as definitive rankings. Computed figures show formulas and boundaries. Reported policy discussions are not treated as enacted policy.
The evidence base is strongest on what was released, under which licence, at what list price, and on Alibaba’s segment results. It is weaker on private-company economics, causality between open releases and cloud revenue, production market share, enterprise security outcomes and total cost of self-hosting. Conclusions in those areas are explicitly conditional.
Sources & evidence
Evidence cut-off: 21 August 2026. Access dates are the same unless otherwise noted. Vendor benchmark and adoption claims are attributed; they are not treated as independent verification. Adjacent references are separated as [10], [11] for legibility.
- Qwen3.8 official release repository and release log. Source
- Qwen3.8-2.4T-A95B official model card and license. Source
- Alibaba Cloud: Qwen3.8-Max launch announcement, 3 Aug. 2026. Source
- Moonshot AI: Kimi K3 launch, architecture, evaluations and API pricing. Source
- Kimi K3 official pricing page. Source
- Z.ai: GLM-5.2 official release. Source
- Z.ai: GLM-5.3 official release. Source
- Z.ai developer pricing. Source
- DeepSeek official Hugging Face model catalogue. Source
- DeepSeek-V4 official collection. Source
- DeepSeek API models and pricing. Source
- MiniMax-M3 official model card. Source
- MiniMax-M3 community license. Source
- Hugging Face: State of Open Models — Summer 2026. Source
- ATOM Project report on open-model adoption. Source
- Associated Press: Chinese-model consumer and router adoption, July 2026. Source
- Alibaba Group June-quarter 2026 results. Source
- Alibaba Group June-quarter 2026 results, exchange-hosted copy. Source
- Alibaba Cloud: Joe Tsai on open-source monetization. Source
- Alibaba Cloud full-stack AI roadmap and RMB380bn investment plan. Source
- Alibaba Cloud Model Studio token pricing. Source
- OpenAI official API pricing. Source
- Anthropic: Claude Fable 5 availability and pricing. Source
- U.S.-China Economic and Security Review Commission: Two Loops. Source
- Stanford HAI/DigiChina: China’s diverse open-weight ecosystem. Source
- Reuters: reported discussions about possible Chinese restrictions on overseas access. Source
- China NDRC: action plan on global AI cooperation. Source
- ArXiv: U.S. policies and China’s open-AI ecosystem. Source
- ArXiv: Use of Chinese open-weight models in scientific research. Source
- ArXiv: Chinese open models on financial-language tasks. Source
- Z.ai GLM-5.2 official Hugging Face model repository. Source
- MiniMax investor-relations release index. Source