Back to blog

50+ Open Source LLM Statistics & Trends (2026)

Explore 50+ open source LLM statistics for 2026, covering ecosystem growth, model usage, performance, inference costs, and leading developers.

Sherlock Xu

Written by Sherlock Xu

Last updated on Aug. 21, 2026

Open source LLM token usage on OpenRouter, led by DeepSeek at 14.37 trillion tokens

Open source LLMs went from a research curiosity to the backbone of real production systems in under three years. They now power coding assistants, agents, and enterprise pipelines, and in some fields they’ve overtaken the proprietary models that defined the early days of the boom.

So how big is the open source LLM ecosystem today? And how fast is it actually growing?

I pulled together the most useful open source LLM statistics I could find, mainly from sources like the Stanford AI Index, Hugging Face, Mozilla, Meta, OpenRouter, Artificial Analysis, Epoch AI, and the Stack Overflow Developer Survey. Each section links to the original sources so you can cite them directly.

Top open source LLM statistics

This article prioritizes first-party reports, research papers, official documentation, and original datasets.

StatisticValueScope and dateSource
Public models hosted on Hugging Face2.96 millionAll public model repositories, not only LLMs; August 2026Hugging Face
Downloads captured by the top 1.5% of repositories99.2%Hugging Face downloads; August 2026Hugging Face
Derivative models built on Qwen151,448Hugging Face repositories; August 2026Hugging Face
Tokens processed by DeepSeek-V4-Flash in one week7.22 trillionOpenRouter weekly ranking, July 27-August 2, 2026TechNode
Open source share of token usageRoughly one-thirdOpenRouter, late 2025OpenRouter
OpenRouter annualized token run rateMore than 4.5 quadrillionAll models on OpenRouter, August 2026Menlo Ventures
Tokens processed by DeepSeek models14.37 trillionOpenRouter, November 2024-November 2025OpenRouter
Developers using open models to add AI features79%Mozilla and SlashData developer survey, May 2026Mozilla
Median inference price declineAbout 50x per yearEpoch AI regression, 2020 to early 2025Epoch AI

What is an open source LLM?

An open source LLM generally refers to a language model that people can download, run, and modify. Compared with proprietary models that are only accessible through an API, open source LLMs give developers much greater control over deployment, customization, and infrastructure.

The term open source is often used loosely. Many models described as open source are actually released as open-weight models. Their weights are publicly available but the license may include restrictions that differ from a traditional open source software license. Because the industry commonly uses “open source LLM” to refer to both categories, this article follows that convention for simplicity.

How many open source LLMs are there?

There is no authoritative global count of open source LLMs. Hugging Face hosted 2.96 million+ public model repositories in August 2026, but that total includes models for text, image, audio, robotics, and other tasks, as well as fine-tunes, adapters, quantizations, and derivatives.

The scale of Hugging Face in August 2026: 2.96 million public model repositories, 1 million datasets, and 1.44 million Spaces

The Hugging Face Hub grew on every axis in the first eight months of 2026:

Hugging Face statisticJanuary 2026August 2026
Public model repositories2.43 million2.96 million
Public datasets711,0001 million
Spaces1.00 million1.44 million

Source: Hugging Face

Almost all of that growth sits in the long tail. 85.6% of models on the Hub have fewer than 200 lifetime downloads, and just 1.5% of repositories account for 99.2% of all downloads. Repository counts measure publishing activity, not usage.

Small models dominate what people actually pull. Among models that declare a parameter count, everything under 1B takes 83% of all-time downloads, while everything above 100B takes 1%. Models larger than 70B made up only 3% of 2026 download volume.

Models under 1B parameters take 83% of all-time Hugging Face downloads, against 1% for models above 100B

Two more signals from the same report are worth noting. The two organizations that published the most new model repositories in 2026 were AMD and Nvidia, each with more than 200, mostly hardware-optimized conversions rather than new architectures. And GGUF library repositories, the format that makes large models runnable on consumer hardware through llama.cpp, grew 464% between January and August 2026.

An earlier Hugging Face report, published in spring 2026, added community and contributor detail that the summer edition does not restate:

Hugging Face statistic (spring 2026 report)ValueWhat it measures
Registered users13 millionPlatform community size
Fortune 500 companies with verified accountsMore than 30%Organizational presence, not confirmed production adoption
Industry share of model development37%Down from roughly 70% before 2022
Downloads attributed to unaffiliated developers39%Up from 17% before 2022
Mean size of a downloaded model20.8B parametersUp from 827M in 2023
Median size of a downloaded model406M parametersUp from 326M in 2023
Mean engagement period after releaseAbout 6 weeksHow long models typically sustain attention

Source: Hugging Face

Open source LLM adoption and usage

Open models are no longer the minority. They handle a majority of all tokens routed through OpenRouter as of mid-2026, and the seven highest-volume models on the platform all ship open weights. That is a reversal, not an increase: open models held roughly one-third of OpenRouter token volume in late 2025.

The geographic split moved just as fast. Chinese open source models serve 46.4% of routed tokens, against 35.7% from US providers. US models from Google, OpenAI, and Anthropic held about 70% of token share in June 2025 and roughly 30% a year later.

Chinese open models rose from 1.2% of OpenRouter token volume in late 2024 to 46.4% in mid-2026

A single model shows the scale involved. DeepSeek-V4-Flash topped the OpenRouter weekly usage ranking for July 27 to August 2, 2026 with 7.22 trillion tokens processed in one week, roughly half of what the entire DeepSeek family processed over the preceding year.

The year-long study behind the late-2025 baseline analyzed more than 100 trillion tokens, covering November 2024 through November 2025.

Open model developerTokens on OpenRouter, Nov 2024 to Nov 2025
DeepSeek14.37 trillion
Qwen5.59 trillion
Meta Llama3.96 trillion
Mistral AI2.92 trillion
OpenAI1.65 trillion
MiniMax1.26 trillion
Z.ai1.18 trillion
TNGTech1.13 trillion
Moonshot AI0.92 trillion
Google0.82 trillion

Source: OpenRouter

DeepSeek processed about 2.6 times as many tokens as Qwen, the second-largest open source family in the study. By late 2025, however, no individual model consistently accounted for more than 20%-25% of open source model tokens. Usage had spread across five to seven competitive models.

The same study found that:

  • Models developed in China rose from as little as 1.2% of weekly token volume in late 2024 to nearly 30% in some weeks.
  • Chinese open source models averaged 13.0% of weekly OpenRouter token volume over the full period.
  • Open source models developed outside China averaged 13.7%.
  • Proprietary models developed outside China retained an average share of about 70%.
  • Roleplay represented about 52% of open source model tokens, while programming was the second-largest category at roughly 15%-20%.

OpenRouter itself kept compounding after the study period. By August 2026 the platform reported 10 million users, more than 500 models, more than 80 providers, and an annualized run rate above 4.5 quadrillion tokens. That is three years of roughly 33% month-over-month growth, a doubling every 11 weeks, and a 30,000x increase since launch.

The routing layer is now valuable enough to buy. Stripe agreed to acquire OpenRouter on August 19, 2026, three years after the platform launched. Terms were not disclosed.

Developer adoption of AI tooling is now near-universal. 84% of respondents to the 2025 Stack Overflow Developer Survey use or plan to use AI tools, up from 76% the year before, and 51% of professional developers use them daily. The 2026 edition of the survey was still in the field as of mid-2026, so the 2025 results remain the most recent published figures.

Sources: OpenRouter, Mozilla, Menlo Ventures, TechNode, Stack Overflow, Yahoo

The gap between adoption and production

High adoption does not mean high deployment. Mozilla published the first edition of The State of Open Source AI in July 2026, combining a SlashData developer survey fielded in May 2026 with an analysis of the OpenRouter token dataset. The headline finding is a split between who tries open models and who ships them.

Mozilla findingOpen modelsClosed models
Developers using them to add AI functionality79%71%
Teams that reach productionAbout 53%63%

About half of developers use both categories together, so the two are largely complementary rather than competing. The report attributes the production shortfall to deployment, tooling, and operational infrastructure rather than model quality.

Source: Mozilla

Popularity depends on the metric. Downloads measure distribution, derivative repositories measure how often developers build on a model, and routed tokens measure hosted API usage.

On Hugging Face, Qwen has become by far the largest ecosystem for derivative work: 151,448 derivative repositories as of August 2026, which is 2.6 times the total Meta footprint and 4.7 times the Llama repositories specifically. That count grew by 180 to 210 new repositories per day through the first seven months of 2026.

The gap is widest in the formats people run locally. Qwen has 28,531 GGUF conversions on the Hub, published almost entirely by the community rather than by Alibaba, which released only 54. Those conversions pull 39.6 million downloads per month, against 20.8 million for Gemma and 7.5 million for Llama.

These figures count downloads and repositories, not unique users or production deployments. Automated downloads, mirrors, quantizations, and repeated pulls can all affect the totals.

You can compare current model releases, licenses, context windows, and hardware requirements on the OpenLLMStack models page.

Sources: Hugging Face

Open vs. closed model performance

The leading open model trailed the leading closed model by 3.3% on the Arena leaderboard in March 2026, according to the Stanford AI Index. The gap had been only 0.5% in August 2024 before reopening during 2025.

DateTop closed-vs-open performance gap
August 20240.5%
March 20263.3%

Source: Stanford AI Index 2026

The gap narrowed again in July 2026. In the Kimi K3 technical report, Kimi K3 scored 57.1 on the Artificial Analysis Intelligence Index, against 59.9 for Claude Fable 5, a gap of 2.8 points. Kimi K3 ranked fourth among 580 leaderboard entries, or third when GPT-5.6 Sol effort variants were counted as one model, behind Claude Fable 5 and GPT-5.6 Sol.

Kimi K3 also became the first open model ever to top WebDev Arena, ranking first of 99 models at 1,678 Elo, and it placed first in the Arena Frontend Code evaluation at 1,679 points in blind developer testing.

The correct conclusion is not that open models have permanently reached parity. Aggregate indices hide task-level gaps that remain wide:

BenchmarkLeading open source modelsLeading proprietary model
HLE (expert-level questions across disciplines)34-36%44-45%
TerminalBench Hard (agentic coding)43-46%61%
CritPt (research-level physics)4-12%27%

Open models are competitive on mainstream coding and reasoning work and still behind on the hardest research and long-horizon agentic tasks. The gap also moves with every release cycle in both directions.

Sources: Artificial Analysis, Kimi K3 technical report, Moonshot AI

Where open source models come from

Models developed in China made up 41% of Hugging Face downloads in 2025, the largest share attributed to a single country. China surpassed the United States in both monthly and overall downloads during the year.

The contributor mix changed at the same time. The share of model development attributed to industry fell from roughly 70% before 2022 to 37% in 2025, while unaffiliated developers grew from 17% to 39% of downloads.

OpenRouter recorded a similar geographic shift, although through a different metric. Chinese open source models rose from 1.2% of weekly token usage in late 2024 to nearly 30% in some weeks of 2025, then to more than 45% of weekly traffic by April 2026.

The size distribution now splits along the same line. Through the first seven months of 2026, Chinese labs released models with a monthly parameter ceiling between 754 billion and 2.78 trillion, while US labs stayed under 130 billion in five of seven months, with the Nvidia Nemotron 3 Ultra at 550B and Inkling at 952B as the exceptions.

Licensing has stayed permissive. Across more than 178 Chinese model releases above 20B parameters, 59% shipped under Apache 2.0 and 22% under MIT, with non-commercial restrictions rare outside the very largest models.

Sources: Hugging Face, Mozilla, OpenRouter

The falling cost of LLM inference

An Andreessen Horowitz analysis found that the price of inference at a fixed level of MMLU performance fell by roughly 10x per year. At an MMLU score of 42, the cheapest observed price dropped from $60 per million tokens in 2021 to $0.06 in 2024, a 1,000-fold decline.

LLM inference cost at an MMLU score of 42 fell from $60 to $0.06 per million tokens between 2021 and 2024

That historical estimate has limitations: it relies on MMLU, averages input and output pricing, and covers selected models from OpenAI, Anthropic, and Meta. It is useful as a directional cost trend, not a law that guarantees another 10x decline every year.

Epoch AI measured the same trend with a wider dataset and found the rate depends heavily on which capability you hold fixed. Across benchmarks, the price decline ranges from 9x to 900x per year, with a median of about 50x per year between 2020 and early 2025. Matching GPT-4 performance on PhD-level science questions got 40x cheaper per year.

A separate MIT FutureTech study, which combined benchmark data from Epoch AI with historical price data from Artificial Analysis, found that models in the highest GPQA-Diamond performance bin fell 31x in price per year, while models in the lowest bin fell 1.7x per year.

The lesson is that “inference is getting cheaper” is not one number. In the GPQA-Diamond analysis, prices fell much faster in higher-performance bins than in lower-performance bins; the study says different economic or technical factors may drive the gap.

Prices do not only fall

August 2026 supplied the counterexample. DeepSeek raised API prices on August 16, 2026 and replaced flat pricing with a peak and off-peak schedule. Peak hours run 01:00 to 04:00 and 06:00 to 10:00 UTC, and off-peak rates are half the peak rate.

ModelCached inputInput, cache missOutput
DeepSeek-V4-Flash, off-peak$0.007$0.22$0.66
DeepSeek-V4-Flash, peak$0.014$0.44$1.32
DeepSeek-V4-Pro, off-peak$0.022$0.66$1.98
DeepSeek-V4-Pro, peak$0.044$1.32$3.96
GPT-5.6 Luna$0.02$0.20$1.20
GPT-5.6 Terra$0.20$2.00$12.00
GPT-5.6 Sol$0.50$5.00$30.00

Prices are per 1 million tokens. Measured against the previous flat rates, DeepSeek-V4-Flash output went up 2.4x off-peak and 4.7x at peak, from $0.28. The steepest increase hit cached input on V4-Pro, which rose from $0.003625 to $0.044 at peak, an increase of more than 1,100%.

DeepSeek-V4-Flash output price rose from a flat $0.28 to $1.32 per million tokens at peak hours

The comparison against OpenAI changed with it. DeepSeek-V4-Flash is still far cheaper than the OpenAI flagship: off-peak input is 95.6% cheaper and output is 97.8% cheaper than GPT-5.6 Sol. But at the low end of the lineup the advantage has now disappeared. Against GPT-5.6 Luna, off-peak V4-Flash input at $0.22 is 10% more expensive than Luna at $0.20, and at peak hours V4-Flash costs more than Luna on both input and output.

An open model is no longer automatically the cheapest token on the market. It depends on the tier you compare against and, now, the hour of the day.

Serving efficiency comes from more than cheaper hardware. Quantization, batching, caching, sparse architectures, and optimized kernels all contribute. OpenLLMStack tracks major techniques on the inference optimizations page and the engines that implement them in the inference directory.

Sources: Andreessen Horowitz, Epoch AI, MIT FutureTech, DeepSeek API pricing, OpenAI API pricing

Top open source LLM API providers

You do not need your own GPUs to run an open model. A growing set of serverless inference providers host the leading open weights behind a single API, billed per token, so you can switch models without managing infrastructure.

ProviderKnown for
Together AIBroad catalog of 200+ open models behind one unified API
Fireworks AIFast, production-grade serving of popular open models
GroqCustom LPU silicon for very low latency and high throughput
DeepInfraLow-cost, pay-per-token hosting of open models
BasetenCustom deployment and autoscaling for open weights
ModularShared endpoints and reserved dedicated GPU capacity, optimized by MAX across GPU vendors
Hugging FaceHub-native inference endpoints next to the models

For the cheapest access to a single model, the first-party API from the model maker is often the lowest-priced option, such as the APIs from DeepSeek and Alibaba for their own models. Aggregators like OpenRouter route one request across many of these providers so you can compare price and speed.

DeepSeek R1: the breakout moment for open source LLMs

No single release did more for the profile of open source LLMs than DeepSeek R1. The model launched on January 20, 2025, and within days the DeepSeek app climbed to No. 1 on the U.S. Apple App Store, displacing ChatGPT and topping the charts in more than 50 countries.

The download surge was almost vertical. The app reached 2.6 million downloads across the App Store and Google Play by the Monday after launch, with more than 80% of all downloads coming in the previous seven days, and Appfigures data ranked the app No. 1 worldwide.

The market reaction was just as dramatic. On January 27, 2025, Nvidia lost about $589 billion in market value, the largest single-day loss for any company in history, after DeepSeek showed that a frontier-grade open model could reportedly be trained for around $5.6 million.

Nvidia lost about $589 billion in market value on January 27, 2025, the largest single-day loss for any company

You can trace these milestones on the OpenLLMStack timeline.

Sources: TechCrunch, Bloomberg, Forbes

Frequently asked questions

How many open source LLMs are there?

No authoritative organization maintains a global count of open source LLMs. Hugging Face hosted 2.96 million public model repositories in August 2026, but that figure covers all model types and includes derivatives, adapters, and quantizations. It is not a count of unique LLMs, and 85.6% of those repositories have fewer than 200 lifetime downloads.

Which open source LLM is the most popular?

It depends on the metric.

  • Meta reported 1.2 billion cumulative Llama downloads in April 2025.
  • Qwen has the largest derivative ecosystem on Hugging Face, with 151,448 derivative repositories in August 2026, 4.7 times the Llama total.
  • DeepSeek leads routed usage. DeepSeek-V4-Flash was the single most-used model on OpenRouter for the week of July 27 to August 2, 2026, at 7.22 trillion tokens.
What is the best open source LLM in 2026?

There is no single best open source LLM in 2026. Some models are good at coding, others at reasoning, long-context processing, or multilingual tasks. The right model depends on your workload, hardware, latency requirements, and budget.

If you self-host an open source LLM, you can also adapt it to your domain by fine-tuning the model on proprietary data. This can significantly improve performance for specialized tasks, such as legal analysis, healthcare, finance, or customer support, helping the model outperform a general-purpose foundation model in your specific domain.

By mid-2026, several open families compete at or near the proprietary frontier, each with a different strength.

ModelMakerReleasedStrongest at
Kimi K3Moonshot AIJuly 2026Frontend and agentic coding, first open model to top WebDev Arena, 1M token context
GLM-5.3Z.aiAugust 2026State-of-the-art coding and agentic engineering
MiniMax-M3MiniMax2026Frontier coding and agentic work, native multimodal and computer use
DeepSeek-V4-Pro-0813DeepSeek2026Reasoning and coding with adaptive effort modes and strong world knowledge
Qwen3.8-2.4T-A95BAlibabaAugust 2026The largest open Qwen release, a 2.4T mixture-of-experts model with 262K native context
MiMo-V2.5-ProXiaomi2026Token-efficient coding agents with long-context reasoning

Kimi K3 is the largest open source model released to date. Moonshot AI published the full weights on July 27, 2026: 2.8 trillion total parameters in a mixture-of-experts design that activates 16 of 896 experts per token, 104 billion active parameters, with native visual understanding and a 1 million token context window.

A common thread runs through the 2026 leaders: state-of-the-art coding, agentic tool use, and context windows that now stretch to 1 million tokens or more.

For the full list with parameters, licenses, context windows, and recommended GPUs, see the OpenLLMStack models page. For a head-to-head comparison of the current leaders on price, license, and context, along with when to skip each one, see the best open source LLMs in 2026.

How much open source LLM usage is there?

Open source models handle a majority of token volume on OpenRouter as of mid-2026, up from roughly one-third in late 2025. The seven highest-volume models on the platform all ship open weights. This is strong evidence of adoption on that platform, but self-hosted usage and traffic on other providers are not included.

Are open source models actually used in production?

Not as often as adoption figures suggest. Mozilla found that 79% of developers adding AI functionality use open models, but only about 53% of those teams reach production, compared with 63% for closed models. The report attributes the shortfall to deployment, tooling, and operational infrastructure rather than model quality.

Is DeepSeek still the cheapest option?

Not always. DeepSeek raised API prices on August 16, 2026 and moved to peak and off-peak rates. DeepSeek-V4-Flash remains roughly 96% cheaper than GPT-5.6 Sol on input, but off-peak input at $0.22 per million tokens is now 10% more expensive than GPT-5.6 Luna at $0.20, and at peak hours V4-Flash costs more than Luna on both input and output.

Conclusion

Open source AI in 2025 and 2026 is defined by four trends: models that now rival closed systems on quality, open models taking a majority of routed token traffic, a center of gravity shifting toward Chinese and independent developers, and a first real test of the low-price assumption as DeepSeek raises rates. Open source models are no longer the budget option. For a growing share of teams, they’re the default.

If you’re building with open models, OpenLLMStack tracks current releases, inference engines, optimization techniques, and agent frameworks in one place.