OpenRouter releases "AI Status Quo Report Based on 100 Trillion Tokens of Data"
In today's digital age, artificial intelligence (AI), especially large language models (LLMs), is changing our world at an unprecedented pace. However, we still lack sufficient information regarding the real-world applications and impact of these models...
In today's digital age,artificialintelligent(AIEspecially large language models (LLM(s) is changing our world at an unprecedented pace. However, we still lack systematic empirical research on the real-world use and impact of these models. Based on this, OpenRouter and a16z jointly released the "State of AIThe report, "An Empirical 100 Trillion Token Study with OpenRouter," analyzes the authenticity of over 100 trillion tokens on the OpenRouter platform.practicalUser interaction data, in-depth exploration LLMThe report reveals the actual usage of [the technology/product] globally.open sourceThe competitive landscape with closed-source models, the rise of inference optimization models, and the dominance of programming and role-playing.AI The globalization trend of usage, and key issues such as user retention and cost dynamics, are crucial for understanding... LLMThis provides a fresh perspective and data support for the current status and future development direction of s.
Background and Significance of the Research
An in-depth investigation jointly released by OpenRouter and a16z reveals...AIThe field is undergoing an unprecedented "great divergence." The report is based on the true value of 100 trillion tokens on the OpenRouter platform.practicalUser interaction metadata, covering the period from the end of 2023 to November 2025 (with a core focus on nearly one year), encompasses over 300 models and more than 60 providers globally, making it the largest dataset to date.LLMEmpirical research. Previously, assessments...AIMetrics of model impact are often limited to academic benchmark tests or claimed user numbers. OpenRouter offers the first God's-eye view based on real computing power consumption, revealing exactly how developers and enterprises are using [the system].AI.
open sourceThe Rise of Models
open sourceComparison with closed-source models:
- Closed-source modelOpen still dominates high-value scenarios, accounting for approximately 70% of total token usage, especially in enterprise-level and regulated tasks (such as financial compliance and medical consultation), where users tend to choose Open.AIProprietary models from vendors such as Anthropic and Google (e.g.)Claude 3.7 SonnetGPT-5 Pro).
- open sourceModelBy the end of 2025, its share will be stable at 30%, and the growth is "sustainable". It is not a short-term experimental use, but a deep integration into the production environment (such as the long-term maintenance of peak traffic after the release of DeepSeek V3 and Qwen 3 Coder).
Chinaopen sourceModel explosion:
- dataAt the end of 2024, the usage share of Chinese models was only 1.2%; by the second half of 2025, in some weeks, the usage of Chinese OSS models (such as DeepSeek, Qwen, MiniMax, Kimi, GLM, etc.) even accounted for nearly 30% of all model traffic.
- Core advantages:
- Fast iteration speedDeepSeek and the Qwen family of software utilize "high-frequency updates" (such as 1-2 new versions per month).fastIt adapts to different scenarios (such as long context programming and Chinese role-playing).
- Strong scene adaptabilityIn areas such as Chinese language processing, role-playing, and code generation (e.g., Qwen 3 Coder), its performance is close to or even surpasses that of other code generators.open sourceModels (such as Meta LLaMA 3.3).
Model size preference: "Medium-sized models" become the new mainstream
- Small model (<15B parameters)Its market share continues to decline. Although there are new products such as Google Gemma 3.12B, their limited capabilities make it easy for users to "switch frequently" and it is difficult to form stable stickiness.
- Medium-sized model (parameters 15B-70B)Emerging in 2025, representative models such as Qwen2.5 Coder 32B and Mistral Small 3 achieved an optimal balance between "capabilities (reasoning, code)" and "efficiency (cost, latency)," becoming the preferred choice for developers.
- Large model (>70B parameters)Diversified needs mean it's no longer the "only option" – Qwen3 235BGPTWhile high-performance devices like the OSS-120B are powerful, their high cost limits their use to complex tasks such as system architecture design.
The use cases are "polarized," with programming and role-playing dominating traffic.
open sourceModels: Role-playing games make up half the market.:
- data:open sourceIn the model, 52% of the tokens are used for role-playing, far exceeding the "productivity scenarios" (programming accounts for 15%-20%, and writing accounts for 5%).
- Scene detailsThis includes game NPC dialogues, fan fiction creation, virtual companion interactions, etc., with the core requirements being "flexible responses, emotional subtlety, and low content restrictions".open sourceThe model can be freely fine-tuned and is not constrained by commercial security filters (such as DeepSeek Chat V3 which supports custom character settings, and Qwen role-playing model which can maintain consistency in long conversations).
All platforms: Programming becomes the "number one scenario," with the fiercest competition.:
- Explosive growthThe percentage of tokens used in programming tasks surged from 11% at the beginning of 2025 to over 50% by the end of the year, becoming...LLMThe most core productivity applications (such as code generation, debugging, and code library understanding).
- Market structure:
- AnthropicClaudeseries)It had long monopolized the programming scenario with a market share of over 60%, but in November 2025, it fell below 60% for the first time.
- Rise of the ChaserOpenAIFrom 2% to 8%, Google remained stable at 15%, and China's OSS (Qwen Coder, DeepSeek R1) also saw growth.fastPenetration, MiniMax and other new players saw significant fluctuations in weekly market share (even small changes in model quality/latency affected selection).
AgentIC reasoning becomes a new paradigm.AIFrom "Generator" to "Analysis Engine"
Inference models: accounting for over 50% within six months:
- Paradigm shiftOpen in December 2024AI The o1 model (codenamed "Strawberry") has been released, marking...LLMFrom "single-channel text generation" to "multi-step internal reasoning"—O1 enhances mathematical logic and multi-step decision-making capabilities through "latent planning and iterative optimization," and subsequently...GPT-5、Claude 4.5Gemini 3. Follow up, and by the end of 2025, the proportion of tokens in the reasoning model will exceed 50%.
- Head model:xAIGrok Code Fast 1 has emerged as a dark horse, accounting for approximately 25% of the tokens used in reasoning scenarios, surpassing Google. Gemini 2.5 Pro (20%), OpenAI GPT-OSS-120B (15%).
Tool calls and long contexts:AgentThe two pillars of IC:
- Tool usage becomes routineThe proportion of tool call requests steadily increased in 2025 (excluding the abnormal peak in May), Anthropic Claude 4.5 Sonnet (after the end of September)fast(30%)AI Grok Code Fast (15%) is the main contractor, marking...AIThe role has shifted from "interlocutor" to "system component" (such as calling APIs to retrieve data or executing code).
- Context length skyrocketed:
- Average prompt lengthThe number of tokens will increase from 1,500 at the beginning of 2024 to 6,000 at the end of 2025 (a 4-fold increase).
- completion lengthIncrease from 150 Tokens to 400 Tokens (3x increase).
- Core DriverProgramming tasks (code library understanding and debugging require 20,000+ token inputs), while other scenarios (such as document analysis) have a gradual increase in context.
LLMWhat was it used for?
Programming becomes the number one core task:
- dataThe proportion of programming-related request tokens surged from 11% at the beginning of 2025 to over 50% by the end of the year, becoming the most stable growing category. This includes scenarios such as code generation, debugging, and data script writing, marking a significant shift in the market.LLMShifting from exploratory dialogues to application tools, deeply embedded in developer workflows.
- Market competition landscape:
- Anthropic ClaudeseriesIt had long monopolized the programming scenario with a market share of over 60%, but in November 2025, it fell below 60% for the first time.
- OpenAIIts share increased from 2% to 8%.
- GoogleIt remains stable at 15%.
- China OSS (Qwen, Z).AIAnd MiniMax and other rising stars:fastPenetration means that developers are highly sensitive to even minor changes in model quality and latency.
Internal structure of twelve common tasks:
- role play:occupyopen sourceOf the 52% of model token usage, 60% is concentrated in "games/role-playing games," while the proportions of author resources (15.6%) and adult content (15.4%) are similar, indicating that it is not just casual chatting, but rather has a clear type of scenario requirement.
- Programming subdivisionMore than two-thirds of the traffic belongs to "programming/other", indicating a wide range of general needs; the proportion of development tools (26.4%) has increased, showing a trend towards specialization.
- Long-tail domain characteristics:
- scientific field80.4% of queries focused on "Machine Learningandartificialintelligent“With Yuan”AIThe focus is on problems, rather than traditional STEM topics.
- health fieldThe distribution is the most dispersed, with sub-tags accounting for no more than 25% each, covering medical research, consulting, diagnosis, etc., and the demand is complex and sensitive.
- Finance and legal fieldsThe labels are scattered and lack mature, dedicated ones.LLMThe workflow and applications are still in the exploratory stage.
LLMWhat are the differences in usage across different regions?
Regional usage distribution:
- North AmericaIt remains the largest market (accounting for 47.22%), but its share continues to decline.
- AsiaIts share more than doubled from 13% to 31%, making it the fastest-growing consumer market.
- EuropeIt remains stable at 15%-20%.
- National levelThe United States leads by a wide margin with 47.17%, followed by Singapore (9.21%), Germany (7.51%), and China (6.01%). More than 60 countries worldwide participated.LLMuse.
Language distribution:
- EnglishThe percentage of users who use English is 82.87%, reflecting the large user base of developers and the prevalence of the English model.
- Simplified ChineseEnglish accounted for 4.95%, followed by Russian (2.47%) and Spanish (1.43%), indicating a gradual increase in demand for non-English languages.
User retention patterns: The "Cinderella Glass Shoe Effect" determines long-term user engagement.
Core phenomenon:
- Most models show a high user retention rate with "high churn".fast"attenuation" characteristics,However, a basic user base exists.The workload of these users is deeply integrated with the model, creating economic and cognitive inertia that makes it difficult to migrate even when a new model is released.
- “The Cinderella glass slipper effectIf a new model can accurately match high-value workloads that are not being met, it can lock in a basic user base; otherwise, it will fail to establish stable user engagement, and users will continue to explore alternative models.
Typical retention patterns:
- First-mover advantage:likeClaude 4 Sonnet,Gemini The early user base of 2.5 Pro formed a stable match in the early stages of the model's release, and its retention rate has been higher than that of subsequent user groups for a long time.
- Mismatch:likeGemini The 2.0 Flash and Llama 4 Maverick systems did not establish a high-performance user base, resulting in low user retention rates across all batches.
- boomerang effectThe reason why DeepSeek users returned after churn was that after testing competing products, they confirmed that DeepSeek had advantages in professional performance and cost.
Cost and usage dynamics, significant market segmentation
open sourceComparison with closed-source models:
- Closed-source model: Focused on the high-cost, high-usage quadrant, mainly handling high-value tasks.
- open sourceModelIt mainly focuses on low-cost, high-usage areas and mainly handles large-volume, cost-sensitive tasks.
Cost-usage four-quadrant distribution:
- High-end workloads (high cost, high usage)For technical and scientific tasks, users are willing to pay a premium for complex reasoning (such as system architecture design).
- Mass market driven (low cost, high usage)Programming, role-playingopen sourceThe model dominates due to its cost advantage, and user engagement is comparable to that of professional tasks.
- Specialized experts (high cost, low usage)Finance, healthcare, and law are niche sectors with high-risk demands and require extremely high accuracy.
- NichepracticalTools (low cost, low usage)Translation and trivia are highly commoditized, with ample alternatives available.
Market pricing and user behavior characteristics:
- Low demand elasticityAt the macro level, price changes have a relatively small impact on usage; at the micro level, enterprise users are willing to pay higher prices for mission-critical tasks (such as...).GPT-4、Claude 3.7 Sonnet), developers and hobbyists are cost-sensitive.
- Jevons Paradox SignsLow-cost models (such as)Gemini Due to improved efficiency, Flash 2.0 and DeepSeek V3 were widely integrated into more tasks, resulting in a surge in total token consumption.
- Model hierarchical competitionThe market presents four archetypes—high-end leaders (Claudeseries),High efficiencyGiants (Gemini Flash, DeepSeek V3), Long Tail Model (Qwen 2-7B), Advanced Expert (GPTWith the -5 Pro, differentiation (latency, context length, reliability) remains the core competitive advantage.
Core Insights:
- Multi-model ecosystems become mainstreamNo single model can cover all scenarios; closed-source models dominate high-value tasks.open sourceThe model occupies a low-cost, high-capacity scenario, and developers need to flexibly integrate multiple models.
- Applications Beyond ProductivityThe high proportion of entertainment scenarios such as role-playing highlights the potential of narrative and emotional interaction applications for consumers. Model evaluation needs to balance consistency and dialogue experience.
- AgentIC reasoning becomes a new paradigmThe focus has shifted from single-round generation to multi-step planning and tool usage, and the evaluation criteria have shifted from language quality to task completion.
- Globalization and regionalization in parallelThe rise of the Asian market, Chinaopen sourceModels have become an important force.LLMIt needs to be compatible with multiple languages and cultural scenarios.
- Retention is more important than growthUnder the "crystal shoe effect," only models that accurately match high-value workloads can build long-term user stickiness.
limitation:
The data only covers the OpenRouter platform and does not include local enterprise deployments or internal systems; some analyses rely on proxy metrics (such as tool calls to identify inference tasks), and the results are indicative rather than absolute.
Future Trends:
LLMAs it will be deeply integrated into the global computing infrastructure, the focus of competition will shift from the scale of model parameters to task completion efficiency, cost control, and scenario adaptability.intelligentbodyReasoning will gradually mature and drive...LLMUpgrade from a "generation tool" to a "decision engine".
- Report official websitehttps://openrouter.ai/state-of-ai
CodeFlying Overseas Version Test: One-Sentence Application GenerationautomaticDeployment and launch
Real-world testing shows that the Zhipu GLM-4.6V is the most powerful domestically produced GLM.MultimodalAgentBase model