中国开源模型成AI研究默认选择

Nathan Lambert · @natolambert · X·2026-08-24 22:52·1天前
AI 导读

Nathan Lambert用Codex分析50万篇arXiv论文发现,中国开源LLM在AI研究中的提及率从2024年的10%升至如今的约40%,而美国开源模型从30%降至25-30%。Qwen被1/3提及LLM的论文引用,DeepSeek在R1发布后明显跃升;Llama于2025年4月达30%峰值后持续下滑。

Nathan Lambert@natolambert
64AI 编辑部评分,满分 100

中国开源模型成AI研究默认选择

2026-08-24 22:52· 1天前
AI 导读

Nathan Lambert用Codex分析50万篇arXiv论文发现,中国开源LLM在AI研究中的提及率从2024年的10%升至如今的约40%,而美国开源模型从30%降至25-30%。Qwen被1/3提及LLM的论文引用,DeepSeek在R1发布后明显跃升;Llama于2025年4月达30%峰值后持续下滑。

Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research.

In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model.

Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating.

When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power.

Some more observations:

  1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%.
  1. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since.
  1. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research.

The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD

Now over 50% of AI papers, from 10% in 2023.

Other notes: • Gemma and Mistral hover around 5-10%. • Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. • DeepSeek has a clear jump after R1 in Jan. 2025 • Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML

Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.