- by x32x01 ||
The story behind China’s leading AI models did not start with DeepSeek, Kimi, or GLM. Many of the people building these systems were already working on language models, machine learning, and large-scale AI research years before today’s LLM boom.
One useful way to understand this history is to follow the chain:
Research → People → Papers → Models → Companies
Yang earned his Ph.D. from Carnegie Mellon University in 2019. His academic advisors included Ruslan Salakhutdinov and William Cohen, while Jie Tang was his advisor at Tsinghua University.
His name is closely associated with two influential papers from the pre-ChatGPT era:
Years later, Yang became a co-founder of Moonshot AI, the company behind Kimi.
That connection is important because it shows how academic research can eventually become part of a commercial AI product ecosystem.
Tang's work predates the current generative AI wave by many years. He created AMiner, an academic search and knowledge-mining platform that began as ArnetMiner in 2006. AMiner focuses on connecting researchers, publications, conferences, and academic knowledge.
Tang was also involved in the research behind the GLM family of language models.
Z.ai's own history traces the GLM line back to 2021, followed by the open release of GLM-130B in 2022 and ChatGLM in 2023.
So the path looks something like this:
Academic research → AMiner → GLM research → Z.ai
It is another example of how years of research can eventually turn into large-scale AI products.
At Alibaba, Luo was one of the authors of VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation, a 2020 paper focused on cross-lingual pretraining and language understanding.
She later appeared among the authors of DeepSeek's research, including the DeepSeekMoE paper published in 2024. That work explored a Mixture-of-Experts architecture designed to improve expert specialization while reducing the computation required during model use.
Luo later joined Xiaomi's AI efforts and continues to appear on Xiaomi MiMo research papers, including recent technical reports.
Her career illustrates another important pattern:
Research experience can move across companies while the underlying knowledge continues to evolve.
A researcher may work on multilingual NLP at one organization, large language models at another, and foundation models at another.
Rather than coming primarily from an academic NLP background, he built his career around quantitative finance and machine learning.
He co-founded High-Flyer, a quantitative investment firm that used machine-learning techniques for trading. The company later built significant computing infrastructure for AI research. In 2023, Liang launched DeepSeek as an independent AI research organization focused on advanced AI and large language models.
This background helps explain another side of modern AI development.
Building advanced models is not only about algorithms and papers. It also involves:
But a larger pattern appears:
For example, seeing a new model name tells you what was released.
Following the researchers and papers can help you understand where the underlying ideas came from.
When you discover a new model, ask:
For example, Zhilin Yang's research connects earlier Transformer work with Moonshot AI, while Jie Tang's academic work connects AMiner and GLM research with Z.ai. Fuli Luo's publication history shows how research experience can move between Alibaba, DeepSeek, and Xiaomi.
Start with the papers directly related to the models or technologies you want to understand.
Pay attention to:
That is where AI research becomes much easier to understand. 🧠
The modern AI industry is not simply a collection of competing products. It is a growing network of researchers, papers, techniques, models, laboratories, and companies that continuously influence one another.
And once you start following that network, today's AI landscape becomes much easier to connect.
One useful way to understand this history is to follow the chain:
Research → People → Papers → Models → Companies
Zhilin Yang: From NLP Research to Kimi
Before becoming known as a co-founder of Moonshot AI, Zhilin Yang was a researcher in natural language processing.Yang earned his Ph.D. from Carnegie Mellon University in 2019. His academic advisors included Ruslan Salakhutdinov and William Cohen, while Jie Tang was his advisor at Tsinghua University.
His name is closely associated with two influential papers from the pre-ChatGPT era:
- Transformer-XL
- XLNet
Years later, Yang became a co-founder of Moonshot AI, the company behind Kimi.
That connection is important because it shows how academic research can eventually become part of a commercial AI product ecosystem.
Jie Tang: From Academic Search to GLM
Another important name in this story is Jie Tang, a professor at Tsinghua University and a co-founder of Zhipu AI, now branded internationally as Z.ai.Tang's work predates the current generative AI wave by many years. He created AMiner, an academic search and knowledge-mining platform that began as ArnetMiner in 2006. AMiner focuses on connecting researchers, publications, conferences, and academic knowledge.
Tang was also involved in the research behind the GLM family of language models.
Z.ai's own history traces the GLM line back to 2021, followed by the open release of GLM-130B in 2022 and ChatGLM in 2023.
So the path looks something like this:
Academic research → AMiner → GLM research → Z.ai
It is another example of how years of research can eventually turn into large-scale AI products.
Fuli Luo: A Research Path Across Alibaba, DeepSeek, and Xiaomi
Fuli Luo provides another interesting example of how AI researchers move between major organizations.At Alibaba, Luo was one of the authors of VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation, a 2020 paper focused on cross-lingual pretraining and language understanding.
She later appeared among the authors of DeepSeek's research, including the DeepSeekMoE paper published in 2024. That work explored a Mixture-of-Experts architecture designed to improve expert specialization while reducing the computation required during model use.
Luo later joined Xiaomi's AI efforts and continues to appear on Xiaomi MiMo research papers, including recent technical reports.
Her career illustrates another important pattern:
Research experience can move across companies while the underlying knowledge continues to evolve.
A researcher may work on multilingual NLP at one organization, large language models at another, and foundation models at another.
Liang Wenfeng: From Quantitative Trading to DeepSeek
Liang Wenfeng followed a different path.Rather than coming primarily from an academic NLP background, he built his career around quantitative finance and machine learning.
He co-founded High-Flyer, a quantitative investment firm that used machine-learning techniques for trading. The company later built significant computing infrastructure for AI research. In 2023, Liang launched DeepSeek as an independent AI research organization focused on advanced AI and large language models.
This background helps explain another side of modern AI development.
Building advanced models is not only about algorithms and papers. It also involves:
- Large-scale computing
- GPU infrastructure
- Distributed systems
- Data processing
- Model training
- Engineering and optimization
The Bigger Pattern Behind China’s AI Companies
When you look at these researchers individually, their careers can seem unrelated.But a larger pattern appears:
[]A researcher works on an academic problem.
[]The research produces a paper or new technique.
[]Other researchers build on that work.
[]Researchers move between universities, labs, and companies.
[]The ideas become part of larger models.
[]Companies eventually turn those technologies into products.
For example, seeing a new model name tells you what was released.
Following the researchers and papers can help you understand where the underlying ideas came from.
How to Study AI Beyond Model Names
If you are learning AI, machine learning, or NLP, try following the people behind the models.When you discover a new model, ask:
- Who wrote the paper?
- What research came before it?
- Which researchers contributed to it?
- Where did those researchers previously work or study?
- Which other papers did they publish?
- Did the same researchers later build or join an AI company?
For example, Zhilin Yang's research connects earlier Transformer work with Moonshot AI, while Jie Tang's academic work connects AMiner and GLM research with Z.ai. Fuli Luo's publication history shows how research experience can move between Alibaba, DeepSeek, and Xiaomi.
Why Reading Papers Matters
You do not need to read every AI paper from beginning to end.Start with the papers directly related to the models or technologies you want to understand.
Pay attention to:
- The authors
- Their affiliations
- Previous papers they cite
- Earlier versions of the same idea
- The technical problem being solved
- What changed compared with previous approaches
That is where AI research becomes much easier to understand. 🧠
The modern AI industry is not simply a collection of competing products. It is a growing network of researchers, papers, techniques, models, laboratories, and companies that continuously influence one another.
And once you start following that network, today's AI landscape becomes much easier to connect.