When Deepseek launched in January, it disrupted the AI industry. The company rocketed to No.1 on Apple’s App Store and shook the U.S. stock market, causing the NASDAQ to plunge 3.1%. The key to its sensational debut was Deepseek’s innovation in model architecture. Instead of improving the raw computational power of rival U.S. models, DeepSeek invested heavily in optimizing model design and training methodologies, focusing on using limited resources in the most efficient way.

The capabilities of DeepSeek’s large language models (LLMs) nearly match global leaders like Anthropic and Open AI, but at a much lower cost. To achieve this, Deepseek relied on architecture innovations such as the mixture of experts (MoE), multi-head latent attention (MLA), group relative policy optimization (GPRO), and distillation. These made possible Deepseek’s smaller size and lower cost, while maintaining high performance.

In addition to technical breakthroughs, DeepSeek is unique in its open-source approach to its LLMs. Open-source models, like Meta’s LLama, are free and open to everyone. Consequently, people can recreate and redistribute the code, contributing to a transparent and collaborative ecosystem. In contrast, proprietary models like Open AI’s ChatGPT do not release the underlying code, forcing users to pay to access the model.

“Both have drawbacks and benefits,” Jimmy Goodrich, a senior advisor at RAND, told US-China Today. “Some would argue that open-weight, open-source models are more secure because you can know what the code is; others argue they’re less secure because they’re proliferating these powerful capabilities and they might be misused by bad actors.”

While there’s a debate between open and proprietary sources, DeepSeek’s firm stand on open-source reflects its mission statement: “Unravel the mystery of artificial general intelligence (AGI) with curiosity,” words on its webpage, expressing the goal of building cutting-edge models instead of commercializing the applications.

“Open source is more of a cultural behavior than a commercial one, and contributing to it earns us respect,” said the CEO of DeepSeek, Liang Wenfeng, in an interview on Waves.

Liang’s unique background sheds light on Deepseek’s prioritization of R&D over monetization. After graduating from Zhejiang University with engineering degrees, Liang did not enter the tech industry immediately. Instead, he worked as a quantitative trader in the stock market, where he sought to apply AI to finance. In 2015, he co-founded the quantitative hedge fund, High-Flyer, relying on math and AI to predict markets and make investment decisions. By 2019, Liang was the majority owner of High-Flyer, which had earned over 10 billion yuan in assets.

Liang’s career path is “very unusual for someone in a tech space,” said Cory Combs, the associate director at Trivium China, a Beijing-based political economy research group. “Usually, tech companies are making the money by commercializing their AI applications. He skipped all that. He went straight to traditional asset management, made a ton of money; but he had the technical background while he was doing this.”

In May 2023, when Liang dove into AI full-bore and founded DeepSeek, he had the resources to self-fund the team, directing it to focus solely on research, rather than pursuing commercialization or pleasing investors. The organizational culture of DeepSeek also contributed to its rise. Compared to other Chinese tech giants like Alibaba and Huawei, where workers follow the traditional “996 work regime,” DeepSeek applied flat management, academic-style collaboration and autonomy to their workforce to encourage innovation. The team hires young talent who are passionate about technology, as opposed to seasoned workers who seek answers through past experiences.

“It was the world’s best funded academic group, effectively, without being an academic group per se,” Combs said. “I think that’s a large part of what made them very effective.”

Deepseek’s business model and corporate culture created an atmosphere for innovation and disruption. But what does DeepSeek, which may seem like an outlier in the industry, tell about China’s innovation capability in general? Given that DeepSeek’s global success and core team developers, including Liang, are based in China and attended universities there, some argue that China has surpassed the U.S. in its ability to innovate and nurture talent.

The reality is more complicated, however. In October 2022, the Biden Administration cut off exports of the most advanced semiconductors to all Chinese firms, in an effort to impede China’s development of chips and AI. Fortunately for DeepSeek, Liang had acquired at least 10,000 Nvidia A100s chips by 2021, well before the sanction, which he was able to put to use. But lack of access to advanced chips remains a concern for the company. “Money has never been the problem for us,” Liang told Waves. “Bans on shipments of advanced chips are the problem.”

U.S. export controls on semiconductors may have inhibited China’s development in cutting-edge chips, but the controls also catalyzed China’s indigenous innovation. The rise of DeepSeek suggests that even when an AI company lacks the funds and GPUs of a rival like OpenAI, there are other paths toward disruption, such as architectural improvement. In other words, hegemony in computing power does not necessarily guarantee technological hegemony.

In what is certain to be a prolonged AI race, China is endeavoring to produce its own chips to reduce reliance on the U.S. and dull the impact of U.S. sanctions. Although China cannot produce all of its chips domestically yet, Combs believes that Chinese chips are on the near horizon, lower in quality perhaps, but also less expensive. “Second best might be good enough and China can make second best at most things at once,” he said. Part of the reason for China ramping up its chip development so quickly, Combs added, is “this geopolitical impetus.”

Beyond hardware, Beijing has made significant investments in its indigenous technological development. In 2017, China launched the “New Generation Artificial Intelligence Development Plan,” which aims to become a global leader in AI by 2030. Beijing is also demonstrating its AI-commitment through guidance funds. “Over the past few years, there were over 2,000 guidance funds allocated for different technologies,” Combs stated. After the success of DeepSeek, China announced a $138 billion government-backed venture fund to grow hard technology sectors, according to Reuters.

In addition to direct funding, Beijing has also focused on strengthening AI curriculum in higher education. Many universities, middle schools, and even primary schools, have begun offering AI courses. “While the exam-centric environment is still deeply engrained due to the intense competition for college admissions, it’s evident that schools are gradually placing more emphasis on nurturing critical thinking and creative problem-solving,” said Tim Yang, a computer science graduate from a top university in China.

Beijing’s AI-friendly initiatives may not be the only reason for DeepSeek’s success, considering the company’s unique formation. However, Deepseek’s developers and engineers have certainly benefited from these broad government initiatives, including funding, STEM education, and strategic industrial planning.

Recently, tech giants like Baidu and Google launched new AI models to compete with DeepSeek. In this global, tense, and rapid changing AI race, DeepSeek’s trajectory remains uncertain. Nonetheless, its rise shows a break in the AI paradigm, the very definition of innovation.