21 August 2026

The People’s Republic of Data: How China Is Turning Data into Capital

U.S.-China Economic and Security Review Commission

China designated data as a fifth factor of production in 2020, establishing state-led data exchanges, accounting rules, and commercialization incentives to build an integrated national data market. Following a multi-year regulatory crackdown on private technology platforms, Beijing's National Data Administration is orchestrating a top-down data economy to fuel artificial intelligence models, industrial productivity, and military capabilities.

The state-driven initiative focuses on high-value enterprise, operational, and physical-world sensor records that remain scarce following the exhaustion of public web scrapings. By late 2025, Director Liu Liehong announced over 4,000 interconnected exchanges, operators, and merchants offering 13,000 data products, with major exchanges in Guiyang, Shenzhen, Shanghai, and Beijing each exceeding RMB 1 billion in annual trading volume. These official marketplaces integrate third-party valuation, cleaning, and labeling services while implementing standardized "AI-Ready" grading frameworks for more than 500 petabytes of curated datasets. This commercial expansion operates under strict state surveillance, ensuring the Chinese Communist Party retains ultimate oversight and access.

Comment

Centralising dataset curation through institutional platforms like the Guiyang Big Data Exchange converts raw industrial telemetry into structured factor inputs at scale. By treating enterprise operational data as capital rather than proprietary trade secrets, state directives systematically lower the marginal cost of model training for dual-use artificial intelligence applications. This administrative aggregation solves the data-silo bottleneck that typically stalls private-sector algorithmic development.

The downstream effect strengthens China's defence-industrial base by streamlining the feed of physical-world sensor records into military synthetic environments. Enterprise-scale data annotation hubs, such as those operated by Baidu, act as digital force multipliers, providing the continuous, standardised data pipelines required to train autonomous systems faster than Western defence primes relying on fragmented commercial sources.

Strategic Question for Discussion
Which factor will exert a greater influence on the operational maturity of Chinese military AI — the sheer volume of sensor data aggregated through platforms like the Guiyang Big Data Exchange, or the throughput limits of human annotation hubs like those run by Baidu?
The trajectory indicates that throughput bottlenecks at annotation hubs like Baidu's will present a more severe immediate constraint than raw data collection at the Guiyang Big Data Exchange. While state exchanges can rapidly pool massive volumes of telemetry, transforming unstructured sensor streams into high-fidelity training data for military synthetic environments requires domain-specific expert labeling that automated tools cannot yet replace. Consequently, backend processing capacity, rather than raw data capture, will determine the operational deployment pace of sovereign dual-use models.
Share your assessment in the comments below.

No comments: