The relevant person in charge of the National Data Administration disclosed that as of August this year, more than 126,000 high-quality data sets have been built across the country, with the total volume exceeding 1,815PB (approximately 1.77EB), a significant increase of more than 89% from the end of the first quarter of this year.
The so-called high-quality data set refers to a data set that can be directly used for artificial intelligence model development and training after full-process data processing such as collection, cleaning, annotation, and processing, and can effectively improve model performance. As large model technology accelerates iteration, high-quality data has become the core "fuel" for the development of the AI industry. The 126,000 data sets released this time cover many key industries such as government affairs, medical care, finance, transportation, and manufacturing, marking the accelerated formation of my country's data element supply system.
It is worth noting that since the national data set management service system was officially launched in April this year, more than 1,700 data sets have been released, and the scale of data supply continues to increase. The operation of this system effectively breaks through the information barriers on both sides of data supply and demand, and provides infrastructure support for the efficient circulation and compliant use of data elements.
This Digital Expo also set up a special "Data Market" for the first time, where 95 data companies presented 253 high-quality data products, aiming to promote the precise connection between data suppliers and data users. Industry experts point out that the launch of data marts will significantly lower the threshold for AI companies to obtain high-quality training data, and is expected to accelerate the intelligent transformation process of thousands of industries.
From the first quarter to the end of the third quarter, the total number of high-quality data sets nationwide increased by nearly 90%, reflecting my country’s continued increase in investment in the construction of data resources. With the continuous release of policy dividends such as the "Twenty Data Measures" and the successive implementation of data exchanges in various places, my country's data element market is moving from "quantitative accumulation" to a new stage of "qualitative leap". It is foreseeable that the explosive growth of high-quality data sets will inject continuous impetus into the development of my country's artificial intelligence industry.
Comments