AB
AiBoss
project

CCI 3.0 - A large-scale Chinese internet corpus released by the Academy of Artificial Intelligence.

CCI 3.0 is a large-scale Chinese internet corpus released by the Beijing Academy of Artificial Intelligence (BAAI), containing a 1000GB dataset and a 498GB high-quality subset, CCI 3.0-HQ. This version expands the data scale by nearly one... compared to CCI 2.0.

What is CCI 3.0?

CCI 3.0 is a large-scale Chinese internet corpus released by the Beijing Academy of Artificial Intelligence (BAAI), containing a 1000GB dataset and a 498GB high-quality subset, CCI 3.0-HQ. This version nearly doubles the size of CCI 2.0 in terms of data volume, with the number of data sources increasing to over 20, improving the coverage and representativeness of the data. CCI 3.0 includes over 268 million web pages, covering multiple fields such as news, social media, and blogs. CCI 3.0 has performed meticulous classification and labeling of the raw data, covering more than 10 dimensions including grammar, syntax, and education level, filtering out high-value data.

Main functions of CCI 3.0

  • Data scale and sourcesCCI 3.0 boasts a data volume of 1000GB, encompassing over 268 million web pages and covering multiple sectors including news, social media, and blogs. The number of data source organizations has expanded to over 20, enhancing the data's coverage and representativeness.
  • Fine annotationCCI 3.0 meticulously categorizes and labels the raw data, covering more than 10 dimensions such as grammar, syntax, and education level, and filters out high-value data.
  • High-quality subsetsCCI 3.0 includes a 498GB high-quality subset, CCI 3.0-HQ, which is obtained by training a small-size quality model after automatically labeling samples based on the 70B model. This better meets the needs of different industries and application scenarios.
  • Data processing rulesDuring the construction process, CCI 3.0 uses methods including rule-based filtering (such as keyword filtering, spam filtering, etc.), model-based filtering (such as low-quality content filtering), and data deduplication (including deduplication within and between datasets) to ensure data quality and security.

Technical advantages of CCI 3.0

  • Significant training effectExperiments comparing training from scratch on different datasets using 100B data show that CCI 3.0 outperforms other datasets in both training on standalone Chinese corpora and training on mixed Chinese and English corpora, with CCI 3.0 HQ showing even more outstanding performance.
  • The concept of co-construction and sharingThe release of CCI 3.0 promotes data co-construction and sharing, builds a large-scale, high-quality, and high-knowledge-density Chinese dataset, and contributes to the development of China's artificial intelligence industry.
  • Convenient way to obtainThe CCI 3.0 dataset can be downloaded from platforms such as Flopsera, Huggingface, and Datahub, making it convenient for researchers and developers to use.

CCI 3.0 project address

  • Project official websitehttp://open.flopsera.com/flopsera-open/data-details/BAAI-CCI3

Application scenarios of CCI 3.0

  • Natural Language Processing (NLP) ResearchCCI 3.0 can be used for various NLP tasks, such as text classification, sentiment analysis, machine translation, question answering systems, and text summarization.
  • Large model trainingThe large-scale dataset of CCI 3.0 is suitable for training large language models, improving the performance and accuracy of the models in Chinese contexts.
  • Content recommendation systemBased on the corpus data in CCI 3.0, a more accurate user behavior prediction model can be trained for personalized content recommendation.
  • Knowledge Graph ConstructionBy analyzing a large amount of text in CCI 3.0, key information can be extracted to build a knowledge graph, which can be used to enhance search engines, improve the knowledge base of intelligent assistants, etc.
  • Education and academic researchCCI 3.0 can serve as a resource for academic research, helping scholars study the characteristics and changing trends of the Chinese language.