AB
AiBoss
project

REEF - Shanghai AI Lab, in collaboration with the Chinese Academy of Sciences and other universities, launched fingerprint recognition technology for large-scale models.

REEF (Representation Encoding Fingerprints) is a fingerprinting technique used for large language models (LLMs). It generates a unique fingerprint for each model by embedding specific encoded information during model training...

What is REEF?

REEF (Representation Encoding Fingerprints) is a fingerprinting technique used for large language models (LLMs). By embedding specific encoded information during model training, a unique "fingerprint" is generated for each model. This "fingerprint" contains the model's basic features and its evolution at different stages. REEF technology features high accuracy, low overhead, robustness, and compatibility. It can achieve high-precision model recognition without degrading model performance, and the "fingerprint" can still be accurately identified even after multiple model modifications or merging.

The main functions of REEF

  • Model fingerprint recognitionThe REEF technology creates unique "fingerprints" for large language models (LLMs), enabling the identification and differentiation of different large models, even those that have been pruned or merged.
  • Copyright protectionREEF technology effectively prevents models from being "shelled" or disguised, protects model copyright, and prevents unauthorized use and tampering, providing strong support for model copyright protection.
  • High-precision recognitionREEF technology can achieve high-precision model recognition without reducing model performance. Even if the model has been modified or merged multiple times, its "fingerprint" can still be accurately identified.
  • low costThe implementation of REEF technology does not significantly increase the computational and storage costs of the model, and it can be widely used on models of various sizes.
  • compatibilityREEF technology can be seamlessly integrated with existing large-scale language models without requiring major adjustments to the model structure.
  • Combating illegal activitiesREEF technology provides a new approach to addressing copyright infringement issues related to large models, combating unauthorized copying, modification, or merging of models.

REEF's technical principles

  • Feature representation extractionThe REEF system first extracts key features from the internal structure of a large language model (LLM), which can reflect the unique properties of the model.
  • Encoding Vector GenerationThe extracted features are then encoded into a compact vector, or "fingerprint," which contains basic information about the model and reflects its performance characteristics on different tasks.
  • Hash function encodingThe REEF system uses a hash function-based encoding method to convert feature vectors into fixed-length binary strings to reduce storage space and improve recognition speed.
  • Noise robustness mechanismThe REEF system introduces a noise robustness mechanism, which can maintain the consistency of the "fingerprint" even if the model has been pruned or merged.
  • Core Alignment Similarity (CKA)The REEF system compares the CKA similarity of the feature representations of the suspect model and the victim model on the same samples. CKA is a similarity index based on the Hilbert-Schmidt Independence Criterion (HSIC) used to measure the independence between two sets of random variables.
  • No-training methodREEF is a training-free method, which means it does not harm the overall performance of the model, nor does it add additional training costs.
  • robustnessREEF is flexible to various subsequent model development techniques (including fine-tuning, pruning, merging, arrangement, and scaling transformations). Even if the model has undergone extensive fine-tuning or pruning, REEF can still effectively identify the victimized model.

REEF project address

REEF application scenarios

  • academic researchThe REEF system can help researchers quickly identify and verify the source of a model, ensuring the authenticity and reliability of research results.
  • Commercial copyright protectionThe REEF system can provide enterprises with strong copyright protection, preventing competitors from obtaining and using their research and development results through illegal means.
  • Government agencies and regulatory agenciesThe REEF system can be applied to government agencies and regulatory bodies to help them better manage and supervise the use of artificial intelligence technology, ensuring the healthy development of technology and social fairness and justice.
  • Intellectual Property ProtectionThe REEF system can help businesses and individuals effectively prevent their models from being stolen and protect their legitimate rights and interests.
  • Technical supervisionThe REEF system can assist government agencies and regulatory bodies in better managing and overseeing the use of artificial intelligence technologies.