AB
AiBoss
project

Granite-Docling-258M - A lightweight visual language model from IBM

Granite-Docling-258M is a lightweight visual language model from IBM, designed for efficient document conversion. The model can convert documents into machine-readable formats while fully preserving elements such as layout, tables, and formulas.

What is Granite-Docling-258M?

Granite-Docling-258M is a lightweight visual language model from IBM, designed for efficient document conversion. The model converts documents into machine-readable formats while fully preserving elements such as layout, tables, and formulas. With only 258M parameters, it boasts superior performance, high cost-effectiveness, and supports multiple languages (including Arabic, Chinese, and Japanese). The model uses the DocTags format to accurately describe document structure, preventing information loss. Granite-Docling-258M integrates seamlessly with the Docling library, providing powerful customization and error handling capabilities, making it suitable for enterprise-level document processing and a powerful tool in the document processing field.

Main functions of Granite-Docling-258M

  • Precise document parsingThe model can accurately identify and parse various elements in a document, such as text, tables, formulas, and charts, providing a clear and accurate data foundation for subsequent processing.
  • Structure Preservation ConversionWhen converting documents to electronic formats, the layout and structure of the original document are fully preserved, ensuring that the converted document is highly consistent with the original, making it easy to read and further edit.
  • Multimodal input supportIt supports both image and text input, and can process various document formats such as scanned documents, handwritten notes, and electronic documents, thus broadening its application scope.
  • Multilingual document processingIt has multilingual processing capabilities, enabling it to handle documents in different languages, thus facilitating document processing for multinational corporations and in multilingual environments.
  • High-efficiency data extractionIt supports quickly extracting key information and structured data from documents, improving work efficiency and reducing manual processing time.
  • Flexible output formatsIt supports converting documents into various common formats, such as Markdown, HTML, and JSON, making it convenient for users to perform subsequent processing and applications according to their needs.
  • Powerful customization capabilitiesIntegration with the Docling library allows users to customize document processing workflows according to specific needs, enabling personalized document conversion and analysis functions.
  • Enterprise-level stabilityAfter optimization, the model is more stable when processing documents, reducing the occurrence of errors and anomalies, making it suitable for large-scale application in enterprise environments.

Technical Principles of Granite-Docling-258M

  • Model Architecture:
    • Visual encoderUsing siglip2-base-patch16-512 as a visual encoder can efficiently process image input and extract visual features from documents.
    • Visual Language ConnectorBased on the pixel shuffle projector, visual features are connected with the language model to achieve the fusion of visual and linguistic information.
    • Language ModelBased on the Granite 165M language model, it can process and generate natural language text, ensuring accurate conversion of document content.
  • DocTags formatDocTags is a general-purpose markup language that accurately describes various elements in a document (such as charts, tables, formulas, etc.) as well as their context and location. DocTags also optimizes the readability of LLM, allowing the output document to be directly converted to Markdown, HTML, or JSON formats for easier subsequent processing and application.
  • Training dataTraining data includes public datasets and internal synthetic datasets, such as SynthCodeNet (code snippets), SynthFormulaNet (mathematical formulas), SynthChartNet (charts), and DoclingMatix (real document pages). With high-quality labeled data, the model can better learn the structure and content of documents, improving the accuracy and stability of the conversion.

Granite-Docling-258M project address

  • Project official websitehttps://www.ibm.com/new/announcements/granite-docling-end-to-end-document-conversion
  • HuggingFace model libraryhttps://huggingface.co/ibm-granite/granite-docling-258M
  • Experience the demo onlinehttps://huggingface.co/spaces/ibm-granite/granite-docling-258m-demo

Application scenarios of Granite-Docling-258M

  • Enterprise document managementThe model can quickly digitize paper documents, making them easier to store and retrieve, and improving work efficiency.
  • academic researchThe model can efficiently process large amounts of literature, helping researchers quickly acquire and analyze data.
  • Digitization of government recordsIt is used to accurately convert historical archives, ensuring information integrity and facilitating long-term preservation and retrieval.
  • EducationTeachers can quickly organize teaching materials, and students can easily access electronic learning materials.
  • Multilingual document processingMultinational corporations can handle multilingual documents, break down language barriers, and promote international exchange.