AB
AiBoss
project

TITAN - A multimodal whole-slice pathology model developed by Harvard Medical School.

TITAN is a multimodal whole-slice pathology model developed by a research team at Harvard Medical School. Through visual self-supervised learning and visual-language alignment pre-training, it can extract universal slice representations without fine-tuning or clinical labels...

What is TITAN?

TITAN is a multimodal whole-slice pathology model developed by a research team at Harvard Medical School. Through visual self-supervised learning and visual-language alignment pre-training, it can extract universal slice representations and generate pathology reports without fine-tuning or clinical labels. It used 335,645 whole-slice images (WSIs) and their corresponding pathology reports, combined with 423,122 synthetic captions generated by multimodal generative AI collaborators. TITAN performed exceptionally well on various clinical tasks, including linear probing, few-shot and zero-shot classification, rare cancer retrieval, cross-modal retrieval, and pathology report generation.

TITAN's main functions

  • Generate pathology reportTITAN can generate generalized pathology reports in resource-constrained clinical scenarios, such as rare disease retrieval and cancer prognosis.
  • Multitasking performanceTITAN demonstrates superior performance in a variety of clinical tasks, such as linear probing, few-sample and zero-sample classification, rare cancer retrieval and cross-modal retrieval, as well as pathology report generation.
  • Extracting general slice representationTITAN can extract universal slide representations applicable to a variety of pathological tasks, providing a powerful tool for pathological research and clinical diagnosis.
  • Search for similar slices and reportsTITAN excels in rare cancer retrieval and cross-modal retrieval tasks, effectively searching for similar slices and reports to aid clinical diagnostic decisions.
  • Reduce misdiagnosis and inter-observer variabilityTITAN has significant potential in clinical diagnostic workflows, assisting pathologists and oncologists in retrieving similar slides and reports, reducing misdiagnosis and observer variability.

TITAN's technical principles

  • Self-supervised learning and visual-language alignmentTITAN is pre-trained using visual self-supervised learning and visual-language alignment, enabling it to extract general-purpose slice representations without any fine-tuning or clinical labels.
  • Pre-training strategyTITAN's pre-training includes three distinct stages to ensure that the final generated slice-level representations can capture organizational morphological semantics at both the ROI and WSI levels, aided by visual and linguistic supervision signals.
    • Phase 1 (Visual pre-training only)Pre-training was performed on an internal dataset called Mass-340K, which contains 335,645 whole slice images (WSIs) and 182,862 medical reports.
    • Phase Two (Aligning the Region of Interest with the Composite Title)TITANV was pre-trained using 423,122 regions of interest (ROIs) and synthetic captions generated by PathChat to enable the model to capture morphological information at the region level.
    • Phase 3 (Aligning whole slide images with pathology reports)The final model TITAN was obtained by further pre-training 182,862 whole-slice images and their pathology reports, which enabled it to process high-level descriptions of slices.
  • Model DesignTITAN is based on the Visual Transformer (ViT) architecture. Its slice encoder uses pre-extracted image patch features, arranged in a two-dimensional feature grid to preserve spatial context. By increasing the image patch size, the length of the input sequence is effectively reduced. To handle the problem of irregular sizes and shapes in full-slice images, region cropping and data augmentation methods are employed.
  • Language abilityBy comparing the title generator (CoCa) with the pre-training in the second and third stages, the slice representations are aligned with the synthesized titles and pathology reports, respectively. The slice encoder, text encoder and multimodal decoder are fine-tuned to enable the model to have language capabilities, including generating pathology reports, zero-shot classification and cross-modal retrieval.

TITAN's project address

Application scenarios of TITAN

  • Pathological research and clinical practiceTITAN, through visual self-supervised learning and visual-language alignment pre-training, can extract general slide representations and generate pathology reports, providing a more effective tool for pathology research and clinical practice.
  • Clinical scenarios with limited resourcesTITAN is particularly suitable for resource-constrained clinical scenarios, such as rare disease retrieval and cancer prognosis, and can generate pathology reports with generalization capabilities.
  • Clinical diagnostic workflowTITAN can help pathologists and oncologists search for similar slides and reports, reducing misdiagnosis and observer variability.
  • Diverse clinical tasksTITAN excels in a variety of clinical tasks, including linear probing, few-sample and zero-sample classification, rare cancer retrieval and cross-modal retrieval, and pathology report generation.
  • Pathology report generationTITAN generates high-quality pathology reports without any tweaking or clinical labeling, even in resource-constrained situations.
  • Cross-modal retrievalTITAN performs exceptionally well in rare cancer retrieval and cross-modal retrieval tasks, effectively searching for similar slices and reports to aid in clinical diagnostic decision-making.