PUMA - A unified multimodal large language model with multi-granularity strategies
PUMA is an advanced multimodal large-scale language model (MLLM) designed to unify and enhance visual generation and understanding tasks based on integrated multi-granularity visual features. PUMA can handle tasks ranging from text-to-image generation, detailed image editing, and more...
What is PUMA?
PUMA is an advanced multimodal large-scale language model (MLLM) designed to unify and enhance visual generation and understanding tasks based on integrated multi-granular visual features. PUMA can handle tasks ranging from text-to-image generation and detailed image editing to other visual tasks, adapting to varying levels of detail. Based on multimodal pre-training and fine-tuning techniques, PUMA demonstrates cutting-edge capabilities in diverse applications such as text-to-image generation, image editing, conditional image generation, and visual language understanding. The project was updated in October 2024 and is ongoing, a collaborative effort by researchers from CUHK MMLab, HKU MMLab, SenseTime, Shanghai AI Laboratory, and Tsinghua University. The PUMA project pushes the boundaries of AI visual language models, providing a flexible and powerful solution for future exploration of multimodal AI.
PUMA's main functions
- Diverse Text-to-Image GenerationPUMA can generate diverse and high-quality images based on text prompts, enhancing creativity and consistency based on coarse-grained visual features.
- Image editingPUMA enables precise image editing using fine-grained image features, including adding or removing objects, adjusting styles, etc., while maintaining the fidelity of the original image.
- Conditional image generationPUMA excels at image generation tasks based on specific input conditions, such as generating images from edge maps, image inpainting, or colorization, ensuring accurate and context-aware results.
- Multi-granularity visual decodingPUMA is based on five different granularities of image representation and corresponding decoders, enabling a wide range of visual decoding capabilities, from accurate image reconstruction to semantically guided generation.
PUMA's technical principles
- Multi-granularity image codingPUMA uses an image encoder to process input images and extract multi-level visual features from fine to coarse granularity, providing a foundation for generating diverse and controllable images.
- Self-regressive MLLMPUMA’s Autoregressive Multimodal Large Language Model (MLLM) can process and generate multi-scale text and visual tokens, making it suitable for the needs of different tasks.
- diffusion decoderPUMA uses a set of diffusion decoders corresponding to different feature granularities to perform visual decoding of images, supporting highly controllable or highly diverse visual outputs.
- Two-stage training strategyPUMA uses multimodal pre-training and task-specific fine-tuning to optimize the model's performance in multi-task processing, enabling the model to perform well in a variety of visual tasks.
PUMA's project address
- Project official website:rongyaofang.github.io/puma
- GitHub repository:https://github.com/rongyaofang/PUMA
- arXiv technical paper:https://arxiv.org/pdf/2410.13861
Application scenarios of PUMA
- Artistic Creation and DesignPUMA generates diverse images based on text descriptions, providing inspiration for artists and designers or enabling them to create artworks with specific styles and themes.
- Media and EntertainmentIn film, game, and animation production, it generates backgrounds, scenes, or concept art, accelerating the creative process.
- Advertising and MarketingPUMA can quickly generate engaging advertising images based on marketing copy, helping brands create visual content at a lower cost and faster speed.
- Education and TrainingPUMA can generate illustrations and example images for teaching materials, making educational content more vivid and interactive.
- e-commerceOnline retailers create visual representations of their products, for example, by generating product images based on descriptions or by changing product colors and styles.