AB
AiBoss
project

Seedream 2.0 - A native Chinese-English bilingual image generation model launched by ByteDance.

Seedream 2.0 is a native Chinese-English bilingual image generation model developed by ByteDance's Doubao Big Model team, addressing the shortcomings of existing models in text rendering and cultural understanding. The model utilizes a self-developed bilingual big language model (LL...

What is Seedream 2.0?

Seedream 2.0 is a native Chinese-English bilingual image generation model developed by ByteDance's Doubao Big Model team, addressing the shortcomings of existing models in text rendering and cultural understanding. The model uses a self-developed bilingual large language model (LLM) as its text encoder, enabling it to directly learn local knowledge from massive amounts of data and generate high-fidelity images with accurate cultural details and aesthetic expression. Seedream 2.0 employs the Glyph-Aligned ByT5 model for flexible character-level text rendering and uses Scaled ROPE technology to generalize to untrained resolutions.

Main features of Seedream 2.0

  • Strong bilingual comprehension abilityIt supports high-precision understanding and adherence to Chinese and English commands, and can generate Chinese or English aesthetically pleasing images with cultural nuances, breaking down the barriers between different languages and visual styles.
  • Excellent text rendering capabilitiesSignificantly reduces text corruption rate, resulting in more natural and aesthetically pleasing font changes, and delivers high-quality outputs in the generation of traditional Chinese style patterns and elements.
  • Multi-resolution generation capabilityThrough a triple-upgraded DiT architecture, it achieves multi-resolution generation and improved training stability, enabling the generation of image sizes and resolutions that have never been trained before.
  • Human Feedback-Based Reinforcement Learning (RLHF) OptimizationBy developing a reward model and feedback learning algorithm, we improve the overall performance of the model in terms of image-text alignment, aesthetics, structural correctness, and text rendering.

The technical principles of Seedream 2.0

  • Data preprocessing
    • Data compositionThe pre-training data is carefully planned from four parts: high-quality data pairs, distribution-maintaining data, knowledge-injected data, and targeted supplementary data.
    • Data cleaning: A multi-stage filtering method is used to ensure data quality and relevance.
    • Active learning engineOptimize the image classifier to ensure high quality of the training dataset.
    • Image annotationGenerate general and professional titles, covering a variety of description types.
    • Text rendering data: Construct a large-scale visual text rendering dataset for text rendering tasks.
  • Model pre-training
    • Diffusion converter (DiT)It processes image and text tags using a scaled version of 2D rotational position embedding (Scaling RoPE), and supports generalization to untrained resolutions.
    • Text encoder: Self-developed bilingual large language model (LLM) learns local knowledge directly from massive amounts of data and supports high-fidelity image generation.
    • Character-level text encoder: Apply the Glyph-Aligned ByT5 model to achieve flexible character-level text rendering.
  • Post-training of the model
    • Continuous Training (CT)Extend training using high-quality datasets to improve the aesthetics of generated images.
    • Supervisory fine-tuning (SFT): Use a small number of high-quality images to fine-tune the model and enhance its artistic appeal.
    • Human Feedback Alignment (RLHF)By combining preference data, reward models, and feedback learning algorithms, performance can be improved in multiple aspects.
    • Prompt Engineering (PE): Rewrite user prompts using fine-tuned LLM to improve the quality of generated images.
    • refiner: Upscale the image generated by the base model to a higher resolution and fix structural errors.
  • Instructional Image Editing AlignmentSeedream 2.0 can adapt to imperative image editing models, such as SeedEdit, to achieve high-quality image editing while preserving high aesthetics and compositional fidelity.
  • PerformanceSeedream 2.0 excels in cues, aesthetics, text rendering, and structural correctness. After multiple rounds of RLHF optimization, its output is highly consistent with human preferences, achieving an excellent ELO score.

Seedream 2.0 project address

How to use Seedream 2.0

  • Access Platform Use: Use the official website of Doubao or the official website of Jimeng.
  • Register/LoginLog in to the Doubao platform using your account.
  • Input prompt wordsEnter detailed Chinese and English prompts on the image generation interface to describe the content of the image you want to generate.
  • Select generation modeChoose the appropriate generation mode (such as normal generation, high-definition generation, etc.).
  • Adjust parametersAdjust the generation parameters (such as resolution, style, etc.) as needed.
  • Generate imageClick the "Generate" button and wait for the model to generate the image.
  • Download or use imagesThe generated images can be downloaded directly or used for further editing.
  • Using API interface
    • Get API KeyIf you are a developer, you can obtain the API Key through the developer documentation of Doubao or Jimeng platform.
    • Send RequestUse an HTTP request to send the prompt and generation parameters to the Seedream 2.0 API.
    • Receive responseThe API will return links to the generated images, which you can download or use directly.

Application scenarios of Seedream 2.0

  • Poster designGenerates attractive posters, supports complex text rendering and artistic styles, and can generate high-quality poster designs based on user-input prompts.
  • Social media contentGenerate attractive images for social media platforms, supporting multiple styles and themes, helping users quickly create high-quality social media content.
  • Video contentIt generates cover images, keyframes, and other elements for video content, supports various video styles and scenes, and can generate relevant images based on video content.
  • Painting CreationIt generates paintings in various styles, supporting multiple art styles such as oil painting, watercolor painting, and sketching. It can generate high-quality paintings based on prompts entered by the user.
  • Teaching aidsIt generates teaching aid images, supports various teaching scenarios, and can generate relevant images based on teaching content.
  • Game scene generationIt generates game scenes and backgrounds, supports multiple game styles, and can generate relevant images based on game content.