AB
AiBoss
project

FIBO - An open-source image generation model, the first to natively support JSON.

FIBO is the first open-source text-to-image generation model that natively supports JSON, specifically trained on long, structured descriptions. The model was trained on over 100 million structured JSON descriptions (approximately 1,000 characters each), achieving accurate...

What is FIBO?

FIBO is the first open-source text-to-image generation model that natively supports JSON, specifically trained on long, structured descriptions. Trained on over 100 million structured JSON descriptions (approximately 1,000 characters each), the model allows for precise and repeatable control over lighting, composition, color, and camera parameters. FIBO supports three modes: Generate, Refine, and Inspiration, and features feature decoupling, allowing individual attribute adjustments without disrupting the overall scene. FIBO uses 100% licensed data, ensuring compliance and legal transparency, making it suitable for professional workflows.

FIBO's main functions

  • Text to Image GenerationGenerate high-quality images based on the text description entered by the user.
  • Structured JSON hintsExpand short text prompts into detailed, structured JSON descriptions, including details such as lighting, composition, and color.
  • Iterative controllable generationSupports generating images from short prompts, or refining existing JSON prompts in multiple rounds.
  • Feature decoupling controlAdjusting a single attribute (such as camera angle) without disrupting the overall scene.
  • Inspiration Mode: Extract structured cues from input images to generate relevant images and inspire creativity.
  • Enterprise-level compliance100% authorized data is used to ensure legal transparency and repeatability.
  • Production-level integrationThe model supports API interfaces, ComfyUI nodes, and local inference.

FIBO's technical principles

  • ArchitectureBased on an 8B parameter DiT architecture, it adopts a flow matching training method.
  • Text EncoderUsing SmolLM3-3B, coupled with the innovative DimFusion conditional architecture, we achieve efficient training of long descriptions.
  • VAEIt uses WAN 2.2 and is responsible for image encoding and decoding.
  • VLM bootstrapExtend short text hints into detailed structured JSON hints using a Visual Language Model (VLM).
  • Structured supervisionUsing structured JSON descriptions for training promotes feature decoupling and avoids cue word drift.
  • Data complianceTraining was performed on over 100 million authorized long structured JSON descriptions to ensure data compliance.

FIBO project address

  • GitHub repositoryhttps://github.com/Bria-AI/FIBO
  • HuggingFace model libraryhttps://huggingface.co/briaai/FIBO
  • Experience the demo online:https://huggingface.co/spaces/briaai/FIBO

FIBO application scenarios

  • Professional Design and Creative WorkflowGenerate high-quality images for advertising, product design, and graphic design, supporting rapid iteration and precise control to improve creative efficiency.
  • Film and EntertainmentFIBO can generate concept art and scene designs for movies, games, and animations, facilitating visual creation and accelerating the development process.
  • Education and TrainingThe model can generate teaching images and virtual experimental scenarios, assisting in the production of educational content and enhancing the learning experience.
  • Scientific researchModels can transform scientific data into intuitive images, aiding in research presentation and data visualization.
  • Medical and HealthFIBO can generate medical diagrams and virtual surgical scenarios, supporting medical teaching and surgical training.