AB
AiBoss
project

Uni-1.1 - A new generation image generation model from Luma AI

Uni-1.1 is a new generation image generation model and API service launched by Luma AI. It adopts a decoder-only autoregressive Transformer architecture, which integrates text inference and pixel generation into a unified process.

What is Uni-1.1?

Uni-1.1 is a next-generation image generation model and API service launched by Luma AI. It adopts a decoder-only autoregressive Transformer architecture, integrating text inference and pixel generation into a unified process. The model supports joint input of up to nine reference images, sentence-level image editing, complex layout generation, and multilingual text rendering, ranking third globally in Arena.ai's blind test leaderboard. The API offers both pay-as-you-go and reserved throughput models, with a minimum cost of approximately $0.04 per image, targeting enterprise-level scenarios such as advertising, e-commerce, and content creation.

Main functions of Uni-1.1

  • Wensheng TuIt generates high-quality images based on text prompts and can output complex layouts containing more than a dozen types of layout elements, such as headers, navigation, advertisements, and body text, in a single output.
  • Image editingIt enables multi-round editing based on sentence-level instructions, retains unmentioned elements by default, and achieves iterative visual editing like editing a document.
  • Multi-reference graph fusionA single call supports up to 9 reference images as input, and uses brand logos, products, real people, and characters as model-level hard constraints for semantic-level fusion.
  • Space and attitude controlIt supports precise control over rotation, perspective switching, and spatial relationship adjustment, ensuring that the subject's identity and texture are not lost.
  • Multilingual renderingIt supports high-quality text generation using non-Latin characters such as Chinese and Arabic, meeting the needs of global content creation.

Technical Principles of Uni-1.1

  • Unified Autoregressive ArchitectureIt employs a decoder-only autoregressive Transformer, where text tokens and image tokens share the same sequence, enabling cross-modal joint reasoning.
  • Integrating Reasoning and GenerationThe model performs cross-modal inference before generating pixels. Constraints such as composition, space, and brand consistency are solved at the structural level, rather than translating first and then drawing.
  • Two-endpoint API designProvides Reasoning endpoints (deconstruction instructions, planning composition, locking brand/role/product constraints) and Generation endpoints (completing pixel rendering based on reasoning results).
  • Reference graph hard constraint mechanismMultiple reference images are passed in as hard constraints at the model level to ensure that the visual identity remains consistent across all channels and versions.

How to use Uni-1.1

  • Register an accountVisit the Luma AI Developer Platform website (https://platform.lumalabs.ai) to register and log in.
  • Get KeyCreate a project and obtain the API Key in the developer backend.
  • Select billing modeChoose either the Build plan (pay-as-you-go, suitable for flexible scheduling) or the Scale plan (reserved throughput, minimum order of 8 units, suitable for large-scale production) based on usage.
  • Call the Reasoning endpointSend text instructions and reference diagrams to allow the model to deconstruct requirements, plan the layout, and lock in brand/role constraints.
  • Call the Generation endpoint: Perform pixel rendering based on the inference results to obtain the final generated image.
  • Integration SDKIntegrate the API into your existing workflow using the official Python, JavaScript, TypeScript, Go, or CLI SDKs.
  • Upload reference image: Pass up to 9 reference images in the request as a hard constraint to ensure that the output is consistent with the brand's visual identity.
  • Iterative editingUse sentence-level editing commands to make multiple adjustments to the generated results, gradually optimizing them until a satisfactory effect is achieved.

Key information and usage requirements of Uni-1.1

  • Product NameLuma Uni-1.1 / Uni-1.1-Max
  • PublisherLuma AI (core research team of fewer than 15 people)
  • Release timeMay 6, 2026
  • Product PositioningEnterprise-grade AI image generation models and API services
  • Technical Architecturedecoder-only autoregressive Transformer (integrated reasoning and generation)
  • RankingsArena.ai ranks third globally (behind only OpenAI gpt-image-2 and Google nano-banana-2).
  • Price rangeBuild plan raw image: $0.0404–$0.1000 (2048px); Scale plan monthly fee: $2,100–$3,800/unit.
  • Enterprise clientsAdidas, Mazda, Publicis Groupe, Serviceplan, Envato, Comfy, Krea, etc.
  • SDK support: Python, JavaScript, TypeScript, Go, CLI
  • Core TeamJiaming Song (DDIM author), William Shen (CVPR Best Paper)

Uni-1.1's core advantages

  • Third best quality in the worldIt ranked third globally in the Arena.ai user blind ELO rating test, second only to OpenAI gpt-image-2 and Google nano-banana-2.
  • Ultimate cost-effectivenessThe lowest price for a single 2K resolution image is $0.0404, with both price and latency less than half that of top-tier models in the same category.
  • Enterprise-level consistencyBy using hard constraints from reference images and sentence-level editing, it solves the pain points of traditional model character deformation, brand color drift, and inconsistent styles across markets.
  • Completing complex tasks in one goIt can generate complete and readable news website pages and a full set of advertising campaign materials in one go, without the need for splicing multiple modules.

Comparison of Uni-1.1 with similar competing products

Comparison Dimensions Luma Uni-1.1 / Uni-1.1-Max OpenAI GPT-image-2 Google Nano Banana 2
Arena.ai rankings 3rd place (ELO 1193) No. 1 (ELO 1398) No. 2 (ELO 1268)
Publisher Luma AI (15-person Chinese team) OpenAI Google
Core Architecture Decoder-only autoregressive Transformer, integrating inference and generation. The specific architecture was not disclosed (it is speculated to be a diffusion model combined with multimodal operation). The specific architecture has not been disclosed (it is speculated to be a Gemini series multimodal).
Integrating Reasoning and Generation Text and image tokens share the same sequence; reasoning precedes generation. Traditional pipelines separate understanding from generation. Traditional pipelines separate understanding from generation.
Multi-reference graph fusion Up to 9 reference images are input at a time, enabling semantic-level fusion. Reference images are supported, but the fusion accuracy is limited. Reference diagrams are supported, but the constraint capabilities are limited.
Sentence-level editing Edit images according to sentences, retaining unmentioned elements by default. It supports editing but has weak consistency control. It supports editing, but is prone to crashing after multiple iterations.
Complex layout generation It can generate a complete news website/advertising page in one go, and the text is readable. Long texts and complex layouts are prone to errors. Complex layouts require the splicing of multiple modules.
Price of a single 2K resolution image Starting from $0.0404(Less than half the price of competitors) Relatively high (not disclosed, estimated at $0.08+) Relatively high (not disclosed, estimated at $0.08+)
Enterprise-level brand consistency The reference image serves as a model-level hard constraint, locking the visual identity across versions. Character/brand colors are prone to shifting, requiring repeated gacha pulls. Style consistency control is generally good.
Multilingual text rendering Supports non-Latin characters such as Chinese and Arabic. Excellent English, occasional flaws in Chinese. Good multilingual support
Latency performance Low latency (less than half that of competitors) medium medium
Main advantages High cost-performance ratio, consistent across enterprises, complex tasks completed in a single operation, and clear ROI. Generates top-quality, aesthetically pleasing, and ecologically mature products. Google ecosystem integration, stable output, and good multilingual support.
Main disadvantages Small team size, ecosystem still under construction High prices, weak corporate consistency, and poor editorial control. High price, complex layout and limited editing flexibility
Typical corporate clients Adidas, Mazda, Publicis Groupe, Serviceplan Large enterprises and creative agencies Google Cloud customers and advertisers
Applicable Scenarios Localized advertising, bulk e-commerce generation, IP consistency, brand pipeline High-end creativity, artistic exploration, prototype design Multilingual content, produced within the Google ecosystem

Application scenarios of Uni-1.1

  • Advertising localizationThe main visual can be quickly expanded into multilingual and multiregional versions, and brand elements can be locked in by reference images, which greatly shortens the production cycle.
  • E-commerce product visualizationIt generates consistent product images in real time based on product photos, fabric samples, and scene references, replacing the traditional shooting and template application process.
  • Character and IP Consistency: Provides consistency in character design, poses, and lighting across scenes for game promotional materials, comics, and film pre-production.
  • Brand Content PipelineIt integrates with enterprise content production systems to achieve batch generation and style consistency of visual materials across markets.
  • Creative PrototypingCombine hand-drawn sketches with material references to quickly generate realistic product concept images and 3D clothing renderings.