AB
AiBoss
project

Muse Spark - Meta's native multimodal large model

Muse Spark is the first native multimodal large-scale model launched by Meta AI Labs. As the flagship product of Meta AI's reorganization, the model jumped from 18 points to 52 points in the Artificial Analysis benchmark, demonstrating its multimodal...

What is Muse Spark?

Muse Spark is the first native multimodal large-scale model launched by Meta AI Labs. As the flagship product of Meta AI's reorganization, the model jumped from 18 points to 52 points in the Artificial Analysis benchmark, surpassing GPT-5.4 in multimodal understanding and health question answering capabilities. The model supports visual thought chains, multi-agent collaboration, and a "contemplative mode," with pre-training efficiency 10 times higher than Llama 4. The model is available on Meta and the Meta AI App, with an API preview version available to select users.

Main functions of Muse Spark

  • Native multimodal understandingIt supports visual thinking chains and image-to-code conversion, and can directly analyze complex charts, locate screen elements, and convert UI design diagrams into runnable HTML/CSS/JS applications.
  • Multi-agent collaborationBy scheduling multiple sub-agents in "Contemplating" mode to think and work collaboratively in parallel, complex tasks can be decomposed, planned, and executed.
  • Vertical specializationIn the healthcare field, it provides precise Q&A and image analysis based on data from 1,000+ clinicians, and in shopping scenarios, it combines social graphs to make personalized product recommendations.
  • Efficient reasoning mechanismIt employs automatic thought compression technology, which reduces token consumption to one-third of similar models while maintaining high performance, significantly improving inference efficiency.

How to use Muse Spark

  • Use directly on the webVisit the Meta website to experience basic features for free without registration.
  • Mobile AppDownload the official Meta AI App, which fully integrates the Muse Spark model.
  • API AccessDevelopers can apply for access to a private preview version of the API, which is currently only available to select partners.
  • Social platform integrationIn the coming weeks, it will be directly integrated with Facebook, Instagram, and WhatsApp, allowing users to access it directly from the chat interface.

Key information and usage requirements for Muse Spark

  • Product PositioningMeta Superintelligence Labs (MSL) launched its first model (codenamed "Avocado") nine months after its founding. The model is positioned as "personal super intelligence" and is aimed at an ecosystem of 3 billion users.
  • Core performanceArtificial Analysis overall score 52 (Llama 4 only 18 points); Multimodal graph understanding (86.4) and health Q&A (42.8) surpass GPT-5.4; Programming tasks (ARC AGI 2, SWE-Bench) still lag behind.
  • Technical HighlightsNative multimodal reasoning + visual thinking chain; parallel thinking in multi-agent "contemplating mode"; pre-training computing power requirement reduced to 1/10 of Llama 4, and token consumption is only 1/3 of Opus.
  • Team BackgroundLed by Alexandr Wang, the former founder of Scale AI, the core team includes several Chinese researchers (from OpenAI and DeepMind).
  • Access ChannelsMeta.ai web platform (no registration required), Meta AI App (iOS/Android); API preview is only available to partners.
  • Location and CostCurrently, it is fully open to the United States first; for individual users.Free, unlimiteduse.

The core advantages of Muse Spark

  • Native multimodal understandingIt performs exceptionally well on visual tasks such as chart understanding (CharXiv 86.4 points) and screenshot localization (ScreenSpot Pro 84.1 points), significantly outperforming GPT-5.4 and Gemini 3.1 Pro.
  • Medical and health specializationBased on a professional data system built in collaboration with over 1,000 clinicians, it has achieved industry-leading levels in open health Q&A (HealthBench Hard score 42.8) and medical image analysis.
  • Multi-agent cooperative reasoningThe unique "Contemplating" mode supports parallel thinking and task decomposition by multiple agents, and can schedule sub-agents to handle complex processes such as research, planning and execution.
  • Ultimate efficiency optimizationBy reconstructing the pre-training technology stack, the computing power requirement is reduced to one-tenth of that of Llama 4, and the automatic mind compression technology is used to make the token consumption only one-third of that of the top models in the same category.

Comparison of similar products to Muse Spark

Comparison Dimensions Muse Spark GPT-5.4 Gemini 3.1 Pro
Artificial Analysis Overall Score 52 Approximately 51 Approximately 57
CharXiv Chart Understanding 86.4 82.8 80.2
ScreenSpot Pro screenshot positioning 84.1 85.4 84.4
ARC AGI 2 Abstract Reasoning 42.5 76.1 76.5
LiveCodeBench Pro Programming 80.0 87.5 82.9
SWE-Bench Pro code fix 52.4 57.7 54.2
HealthBench Hard Health Q&A 42.8 40.1 20.6
MedXpertQA Multimodal Medicine 78.4 77.1 81.3
HLE (with tools) Deep Thinking 58.4 58.7 53.4
Pre-training computing power requirements 1/10 of Llama 4 Standard level Standard level
Token consumption efficiency 1/3 of Opus benchmark level benchmark level

Application scenarios of Muse Spark

  • Visual creation and developmentThe model supports directly converting application screenshots into runnable front-end code, can parse complex academic charts and engineering drawings, and can generate static images into interactive web games or troubleshooting tools.
  • Health and medical consultationBased on professional data from thousands of clinicians, it provides open-ended health Q&A and medical image interpretation, and can also generate interactive nutrition labels and personalized health management plans based on users' dietary restrictions.
  • Intelligent planning and collaborationIt can process complex tasks in parallel through multiple agents, such as coordinating cultural routes, parent-child activities and logistics for family travel planning, providing personalized shopping recommendations by combining social network data, and autonomously searching and integrating multi-source information to complete in-depth research.
  • Office and ProductivityIt supports office tasks such as document parsing, table analysis, and email composition, and also has screen automation capabilities based on screenshot understanding, enabling it to perform interface operations and form filling.