AB
AiBoss
project

HappyHorse - The top-ranked AI video generation model in blind testing by Artificial Analysis.

HappyHorse is a mysterious AI model that suddenly appeared at the top of the Artificial Analysis video generation blind test leaderboard, leading Seedance 2.0 by a wide margin with an Elo score of 1347, and winning the dual crown of text-generated and image-generated videos.

What is HappyHorse?

HappyHorse is a mysterious AI model that has suddenly topped the Artificial Analysis video generation blind test leaderboard, leading Seedance 2.0 by a significant margin with an Elo score of 1347, achieving dual crowns in both text-based and image-based video generation. The model is suspected to be a product of Alibaba's Taotian Future Life Lab (led by Zhang Di, former head of Keling), employing a 40-layer single-stream Transformer architecture and generating high-quality videos in just 8 denoising steps. Currently, the model remains anonymous and is considered a phenomenal dark horse in the AI video field in 2026.

Main functions of HappyHorse

  • Wensheng VideoHappyHorse supports generating high-quality, cinematic videos based on text prompts, and ranks first globally in the Wensheng video blind test leaderboard with an Elo score of 1347.
  • Image and videoThe model can generate dynamic video content based on reference images, setting a new record for the highest score in the image-generated video track with 1391 points.
  • Video production videoHappyHorse offers video-to-video editing capabilities, allowing users to style-transform or reconstruct content from existing videos.
  • Audio and video collaborationThe model has native audio generation capabilities and can output sound effects that match the visuals in sync, ranking second globally in the combined video and audio rankings.
  • High-definition outputSupports 1080p high-definition video export and provides watermark-free, commercially usable video generation services.
  • Reference Image WorkflowUsers can maintain consistency in roles or scenes through reference image workflows, enabling more precise control over video content.

How to use HappyHorse

  • Access Artificial AnalysisVisit the Artificial Analysis website https://artificialanalysis.ai/ and enter the Video Arena blind test area.
  • Participate in blind test votingThe system will randomly display two videos generated by anonymous models (which may include HappyHorse). Users can choose the better video based on factors such as image quality and motion smoothness by clicking "A is better" or "B is better".
  • View model identityAfter voting, the page will display which model each video comes from. If HappyHorse is selected, you can see its generation effect.

Artificial Analysis only provides comparative evaluation functions and does not support the generation of videos by inputting custom prompts.

Key information and usage requirements for HappyHorse

  • identityAn anonymous AI video generation model has topped both the Artificial Analysis and AI performance charts, and is suspected to be a product of Alibaba's Taotian "Future Life Lab" (led by Zhang Di, former head of Keling).
  • ArchitectureEmploying a 40-layer single-stream Transformer, 8-step denoising, and fusing diffusion models with an autoregressive Transfusion unified multimodal architecture.
  • performanceThe video score is 1347 (Artificial) and the video score is 1391 (Elo score), significantly ahead of Seedance 2.0 by nearly 60-74 points.
  • FunctionSynchronized generation of text-based videos, image-based videos, video-based videos, and native audio; specializing in character consistency, physical logic, and lip-sync.

HappyHorse's core advantages

  • Blind test leaderboard shows a clear lead in generation qualityIn the Artificial Analysis real user blind test, Wensheng Video topped the Elo score charts with 1347 points and Tusheng Video with 1391 points, leading the second-place Seedance 2.0 by nearly 60-74 points, equivalent to the total score difference between the second and nineteenth place.
  • Minimalist and brute-force single-stream architectureIt adopts a 40-layer single-stream Transformer to uniformly process text, video and audio tokens, abandons complex multi-stream structures, and can complete the generation in 8 steps of denoising, without the need for CFG guidance, which significantly reduces inference costs.
  • Native audio and video collaborative generationIt supports the synchronous generation of native sound effects that match the video, ranking second globally in the combined video and audio rankings, achieving a high degree of audio-visual alignment.
  • Commercial-grade personnel consistencyIt excels in simulating facial expressions, lip movements, body movements, and physical logic, making it particularly suitable for commercial scenarios with stringent requirements for character consistency, such as virtual humans and short dramas.

Comparison of HappyHorse and similar competing products

Comparison Dimensions HappyHorse-1.0 Seedance 2.0 (i.e. dream) Kling 3.0
Blind test ranking Wensheng/Tusheng VideoFirst in both rankings(1347/1391 points) Wensheng/Photos and VideosSecond on both lists(1273/1356 points) Ranked 4th-5th (1241 points)
Company Suspected Alibaba Taotian"Future Living Lab" (Zhang Di's team) ByteDance(Jimeng/CapCut Team) quick worker(Keling AI Team)
Core Architecture 40-layer single-flow Transformer8-step noise reductionNo CFG The specific architecture was not disclosed, but it is speculated to be a multi-stream DiT. Omni multimodal architecture, multi-step noise reduction
Audio generation Native audio synchronization, with audio chartssecond Native audio synchronization, with audio chartsFirst Audio is supported, but it's ranked low.
Usage cost $0.83-$1.24/100 credits (approximately $0.05/second) Domestic Premium Member 499 yuan/month(After the price increase) $13.44/minute (Pro version)
Available status Official website available, API "Coming soon" No API, queue required, limited free quota in China. The API is available, but it is expensive.

Application scenarios of HappyHorse

  • Virtual Humans and Digital Human CreationThe model has significant advantages in facial expression, lip-sync, and body movements, and is suitable for commercial scenarios that require a high degree of human consistency, such as virtual anchors, digital human short videos, and AI spokespeople.
  • AI Short Dramas and Film CreationIt supports generating cinematic-quality continuous footage, excels at multi-camera storytelling and character action coherence, and is suitable for producing AI short dramas, commercials, trailers and other film and television content.
  • Physical logic demonstration and product showcaseThe model can accurately simulate physical interactions (such as hula hoop rolling, rubber band ball bouncing, liquid being poured into coffee, etc.), and is suitable for product function demonstrations, educational and popular science videos, and creative content such as physics engines.
  • Audio and video synchronized contentThe model can create immersive videos with ambient sound effects and character dialogue, suitable for audio-visual collaborative scenarios such as audio stories, ASMR content, and dubbing clips.