AB
AiBoss
project

SOLAMI - A VR-based 3D role-playing AI system launched by Nanyang Technological University

SOLAMI is an innovative VR-based 3D role-playing AI system developed by a research team at Nanyang Technological University. It allows users to engage in immersive interactions with virtual characters using voice and body language, based on a social vision-language-behavior model...

What is SOLAMI?

SOLAMI is an innovative VR-based 3D role-playing AI system developed by a research team at Nanyang Technological University. It allows users to engage in immersive interactions with virtual characters using voice and body language. Based on a social visual-language-behavior model, it provides a natural communication experience that transcends traditional text and voice interactions. Driven by an end-to-end VLA model, SOLAMI can recognize and respond to user body language, supporting various character interactions such as dancing and playing games. SOLAMI brings a new level of immersive experience to AI role-playing games.

SOLAMI's main functions

  • Immersive InteractionUsers can interact naturally with 3D virtual characters in a VR environment using voice and body language.
  • Multimodal responseThe system can generate corresponding character voice and action responses based on the user's voice and action input.
  • Role diversityIt supports a variety of characters, including superheroes, robots, and anime characters, providing a rich interactive experience.
  • Interactive gamesSupports simple interactive games with characters, such as rock-paper-scissors.

SOLAMI's technical principles

  • Social Visual-Language-Behavior Model (Social VLA)Using an end-to-end VLA model, it processes the user's voice and motion input to generate the character's response.
  • Multimodal input processingBased on Motion Tokenizer and Speech Tokenizer, the user's voice and actions are converted into tokens that the model can understand.
  • LLM baseUsing a large language model (LLM) as a base, it processes the input tokens and outputs the character's voice and action tokens in an autoregressive manner.
  • Action representationUser actions are represented by 3D rotations in SMPL-X and encoded using VQ-VAE.
  • speech processingThe user's voice is encoded using the RVQ-VAE structure and decoded using SoundStorm to achieve sound cloning.
  • Training processThis includes multi-task pre-training and instruction fine-tuning training, enabling the model to learn the relationships between actions, speech, and text, and to handle multi-turn multimodal dialogues.

SOLAMI's project address

SOLAMI Application Scenarios

  • Virtual socialUsers can socially interact with AI characters in a virtual environment, simulating real conversations and non-verbal communication.
  • Interactive gamesIn VR games, NPCs (non-player characters) can interact with players more naturally, enhancing the gaming experience.
  • Education and trainingIt simulates the roles of teachers or students, providing educational scenarios such as language learning and social skills training.
  • PsychotherapySimulate the role of a therapist in virtual reality to help users with psychotherapy and exposure therapy for social phobia.
  • Entertainment and performanceUsers can interact with virtual singers, dancers, or actors and enjoy an immersive entertainment experience.