SOLAMI - A VR-based 3D role-playing AI system launched by Nanyang Technological University
SOLAMI is an innovative VR-based 3D role-playing AI system developed by a research team at Nanyang Technological University. It allows users to engage in immersive interactions with virtual characters using voice and body language, based on a social vision-language-behavior model...
What is SOLAMI?
SOLAMI is an innovative VR-based 3D role-playing AI system developed by a research team at Nanyang Technological University. It allows users to engage in immersive interactions with virtual characters using voice and body language. Based on a social visual-language-behavior model, it provides a natural communication experience that transcends traditional text and voice interactions. Driven by an end-to-end VLA model, SOLAMI can recognize and respond to user body language, supporting various character interactions such as dancing and playing games. SOLAMI brings a new level of immersive experience to AI role-playing games.
SOLAMI's main functions
- Immersive InteractionUsers can interact naturally with 3D virtual characters in a VR environment using voice and body language.
- Multimodal responseThe system can generate corresponding character voice and action responses based on the user's voice and action input.
- Role diversityIt supports a variety of characters, including superheroes, robots, and anime characters, providing a rich interactive experience.
- Interactive gamesSupports simple interactive games with characters, such as rock-paper-scissors.
SOLAMI's technical principles
- Social Visual-Language-Behavior Model (Social VLA)Using an end-to-end VLA model, it processes the user's voice and motion input to generate the character's response.
- Multimodal input processingBased on Motion Tokenizer and Speech Tokenizer, the user's voice and actions are converted into tokens that the model can understand.
- LLM baseUsing a large language model (LLM) as a base, it processes the input tokens and outputs the character's voice and action tokens in an autoregressive manner.
- Action representationUser actions are represented by 3D rotations in SMPL-X and encoded using VQ-VAE.
- speech processingThe user's voice is encoded using the RVQ-VAE structure and decoded using SoundStorm to achieve sound cloning.
- Training processThis includes multi-task pre-training and instruction fine-tuning training, enabling the model to learn the relationships between actions, speech, and text, and to handle multi-turn multimodal dialogues.
SOLAMI's project address
- Project official website:solami-ai.github.io
- arXiv technical paper:https://arxiv.org/pdf/2412.00174
SOLAMI Application Scenarios
- Virtual socialUsers can socially interact with AI characters in a virtual environment, simulating real conversations and non-verbal communication.
- Interactive gamesIn VR games, NPCs (non-player characters) can interact with players more naturally, enhancing the gaming experience.
- Education and trainingIt simulates the roles of teachers or students, providing educational scenarios such as language learning and social skills training.
- PsychotherapySimulate the role of a therapist in virtual reality to help users with psychotherapy and exposure therapy for social phobia.
- Entertainment and performanceUsers can interact with virtual singers, dancers, or actors and enjoy an immersive entertainment experience.