Gemini Omni Flash is a video generation model introduced at Google I/O, positioned as a unified multimodal generation model that can generate arbitrary outputs from arbitrary inputs.
Gemini Omni Flash is a video generation model introduced at Google I/O, positioned as a unified multimodal generation model that can generate arbitrary outputs from arbitrary inputs.
Qwen3.7-Max is a new generation flagship model launched by Alibaba's Tongyi Qianwen team, designed for the era of intelligent agents and positioned as an all-around intelligent agent foundation. The model features cutting-edge programming, office automation, long-cycle autonomous execution, and cross-framework compatibility...
Qwen3.5-LiveTranslate is a large-scale real-time simultaneous interpretation model developed by the Alibaba Tongyi team. It supports input in 60 languages, output in 29 languages, and over 3500 translation combinations. It uses readable unit streaming technology to compress end-to-end word latency to...
Composer 2.5 is Cursor's self-developed Agentic programming model. It represents a significant improvement over Composer 2 in terms of intelligence and behavioral performance, particularly in SWE-Bench Multilingual (79.8%) and CursorBench benchmarks...
Chronicles-OCR is the industry's first cross-temporal technology jointly launched by Tencent Hunyuan, the Institute of Information Engineering of the Chinese Academy of Sciences, Anyang Normal University, Nankai University, and the Palace Museum, covering the complete evolutionary trajectory of the "seven styles" of Chinese characters...
ESP-Claw is an AI Agent framework for IoT devices launched by Espressif Systems. It adopts the 'Chat Coding' concept, allowing users to define and modify the behavior of hardware devices through natural language dialogue.
Qwen3.7 Preview is a preview version of the next-generation flagship model released by the Ali Tongyi Qianwen team, which includes two versions: Qwen3.7-Max-Preview and Qwen3.7-Plus-Preview.
MemPrivacy is an open-source edge-cloud collaborative agent privacy protection framework jointly developed by the MemTensor team, the Honor AI team, and Tongji University. It addresses the privacy leakage risks in long-term memory scenarios for cloud-based agents...
PPT Master is an open-source, AI-driven standardized workflow (skill) for generating PPT presentations. It runs in an AI IDE with agent-like capabilities and can handle any document format, including PDF, DOCX, XLSX, URLs, Markdown, and PPTX.
Higgs Avatar v1 is a real-time AI digital human model for voice-activated agents launched by BosonAI. The model requires only a single still photo to generate a real-time interactive digital human with synchronized lip movements, facial expressions, and head gestures.
Violin is an open-source, end-to-end AI video translation tool developed by Kevin Lin, a postdoctoral researcher at Oxford University, breaking down language barriers in high-quality video content. It integrates Whisper speech recognition, large language model translation, and TTS speech synthesis...
Intern-S2-Preview is an open-source preview of the next-generation, multimodal scientific model from the Shanghai Artificial Intelligence Laboratory. With only 35 bytes of parameters, it achieves scientific capabilities comparable to models with trillions of parameters. The model integrates general and specialized knowledge across the entire process...
OpenHuman is an open-source personal AI super-intelligent assistant developed by the tinyhumansai team. Positioned as 'Your Personal AI super intelligence,' it emphasizes privacy, simplicity, and extreme power. It's not a traditional chatbot...
Pixal3D is a single-image 3D generation project launched by Tencent ARC Labs in collaboration with Tsinghua University and Victoria University of Wellington. Pixal3D explicitly upscales pixel features into three-dimensional space through back projection, establishing direct pixel...
HiCAD is an open-source AI parametric 3D CAD modeling platform designed for 3D printing enthusiasts. Users describe their needs in natural language, and the AI can generate editable JSCAD parametric code in seconds, along with real-time 3D preview...
Kimi WebBridge is a browser extension developed by Dark Side of the Moon, designed for local AI agents such as Kimi Code, Claude Code, Cursor, and Codex.
TencentDB Agent Memory is an open-source AI agent hierarchical memory management tool developed by the Tencent Cloud database team, licensed under the MIT license. The tool utilizes a unique L0-L3 four-layer progressive memory architecture combined with context offloading and Mermaid tasks...
General365 is a general reasoning benchmark open-sourced by the Meituan LongCat team. It includes 365 original seed questions and 1095 extended variations, covering eight dimensions of reasoning challenges.
AGenUI is the industry's first open-source, cloud-integrated A2UI framework launched by Gaode Maps in collaboration with Alibaba's Qianwen C-end application team, covering iOS, Android, and HarmonyOS.
Xiaomi OneVL is an open-source autonomous driving model launched by Xiaomi's Embodied Intelligence team. It is the first in the industry to unify the three major technical routes of VLA vision-language-action, world model and latent space reasoning into a single framework.
OpenMontage is the world's first open-source agentic video production system, which uses an AI programming assistant to autonomously choreograph the entire process from concept to finished product.
In the AI era, the most scarce resource is undoubtedly tokens, especially for those who are into lobsters (a metaphor for those who are into cryptocurrency trading). They burn through tokens like water, and it's impossible to stop.
9Router is an open-source AI programming routing proxy tool that can unify the access of mainstream AI programming tools such as Claude Code, Codex, Cursor, and Cline to the local proxy layer, and intelligently schedule 40+ vendors and 100+ models.
ELF (Embedded Language Flows) is the first diffusion language model launched by Kaiming He's team. It adopts a continuous diffusion paradigm instead of the traditional autoregressive approach. The model generates text by denoising within a continuous embedding space throughout the entire process...