CoF - DeepMind's Visual Modeling Mindset
CoF (Chain-of-Frames) is a new concept introduced by DeepMind, analogous to "Chain-of-Thought" (CoT) in language models.
CoF (Chain-of-Frames) is a new concept introduced by DeepMind, analogous to "Chain-of-Thought" (CoT) in language models.
Manzano is a new multimodal large language model (LLM) from Apple, capable of simultaneously achieving image understanding and image generation. The model transforms images using a hybrid vision tokenizer...
KAT-Dev-32B is an open-source intelligent large-scale model released by the Kuaishou Kwaipilot team, boasting 3.2 billion parameters. It achieved a 62.4% solution rate in the SWE-Bench Verified benchmark test, ranking 5th. The model has undergone...
KAT-Coder is a closed-source flagship code generation model released by the Kwaipilot team under Kuaishou, possessing powerful programming capabilities. It can efficiently complete tasks such as feature development, defect analysis, and unit test generation, and supports multiple programming languages...
JoySafety is JD.com's open-source large model security framework, providing enterprises with a mature, reliable, and free large model security protection solution. The model is based on various atomic capability modules (such as BERT, FastText, Transformer, etc.)...
Lynx is a high-fidelity personalized video generation model launched by ByteDance. It can generate videos that match the user's identity using only a single portrait photo. It is built on the Diffusion Transformer (DiT) basic model and incorporates an ID-adapter and...
DeepSeek-V3.2-Exp is an experimental artificial intelligence model launched by DeepSeek-AI. By introducing the DeepSeek Sparse Attention (DSA) mechanism, it significantly improves the efficiency of long text processing. The model is based on DeepSeek-V3...
OpenPPT is an open-source PowerPoint tool. Its core service, based on ChatPPT, provides an efficient and convenient PowerPoint creation experience. The tool supports multiple platforms, including Windows, macOS, and Linux, allowing users to create presentations on different devices...
Claude Sonnet 4.5 is Anthropic's latest and most powerful programming model. The model excels in multiple areas including programming, computer operation, reasoning, and mathematics, topping the SWE-bench Verified test and demonstrating its capabilities in various fields...
Ring-1T is a trillion-parameter thinking model open-sourced by Ant Group. Based on the Ling 2.0 MoE architecture, it is pre-trained on a 20TB corpus and trained for inference using the self-developed reinforcement learning system ASystem. It supports 128k iterations...
GLM-4.6 is a new generation of large-scale foundational model launched by Zhipu, with a total of 355 bytes of parameters and 32 bytes of activation parameters. The model demonstrates strong performance in real-world programming, long context processing, reasoning ability, information retrieval, writing capabilities, and intelligent agent applications...
Doubao Big Model 1.6-vision is a visual deep thinking model launched by Volcano Engine, featuring tool-invoking capabilities. The model boasts powerful general multimodal understanding and reasoning abilities, supports the Responses API, and can autonomously invoke tools such as...
RoboBrain-X0 is the world's first open-source embodied model from the Beijing Academy of Artificial Intelligence that supports zero-shot cross-ontology generalization. It can drive various real robots with different architectures to perform basic operations without fine-tuning...
EchoCare is a large-scale ultrasound model developed by the Centre for Artificial Intelligence and Robotics Innovation (CAIR) at the Hong Kong Innovation Research Institute of the Chinese Academy of Sciences. The model is trained on the EchoAtlas dataset, which contains 4.5 million ultrasound images...
Sora 2 is a next-generation AI audio and video generation model from OpenAI. On the web, it supports generating videos up to 25 seconds long (Sora Pro membership required). Technically, it achieves three core breakthroughs: for the first time, through multimodal joint training...
Logics-Parsing is an open-source end-to-end document parsing model from Alibaba, based on Qwen2.5-VL-7B. Through reinforcement learning, it optimizes document layout analysis and reading order inference, enabling the conversion of PDF images into structured HTML...
Tinker API, the first product released by Thinking Machines Lab, is designed specifically for fine-tuning language models. It simplifies the language model fine-tuning process, allowing researchers and developers to focus on algorithms and data without worrying about...
LONGLIVE is a real-time interactive long-form video generation framework jointly developed by NVIDIA and other leading institutions. The framework utilizes a frame-level autoregressive (AR) model, combined with a key-value (KV) caching mechanism, streaming long-form video fine-tuning, and short-window attention...
Dreamer 4 is a new type of intelligent agent developed by DeepMind that solves complex control tasks by being trained through imaginary observation of a world model in a fast and accurate manner. In the game Minecraft, Dreamer...
Mano is a proprietary large-scale model developed by MindLamp Technology, focusing on intelligent operation of graphical user interfaces (GUIs). Based on a multimodal foundational model, the model utilizes innovative technologies such as online reinforcement learning and automatic training data collection within the Mind2W...
SciToolAgent is an open-source tool platform developed by the Innovation Center of Zhejiang University (HICAI-ZJU) to improve research efficiency. It integrates over 500 scientific tools, covering fields such as biology, chemistry, and materials science, and can process data...
xLLM is a high-efficiency intelligent inference framework open-sourced by JD.com, optimized for domestically produced chips and supporting integrated edge-cloud deployment. The framework uses a service-engine separation architecture; the service layer handles request scheduling and fault tolerance, while the engine layer focuses on computational optimization, possessing...
Meta ARE (Agents Research Environments) is a dynamic simulation research platform launched by Meta for training and evaluating AI agents. The platform simulates complex, multi-step processes in the real world by creating environments that evolve over time...
FireRedChat is a full-duplex voice interaction system developed by the Xiaohongshu Intelligent Audio Team. It features real-time two-way dialogue capabilities and supports controlled interruption. It adopts a modular design, including a transcription control module, an interaction module, and a dialogue module...