Google releases Gemini 3.5 Transcribe
Google has released Gemini 3.5 Transcribe, a speech-to-text model that offers APIs for both real-time streaming and recorded files, and supports more than 85 languages.
Google has released Gemini 3.5 Transcribe, a speech-to-text model that offers APIs for both real-time streaming and recorded files, and supports more than 85 languages.
Zhipu has released GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. It adopts a MoE architecture with 320 total parameters and 18 activation parameters, and the model weights have been made public.
Google Labs has launched Play with Putty, a real-time collaboration tool emphasizing a "Vibe Coding" experience. Users need no programming background; they describe their needs using natural language, and AI automatically generates code and renders it in real-time as a usable tool or website. The tool supports collaboration among multiple users via links in the same shared space, with modifications synchronized instantly. Positioned as a lightweight, consumer-grade experimental tool, it transforms software development from individual writing to collective dialogue, enabling rapid prototyping through real-time collaboration.
Alibaba's Tongyi Qianwen open-sourced Qwen3.8-Flash-Next, which uses a 125B main model and 6B single-token activation parameters, and showcases the new architecture planned for Qwen4 in advance.
Jimeng AI has launched its film and television brand, "Jimeng Film Studio," soliciting film and television projects from film companies and AI creators from August to December. Films are offered in two tiers (S/A), and television series in three tiers (S/A/B), with support provided across the entire supply chain, including computing power, technology, and distribution, in collaboration with Tomato Novels IP. A creator development program is also launched simultaneously, including a creative fund and the "Echoes of Talent" award. Outstanding works can receive up to 100,000 RMB in cash and millions in promotional resources, supporting the development of AI-generated film and television IPs.
Stability AI announced the completion of a $76 million (approximately RMB 512 million) Series B funding round. The funds will be used for creative product development, applied research, and expansion of professional services. Investors include EA, Sony Music, Universal Music, Warner Music, and AMD Ventures. CEO Prem Akkaraju stated that the investor group will provide funding, expertise, industry credibility, and direct connections with artists, helping to empower creative professionals with generative AI.
AIsphere has released a technical report on the PixVerse R2 real-time multimodal world model. Building upon the PixVerse R1, the PixVerse R2 further explores the scaling path of real-time world models, supporting multimodal inputs such as text, reference images, audio, and motion through the Omni Causal AR unified architecture, enabling continuous world operation that involves simultaneous generation, interaction, and evolution.
OpenCode Go has added Grok 4.6. The official documentation estimates that approximately 169 requests can be completed every five hours based on a typical request pattern, but the actual number will vary depending on the input, cache, and output tokens.
The OpenAI product manager stated that a regression issue had caused approximately 3% of Pro and Thinking conversations to be unexpectedly routed to GPT-5.5-mini, but the issue has now been fixed.
Google Labs has announced a research experiment called Play with Putty, which allows multiple people to collaboratively build tools and websites through real-time Vibe coding. A shortlist of candidates is currently open.
Stability AI announced the completion of a $76 million Series B funding round, with investors including Electronic Arts, Sony Music, Universal Music, Warner Music, and AMD Ventures.
SenseTime's SenseNova U1.5 Lite has officially integrated with Token Plan, supporting image generation and editing, with a maximum output resolution of 4096×4096, and offering free access during the public beta period.
Tencent's WeChat Vision Team released WeMM-Embedding, offering three versions: 2B, 4B, and 9B, which uniformly support text, images, videos, visual documents, and interleaved multimodal input.
ByteDance officially launched "Doubao Work," an Agent product designed for productivity scenarios. It supports task breakdown, tool invocation, and complex workflows, and is integrated with Lark Enterprise Context.
The Qianwen APP and PC-based work assistant now feature a new Alibaba Cloud Token Plan binding function. After configuring an API Key, users can directly use their subscription quota to execute complex tasks.
MiniMax-AI has launched Awesome MiniMax H3 Integrations, which aggregates ecosystem resources such as checkpoints, quantization schemes, local execution, inference acceleration, and development tools for H3 models.
BreezeBlue released the Breeze TTS 2 speech model with 3 billion parameters, supporting Chinese and English speech cloning, voice design, tone control, and streaming synthesis.
Anthropic has established a $5 million grant program to support independent researchers in developing open-source assessments, benchmarks, and measurement methods to measure the impact of AI on user well-being.
Apple has released a new Mac Studio, available in M5 Max and M5 Ultra versions, with up to 512GB of unified memory, and supports multi-machine AI inference clusters via Thunderbolt 5 and RDMA.
IBM has released Granite 4.2, a language model available in 3B, 8B, and 30B parameter versions, for inference, tool invocation, coding, and enterprise agent workflows, and is licensed under the Apache 2.0 license.
Perplexity introduces Portable Computer, which allows Agent workflows to run locally on NVIDIA DGX Spark; local tasks do not consume credits and will request user authorization before sending content to the cloud.
Anthropic unifies Claude chat and Cowork's memory system, allowing users to view, edit, or delete memories by topic; sensitive topics are not saved by default, and availability for Team and Enterprise is controlled by the administrator.
LiblibAI has launched Newtake, an overseas AI video creation platform focusing on film-level content production. Based on an infinite canvas and node-based workflow, the product integrates features such as a director's console, cinematic lighting, multi-view analysis, texture character tables, and professional shot replication. It is already integrated with the Seedance 2.5 model. The product's core advantage lies in addressing the pain points of AI video face-swapping through character consistency management and lowering the barrier to entry for professional film and television production with a visualized node-based workflow.
Xiaohongshu's FireRed team has open-sourced FireRedTTS3, a unified speech generation and editing model. Based on RedAE semantically enhanced continuous representation and the LLM-DiT architecture, the model can achieve zero-sample timbre cloning, natural language voice design, and precise speech editing of specified regions for 24 languages and 21 Chinese dialects. The model has achieved state-of-the-art results on four public benchmark datasets, providing efficient solutions for audio content, game voice-over, and brand customer service scenarios.