AB
AiBoss
project

k1.5 - Kimi's multimodal thinking model

k1.5 is the latest multimodal thinking model launched by Moonlit Dark Side Technology, possessing powerful reasoning and multimodal processing capabilities. Under the short-CoT (short-chain thinking) mode, the model exhibits significant advancements in mathematical, code, and visual multimodal capabilities, as well as general-purpose abilities...

What is K1.5?

k1.5 is the latest multimodal thinking model launched by Kimi from the Dark Side of the Moon, possessing powerful reasoning and multimodal processing capabilities. In short-CoT (short-chain thinking) mode, its mathematical, code, visual multimodal, and general capabilities significantly outperform the globally leading short-coordinated models GPT-4o and Claude 3.5 Sonnet, with a lead of up to 550%. In long-CoT (long-chain thinking) mode, k1.5's performance reaches the level of the OpenAI o1 official release, becoming the first multimodal model globally to achieve this level.

The design and training of k1.5 incorporates four key elements: long context expansion, improved policy optimization, a concise framework, and multimodal capabilities. By expanding the context window to 128k and employing partial expansion techniques, the model significantly improves inference depth and efficiency. k1.5 further optimizes performance by transferring the advantages of long-chain thinking to short-chain thinking models through the long2short technique.

Main functions of k1.5

  • Multimodal reasoning abilityk1.5 can process both text and visual data simultaneously, has joint reasoning capabilities, and is suitable for fields such as mathematics, code, and visual reasoning.
  • Short chain and long chain thinkingIn the short-chain thinking mode, k1.5 significantly outperforms leading global models (such as GPT-4 and Claude 3.5) in terms of mathematical, coding, visual multimodal, and general capabilities, with a lead of up to 550%. In the long-chain thinking mode, its performance reaches the level of the official version of OpenAI o1.
  • Excellent mathematical and coding skillsk1.5 performs exceptionally well in mathematical reasoning and programming tasks, especially in inputting mathematical formulas in LaTeX format.
  • Efficient training and optimizationThrough long context expansion (context window expanded to 128k) and improved policy optimization, k1.5 achieves more efficient training and exhibits inference characteristics of planning, reflection and correction.
  • Deep reasoning abilityk1.5 excels at solving complex reasoning tasks, such as difficult mathematical problems, programming debugging, and work challenges, and can help users unlock more complex tasks.

Technical principles of K1.5

  • Long Context ScalingKimi k1.5 extends the context window of reinforcement learning to 128k, significantly improving the model's inference ability by increasing the context length. The core of this approach is based on a partial rollout strategy, which reuses previous trajectory fragments to generate new trajectories, avoiding the high computational cost of generating complete trajectories from scratch.
  • Improved Policy OptimizationThe model employs a reinforcement learning formula based on Long-CoT (Long-CoT) and incorporates a variant of Online Mirror Descent for policy optimization. Effective sampling strategies, length penalties, and data formulation optimization further enhance the algorithm's performance.
  • A simple frameworkKimi k1.5's design abandons complex techniques such as Monte Carlo tree search, value functions, and process reward models. Instead, it achieves powerful reasoning capabilities by extending the context length and optimizing strategies. This enables the model to perform exceptionally well in long-context reasoning while possessing the ability to plan, reflect, and correct.
  • Multimodal joint trainingThe model was jointly trained on text and visual data, enabling it to process both text and visual information simultaneously and possess cross-modal reasoning capabilities.
  • Long2Short technologyKimi k1.5 proposed a method to transfer the reasoning ability of long-chain thinking models to short-chain thinking models, including model fusion, shortest rejection sampling, DPO (pairwise preference optimization), and Long2Short RL (reinforcement learning).

Project address for k1.5

How to use k1.5

  • Web versionYou can use it directly by visiting the Kimi official website.
  • MobileSearch for "Kimi Smart Assistant" in the app store and download it, or search for "Kimi Smart Assistant" via WeChat mini program.
  • API callsDevelopers can use the Kimi API to make calls.

Application scenarios of k1.5

  • Complex reasoning tasksKimi k1.5 performs exceptionally well in deep reasoning tasks, handling complex mathematical problems, programming debugging, and reasoning challenges.
  • Cross-modal reasoningThe model supports joint reasoning with text and visual data, and can handle tasks involving mathematical problems and graphical analysis, as well as comprehensive understanding of code and images.
  • AI intelligent assistantKimi k1.5 can act as an intelligent assistant, providing users with efficient reasoning capabilities to help solve a variety of complex problems. It can understand user needs through multi-turn dialogues and provide detailed answers.
  • EducationIn educational settings, Kimi k1.5 can be used to support instruction and help students solve mathematical problems, practice programming, and solve logical reasoning problems.
  • Scientific research and developmentFor researchers and developers, Kimi k1.5 can serve as a tool to assist in complex theoretical derivations, code generation, and algorithm optimization. Its support for LaTeX format mathematical formula input further enhances its applicability in the research field.
  • Multimodal data analysisKimi k1.5 can handle multimodal data and is suitable for analysis tasks that require combining text and image information, such as image annotation and visual question answering.