AB
AiBoss
project

OpenAI o4-mini - A small inference model launched by OpenAI

OpenAI o4-mini is a small inference model from OpenAI, optimized for fast and cost-effective inference. OpenAI o4-mini excels in mathematical, programming, and visual tasks and is a competitor in AIME 2024 and 2025...

What is OpenAI o4-mini?

OpenAI o4-mini is a small inference model from OpenAI, optimized for fast and cost-effective inference. OpenAI o4-mini excels in mathematical, programming, and visual tasks, and was the top-performing model in the AIME 2024 and 2025 benchmarks. OpenAI o4-mini supports high-volume, high-throughput inference tasks, suitable for rapidly processing large numbers of questions. OpenAI o4-mini is multimodal, incorporating images into thought processes for reasoning, supports tool usage, and can quickly generate detailed and thoughtful answers. Compared to its predecessor, OpenAI o4-mini offers significant improvements in performance and cost-effectiveness. Currently, ChatGPT Plus, Pro, and Team users can see OpenAI o4-mini and OpenAI o4-mini-high in the model selector, replacing o1, o3-mini, and o3-mini-high. ChatGPT Enterprise and Edu users will gain access within a week. Developers can use the model via the Chat Completions API and Responses API.

Main functions of OpenAI o4-mini

  • Rapid reasoningIt excels at rapidly processing mathematical, programming, and visual tasks, making it suitable for high-throughput scenarios.
  • Multimodal capabilitiesIt combines images and text for reasoning and supports image processing.
  • Tool usageUse tools such as web search and Python programming to help solve problems.
  • High cost performanceWith superior performance compared to its predecessor, the o3-mini, and the same price, it is the top choice for upgrades.
  • Safe and reliableAfter undergoing security training, we support rejecting inappropriate requests.

Performance of OpenAI o4-mini

  • Mathematical reasoningIn the AIME 2024 and 2025 benchmark tests, OpenAI o4-mini achieved an accuracy of 93.4% without using any tools, and its accuracy soared to 98.7% after integrating Python, approaching a perfect score. In solving complex mathematical problems, OpenAI o4-mini outperformed its predecessor, o3-mini, and in some tasks, it approached the performance of the full-fledged o3.
  • Programming skills:
    • SWE-LancerThe OpenAI o4-mini performs exceptionally well, supporting the efficient completion of complex programming tasks and delivering outstanding benefits.
    • SWE-Bench Verified (Software Engineering Question Bank)The OpenAI o4-mini performs exceptionally well in common tasks such as algorithm, system design, and API calls, with higher accuracy and efficiency than the o3-mini.
    • Aider Polyglot Code Editing (Multilingual Code Editing Benchmark)The OpenAI o4-mini performs exceptionally well in code editing tasks, including both complete rewrites and patch modifications, outperforming the o3-mini.
  • Multimodal capabilities:
    • MMMU (University-Level Visual Mathematics Problem Bank)OpenAI o4-mini supports combining images and mathematical symbols to solve problems, achieving an accuracy rate of 87.5%, far exceeding the 71.8% of its predecessor, o1.
    • MathVista (Visual Mathematical Reasoning)The OpenAI o4-mini performs exceptionally well in visual mathematical reasoning tasks such as geometric figures and function curves, achieving an accuracy rate of up to 87.5%.
    • CharXiv-Reasoning (Scientific Chart Reasoning)OpenAI o4-mini can understand charts and diagrams in scientific papers with an accuracy of 75.4%, which is significantly better than o1's 55.1%.
  • Tool usage:
    • Scale MultiChallenge (Multi-round instruction follow-up)OpenAI o4-mini supports handling complex multi-turn instruction tasks and correctly understands and executes multi-turn instructions.
    • BrowseComp Agentic Browsing (Browser Task)Based on virtual browser search, click, page turning and information integration, its performance is close to O3 and far exceeds the traditional AI search capabilities.
    • Tau-bench function callsIt performs stably in function call tasks and supports accurate generation of structured API calls, but further optimization is needed in complex scenarios.
  • Comprehensive test:
    • Humanity’s Last Exam (Expert-level Comprehensive Test)The accuracy rate is 14.3% without tools, which improves to 17.7% with the help of plugins. It is lower than o3's 24.9%, but performs well in small models.
    • Interdisciplinary PhD-level science questions (GPQA Diamond)The accuracy rate on science questions was 81.4%, slightly lower than o3's 83.3%, which is already very good for a small model.

OpenAI o4-mini project address

Application scenarios of OpenAI o4-mini

  • Educational guidance: To help students solve math and programming problems.
  • Data AnalysisQuickly generate data charts and analysis results.
  • Software developmentGenerate code snippets to assist in code debugging.
  • Content creationProvides creative inspiration and generates descriptions by combining images.
  • Daily InquiryAnswer questions based on search and image analysis.