AB
AiBoss
Tutorials

Step 3.7 Flash Open Source Model Real-World Testing - Multimodal Agent: Less Brainpower Token

Flash models are no longer just faster and cheaper alternatives to flagship models. They can now be integrated into Agent workflows, making each step faster, more stable, and more cost-effective. Recently, LeapStar launched a new generation of high-efficiency Flash...

Step 3.7 Flash开源模型实测 - 多模态 Agent 大脑更省Token

It's hard to imagine that businesses would use it. AI The cost has far exceeded the cost of hiring employees.

Last week, Axios reported that a AI The consultant revealed that one of his corporate clients was penalized for not providing its employees with... Claude The licenses had usage caps and cost a staggering $500 million in just one month.

miHoYo employees are testing AI Agent At that time, because dozens of them were built Agent If it wasn't shut down in time, approximately 2 million RMB worth of tokens were burned up overnight.

Multiple Agent The collaborative production chain, with its multiple rounds of calls and high-frequency tool triggering, is becoming an unbearable burden for enterprises due to token consumption and latency overhead.

That's why everyone's been pushing Flash models lately.

Flash models are no longer just faster and cheaper alternatives to flagship models. They can now be placed in… Agent In the workflow, make every step faster, more stable, and more economical.

Recently, Jieyue Xingchen launched its new generation.High efficiencyFlash rate open sourceModel Step 3.7 FlashAccording to the official documentation, Step 3.7 Flash is a 198B parameter-sparse MoE. MultimodalThe model activates approximately 11 bytes of parameters per token, supports 256K contexts, and has a maximum throughput of 400 tokens/s. It also supports three inference intensities: low, medium, and high.

We are more concerned about its performance in real-world, complex scenarios. Agent Link efficiency. Today, let's set aside ratings and leaderboards and conduct a real-world test.

The test used was Claude Code + StepFun’s Coding Plan.

Case 1 MultimodalPerception and UI Execution

I quickly sketched a draft to let Step 3.7 Flash create an e-commerce operations review dashboard.

Use the draft diagram as a reference to create an e-commerce operations review dashboard.

Step 3.7 Flash integrates visual understanding. Agent The workflow model can accurately recognize handwritten text and spatial layout in sketches.Transform sketches into modern, responsive HTML/CSS/JS Kanban applications.

The generated webpage has an extremely high degree of fidelity, almost exactly the same as my hand-drawn sketch. The page sections and text are recognized very accurately, and even the small arrows and icons I drew are reproduced.

However, there should be an "All" option at the top of the channel sales section, which Step 3.7 Flash omitted.

We continued to let it optimize the page based on the sketch:

We've continued to optimize the page, specifically the channel sales section, which differs from the original image. We've added an "All" option at the top to match the original layout.

Step 3.7 FlashMultimodalThe ability goes beyond simply understanding images; it includes being able to pinpoint the areas that need modification and make accurate changes.

Case 2 Visual Search and Tool-Enhanced Reasoning

BYD released its production and sales report for May today. Let's test it with Step 3.7 Flash:

Extract key information from images and generate analysis reports online.

This task is not simply OCR character recognition, but rather to see if Step 3.7 Flash can first extract key data, then verify the background online, and finally output a readable analysis report.

Step 3.7 The information recognized by Flash is very accurate.

Let's take a look at the generated report. Step 3.7 Flash captured several key points, and the content is very accurate:

BYD's sales of new energy vehicles in May 2026 were 383,453 units, and its production of new energy vehicles was 380,549 units.

The cumulative year-on-year decline from January to May was 20.32%, but production increased by 8.78% and sales increased by 0.26% in May, showing a clear recovery and marking an important turning point, with both production and sales showing restorative growth.

In May, exports accounted for 41.9% of total new energy vehicle sales, making exports one of BYD's most important growth engines.

Case 3 Visual Understanding

I uploaded a picture of a mixing console and asked it:

How do I adjust the microphone?

Step 3.7 Flash identified this as an NFM M-series professional mixing console and also explained that adjusting the microphone requires checking the channel and G.AIKey locations include N, FADER, MUTE, AUX, and main output.

For the average beginner, the steps provided in Step 3.7 Flash can basically guide people to troubleshoot problems such as "why the microphone has no sound", "the sound is too low", and "there is feedback".

In particular, the reminder to first look at the MUTE, then the gain, then the channel fader, and then check the main output is very impressive in terms of visual understanding and logic.

Case 4: Image to Interactive Map

Please use the images in the folder as input directly, without providing any additional background information. Please complete the entire workflow in one go.

Objective: Create a complete, demonstratorable single-page HTML city tour page, named ucsd-tour.html. The page should be able to:

1. Identify landmarks in the provided image.

2. Verify the recognition results through web page search.

3. Copy the image to the current working directory and save it with a suitable name.

4. Build a beautiful, interactive map-style city guide.

Important input rules:

  • Only use the directly provided images as input.
  • Do not scan folders or directories to find additional images.
  • Do not import irrelevant images from the current directory.
  • The provided images will be considered as a complete image set.

Step 3.7 Flash was able to accurately identify 7 locations, indicating that its visual understanding and web search capabilities are up to standard.

However, upon closer inspection, the landmark names and images do not match, suggesting that the model may not be rigorous enough in terms of multi-file management, path mapping, and resource naming.

Looking at the map generated by Flash in Step 3.7, it only roughly draws a direction and does not actually represent a map. The locations of the landmarks also deviate from their actual geographical locations.

Overall, Step 3.7 Flash only completed the core recognition task, and there is still room for improvement in the details.

Step 3.7 The most intuitive feeling Flash gives me in actual interaction is its fast response speed.

While there is still room for improvement in some details when dealing with complex tasks such as multi-file mapping and precise spatial logic, Step 3.7 Flash's high response speed and...MultimodalThe integration of perception was demonstrated in multiple rounds of interaction.High efficiencyError correction abilityThus, with lower latency and cost, it can provide complex... Agent The link has been traded for greater fault tolerance.

The actual tokens used in this evaluation accounted for only about 15% of the weekly Coding Plan credit limit.Thanks to the cost advantages of the MoE architecture, even Agent Even with high-frequency multi-round iterations, retrieval, and error correction in long workflows, the computing cost can still be kept within a range that is completely affordable for enterprises.

With Step 3.7 Flash, a production-ready technology... Agent ofHigh efficiencyRate Flash model,Agent When dealing with real-world tasks, it can run the entire workflow faster, more stably, and more cost-effectively, instead of being a daunting token-devouring beast.

Large ModelApplications are becoming more pragmatic. When businesses no longer need to worry about exorbitant bills and delays,AI Only then can it truly transform from a toy displayed in a single location into a productivity tool that operates stably on an industrial-grade production line.

Original link:Real-world step 3.7 Flash performance: More stable, faster, and more energy-efficient. Agent brain