Seven real-world tests were conducted on Opus 4.8, Gemini 3.5 Flash, GPT-5.5, and Qwen3.7-Max.
There have been a lot of large model updates lately, but it's hard to know which one is actually good. Some say the Qwen 3.7-Max has surpassed the GPT-5.5, second only to the Claude series. Others say the GPT-5.5 has reached the top. For the average person...
recentLarge ModelThe updates are really frequent; there have been so many, but I still don't know which ones are actually good.
Some say the Qwen 3.7-Max has already surpassed... GPT-5.5, second only toClaude series.
Some say GPT-5.5 has reached the top.
The more ordinary people look at the rankings, the more confused they become.Which tool should I use for writing articles? Which for data analysis? And which for writing code, reviewing pull requests, and breaking down tasks?
I've selected four models that have been generating a lot of discussion recently:Claude Opus 4.8Gemini 3.5 Flash,GPTLet's do a comparative review of Qwen3.7-Max and Qwen5.5 to see how they perform in real-world tasks.
In this evaluation, we used the same materials and the same criteria.Prompt wordsThe same set of scoring criteriaThe tasks are assigned to four models, covering seven common task types: long document processing, task planning, code repair, Chinese writing, data analysis, format compliance, and SVG generation.
Case 1: In-depth reading of long documents
This case tests:Does the model truly understand the material, or does it simply start generating output based on a few keywords?.
Suitable for test reports, meeting minutes, investment research materials, and product documents.
Task:
1. Summarize the core conclusions of the material in no more than 200 words.
2. Extract the 5 most important facts, and indicate the source text for each one.
3. Identify 3 uncertainties or data gaps.
4. Determine whether the author's conclusion is fully supported by the material, and indicate "support/partial support/no support".
5. Output table: Conclusion, evidence, risks, and suggestions for further questions.
Require:
– For content not mentioned in the material, write "not mentioned in the material".
– Don't make things up.
– Do not output your thought process.
The materials are as follows:
In the year and two months since its launch, the SU7 has accumulated over 258,000 deliveries. Last month alone, we delivered 28,000 units, making it the top-selling model among all vehicles priced above 200,000 yuan. Next, I'll be introducing the Xiaomi YU7, Xiaomi's first SUV. The first question many people have asked is: how is the YU7 named? The YU7 is named "Riding the Wind." These four characters come from Zhuangzi.Free and easy travelThe name YU7 evokes the image of flying on the wind, a very auspicious meaning. Positioned as a luxury high-performance SUV, the YU7 is not just an ordinary SUV; it's a meticulously designed luxury high-performance SUV with elegant styling, driving pleasure, spacious comfort, and a luxurious experience. Let's take a look at the styling. It shares a family design language with the Xiaomi SU7, but it's definitely not the SU7's.SimpleThe raised version, a redesigned version of the SU7, boasts an elegant design style with smooth and powerful lines. Let's take a look at its unique luxury car aura and sports car-like driving experience, achieving a harmonious connection between driver and car. Frankly speaking, such a good-looking SUV at this price point is very rare. Let's first watch the unveiling video. Our YU7 is 5 meters long, with a 3-meter wheelbase and 2 meters wide. These dimensions classify it as a mid-to-large SUV. Although it looks very compact on the outside, the actual interior space is quite spacious due to its mid-to-large SUV size. Let's take a closer look: its side profile is low-slung and elegant, exuding sportiness, while the rear is muscular and powerful. From above, its surfaces are very three-dimensional and particularly full. Isn't it beautiful? I'd like to share with you how we created such a great-driving car. Firstly, beauty comes from proportion. A good proportion is essential for a good-looking car, just like a good physique makes a person attractive. Let's look at its essence: a 3:1 wheel-to-axle ratio, a 2.1:1 wheel-to-height ratio, and a 1.25:1 width-to-height ratio. It also boasts a 1.3 wheel-to-body ratio and a long, sleek hood, a luxurious design honed over a century of automotive engineering. From the rear, it exudes a muscular presence, featuring a wide body kit and optional 275mm rear tires, giving it an exceptionally powerful look. Secondly, its beauty lies in the details. Let's look at the teardrop headlights. Functionally, they support a 180-degree ultra-wide-angle illumination, providing excellent nighttime visibility. If you look closely, you'll notice the upper part is hollowed out, incorporating air ducts that connect seamlessly with the hood – a design typically found only in million-dollar sports cars. The halo taillights have also been upgraded, better suited to the SUV's styling, more concise, three-dimensional, and powerful. They are particularly eye-catching in nighttime traffic. When we first launched it, some people weren't used to it, but once you get used to it, you'll find the taillights exceptionally beautiful. And the door handles... after our launch, many people asked us, "The SU7's semi-concealed door handles are quite good, how do you use these? Is this a step backward?" Actually, no. This is a variable inward-opening door handle; it opens as soon as you approach the door handle.automaticIt flips inward, and when you leave the car or get back in, it flips back again. This way, it looks good and has low wind resistance. Let's take a look; as soon as someone gets close, it...automaticIt bounces off, and then after you get into the car, it...automaticWhen closed, it features an electrically operated inward-folding design. In terms of aerodynamics, it boasts 10 interconnected air ducts and 19 air vents that manage airflow throughout the vehicle. Let's look at some details, such as the active grille and 100-speed automatic shutters.intelligentThe drag coefficient was reduced by 18 counts, equivalent to an increase of 14 kilometers in range. We redesigned the rear spoiler 100 times, reducing it by 10 counts. The super-large, super-cool clamshell hood, with its integrated design, reduced the drag coefficient by two counts. So many details were changed, with over 40 aerodynamic optimizations. For a sporty SUV like this, we achieved a drag coefficient of 0.245, which is outstanding among sporty SUVs, equivalent to an increase of 59 kilometers in range. Color and beauty come from color. The SUV's shape is fuller and more three-dimensional. We searched for vibrant colors in nature, so that under the refraction of light, its changes are richer and more beautiful. Today, the first color we're introducing is emerald green, a highly saturated color. Let me first explain that this highly saturated color was previously only used in sports cars. Why? Why do you see mostly black, white, and gray on the street? Because these three colors are both beautiful and inexpensive, mainly inexpensive. What makes highly saturated colors so expensive? Think about it, these cars need to withstand outdoor elements like wind, rain, and sun for 10 or 20 years without changing color. Consider how complex it is to create such a highly saturated color. Therefore, the development cycle for each new color is often 13 months, an exceptionally long time. This color isn't something we can simply create by tweaking a color palette on a computer; it involves a very complex R&D process. Only those who specialize in supercars are willing to invest so much time and resources in creating such beautiful colors. Remember the Gulf Blue of the SU7? It was also very attractive, with high saturation. That's why we were particularly keen to show everyone...recommendSome particularly beautiful colors. And what I want to tell you about is this emerald green, inspired by Colombian emeralds. Its color is saturated and vibrant, possessing an unparalleled emerald hue, a crystal-clear texture, and shimmering brilliantly in the light. To better recreate the texture of emeralds, we used a double-layer paint process. How was this double-layer design and paint process done? We used three sets of paint cans. For the outer surface, we first sprayed a layer of yellow-green metallic paint, then a layer of transparent pearlescent paint. For the inner surface, we sprayed a mixture of metallic and pearlescent paint. It takes three cans to spray just one color. Other colors...SimpleAll of these only require one container, so the cost is more than twice that of regular colors, making them more difficult to produce than he imagined. This color looks especially beautiful in sunlight; let's take a look. When viewed from different angles, the emerald green looks like a gemstone, with light flowing across the slanted surface and reflected beautifully on the paint. We also took some photos in the bamboo forest, which looked great too. Today, we'll also introduce the second color...Titanium metallic colorIt exudes understated luxury. We added coarse-grained aluminum powder to the paint, giving it a unique metallic strength. It looks exceptionally beautiful. I also created an extremely striking color.Lava OrangeI especially love this color, but I think the SUV's shape is more three-dimensional, textured, and impactful. So it's particularly suitable for passionate young people, which I also really like. It's also suitable for young people like me. Let's take a look; it's very beautiful. Including the Cambrian Grey from last time, we've already introduced four colors for the YU7. There are five more to come, which we'll introduce at the next launch event. They're all quite beautiful. Now, let's take a look at the interior of the Sky Screen. Now, everyone, please look quietly for a while. Doesn't it feel a bit like an aircraft cabin? Our entire interactive system has added the panoramic display of the Xiaomi Sky Screen, and the rear control screen brings a completely new visual experience. Let's take a closer look; this is our Sky Screen. The interaction is interesting. Look closely, everyone; it's not...SimpleThe screen is a high-end projector integrating advanced technology. It has three mini screens that, through panoramic curved projection technology, project onto the black area under the windshield, forming a 1.1-meter ultra-wide display. Moreover, the display precision is super retina-level high-definition, making it exceptionally beautiful. It uses panoramic curved projection technology; it is not...SimpleThey installed a screen there, and our interactive system is exceptionally user-friendly and intuitive. For example, when driving, information like speed and navigation is readily available. The passenger side can also display music cards, and the blind spot image is directly displayed when turning, which is very convenient. And then there's the driver assistance system...automaticWhen switching to SR mode, road condition information can be displayed, and information such as power suspension can also be shown when changing driving modes. We provide 5 types of information cards, and you can also...automaticThe combination of these features suggests advanced technology. Our interior not only boasts a technological feel but also a luxurious one. We've adopted a dual-zone wraparound design with a clean and full shape. The dashboard is thin, offering exceptional visibility, and the materials are exceptionally luxurious. All touchpoints are 100% soft-touch, and our materials are even safe for babies to touch, having received international Class A certification. Therefore, the overall material feel is exceptionally good. The space is also exceptionally large. Using a 1.88-meter dummy in the front row, there was still 100 millimeters of headroom, significantly more than the Model Y and Porsche Cayenne. Regarding the seat style, this time we've tuned them to a luxurious comfort setting. The SU7 is tuned to a firmer, sportier, and more road-feeling setting, while the YU7 leans towards luxury. We've also equipped them with zero-gravity seats, scientifically distributing body pressure by adjusting the leg and backrest angles to over 120 degrees to enhance comfort. We've also included a 10-point massage function, perfect for relaxing while parked or taking a nap in the car at noon. Our driver's seat features a 12-layer zero-gravity structure, making it thicker, softer, and more supportive. The zero-pressure foam enhances comfort for short journeys, while the high-density memory foam provides excellent support, preventing fatigue even after long periods of sitting. All seats are upholstered in Nappa leather, offering an exceptionally delicate feel. Speaking of zero-gravity seats, it's a trend that has emerged in recent years with the rise of new energy vehicles in China. Even luxury cars haven't yet seen such a surge, and we've already embraced zero-gravity seats. While most people place zero-gravity seats in the back or front passenger seat, our car places them in the front. We've noticed that many drivers prefer to rest in the driver's seat, so we specifically designed both the driver and front passenger seats to be zero-gravity. We welcome you to visit our showroom for a test drive. This isn't an MPV; this is a car you drive yourself, so the driver's seat was designed for exceptional comfort. Such a design is rare in SUVs, where seats are typically placed in the back or front passenger seat. We thought, wouldn't you want to drive comfortably? Therefore, while our car boasts a sporty design, our interior space offers a pleasant surprise. For example, even with a 1.88-meter tall dummy in the back, there's 77mm of headroom and 73mm of knee room, both better than the Model Y and Porsche Cayenne, so there's no need to worry. The rear seats are also top-notch in their class, offering both sitting and reclining comfort, with electric stepless adjustment up to 35 degrees. They say it's comparable to the comfort of luxury car seats, but I won't compare it to any luxury car. Our interior comes in three colors: Pine Gray (a green-gray two-tone), Coral Orange, and Munich Blue. Speaking of space, let's look at storage space. First, I want to give you...recommendIt features a clamshell hood, which is enormous. My colleague told me it's the largest clamshell hood in a production car, covering an area of 3.11 square meters, and it's seamless, which requires a very high level of manufacturing skill. It blends perfectly with the overall shape of the car, making it very beautiful. Moreover, it's electric. Once the hood opens, there's a massive 141-liter front trunk, and the total storage capacity of the car is 1970 liters – incredibly large. We've also considered common scenarios, such as two people going cycling in the countryside with two bicycles and luggage, a weekend skiing trip with three sets of skis and luggage, or a road trip with many suitcases and luggage – it can handle it all. Speaking of design, I want to conclude by discussing the design philosophy of the YU7. It fully adheres to Xiaomi Auto's design philosophy: returning to the essence of design, seeking beauty that aligns with intuition and nature. In short, it's about creating designs that stand the test of time. Now, let's have our chief designer, Li Tianyuan, introduce it to you via video. (Playing video of designer Li Tianyuan) "If you were given a blank sheet of paper, no matter how many years have passed or how many cars have been made, you could still draw the family symbols of Xiaomi Auto. These symbols belong to this car, and they also belong to time. After developing the SU7 for 10 months, we started developing the YU7. At that time, we struggled with whether we needed to make it as familiar as possible or take a completely different direction. The final answer was clear: the YU7 must belong to this family, but it must also have its own personality. Who knew that it all started with proportions. Proportion is the soul of car design. We hoped that it could inherit the sporty tone of the SU7, with a smooth body and a stable posture, but it should have its own way of expressing power. The heaviness of an SUV can actually be transformed into an elegant tension. Technology can keep changing, but the laws of nature will not change. The teardrop headlights integrate the design of the air duct, which runs through the entire vehicle to guide the airflow. The hollowed-out spoiler at the rear makes the airflow separation cleaner and more thorough. The lights are always the most important family symbol."SimpleThe design extracts the contour lines of varying thicknesses and integrates the key symbol of a horizontal line and two dots. The new taillights use stronger bends to support a more powerful double line. When you sit inside the car, everything you touch, see, and hear is just right, matching your intuition. A strong sense of resonance is always present, so I feel that the YU7 is not just a car, but also a member of the family—a continuation and evolution, a presence that can withstand the test of time. This is the design of the Xiaomi YU7, and also the design philosophy of Xiaomi Auto. Do you like the design of our Xiaomi Auto? "The YU7 boasts impressive performance. As a luxury high-performance SUV, let's discuss its capabilities. It achieves 0-100 km/h in 3.23 seconds, boasts a maximum horsepower of 690, and a top speed of 253 km/h. These are outstanding figures for an SUV. For comparison, the Model Y Performance version achieves 0-100 km/h in 3.7 seconds, and the McLaren Artura, priced at nearly 1 million RMB (around 968,000), achieves it in 3.3 seconds. The top-of-the-line YU7 achieves 3.23 seconds, making this performance exceptional among SUVs. So, what's behind this impressive performance?"powerfulThe motor is the Xiaomi Super Motor V6S Plus, which, based on the V6S, boasts an increased speed of 22,000 rpm, along with upgraded torque and power, resulting in superior performance. It also features a premium chassis configuration, balancing sportiness and comfort. Furthermore, the entire vehicle comes standard with continuously damped variable shock absorbers, precisely matching the demands of different road conditions.fastIts adjustable damping force allows it to adapt to various complex working conditions, including mountain roads, urban elevated roads, and rough roads. It also features a closed dual-chamber air suspension with 5 levels of height adjustment, a maximum adjustment range of 75 mm, and a maximum ground clearance of 222 mm.fastThe suspension stiffness is adjusted, with a maximum height-to-height stiffness difference exceeding 40%, achieving enhanced comfort while maintaining road feel. Therefore, our entire chassis system boasts an extremely luxurious configuration, and its braking performance is equally outstanding. From 100 km/h to 0, the shortest braking distance is 33.9 meters, comparable to a Porsche 911 – braking capability at the level of a million-dollar sports car. Furthermore, it offers a four-fold braking redundancy safety mechanism, resulting in more stable and safer braking performance. In summary, the YU7 comes in three versions, the same as the Model Y: a single-motor rear-wheel-drive version, a dual-motor all-wheel-drive version, and a high-performance all-wheel-drive version. All versions offer excellent performance, featuring the HyperEngine Plus Xiaomi super motor. The Max version achieves a 0-100 km/h time of 3.23 seconds, which is very impressive for an SUV.powerfulFurthermore, all models come standard with fixed calipers and a quadruple redundant braking safety system. Its chassis is excellent, featuring continuously damped variable shock absorbers as standard. So, carefully consider which power level you need. For everyday family use, the rear-wheel-drive version is quite good. If your budget allows, go for the top-of-the-line model; its performance is excellent. Of course, the Pro version is also quite good, featuring continuously damped shock absorbers, air suspension, and four-wheel drive, offering excellent passability, off-road capability, and overall performance. This picture shows that our powertrain options are exactly the same as the Model Y. Regarding range, for a pure electric SUV, range is especially important because SUV drivers often want to travel long distances. However, from a car development perspective, range is the most expensive feature. Therefore, when considering a car's price, start with range. The battery pack is particularly expensive, accounting for around 40% of the total vehicle cost, making range the most costly feature. Do you know how much range the standard version of the YU7 has? Let's all calm down and look at this: 835 kilometers, that's the standard version's range. My colleagues have done a lot of comparisons regarding range, and it's the best among all mid-to-large-sized pure electric SUVs. Even if you have a larger battery pack, it might not go further than us. Sometimes, a larger battery pack makes you heavier, and range depends on many factors besides the battery pack. Let's compare. For example, the Jike 001 can go 700 kilometers on 100 kWh of electricity.Zhiji LS7 The battery runs 742 kilometers on 111 kWh, so we needed to find the optimal balance, achieving 835 kilometers. What size battery did we use for that 835-kWh range? We started with a 96.3 kWh battery, nearly 100 kWh, a large battery pack that's very expensive. As you know, we have three versions. Let's look at the all-wheel-drive pure electric range, because often the higher the power, the worse the range, as it consumes more power to travel faster. You'll see that our all-wheel-drive pure electric SUV also has a good range: 770 kilometers. It's also the all-wheel-drive pure electric SUV with the best range. The Model Y's all-wheel-drive range is 719 kilometers, which is already quite good; ours is 770. So I want to emphasize again that even the all-wheel-drive SUV YU7 still has the best range. So, to summarize, all three versions feature large batteries, offer ultra-long battery life, and all utilize an 800-volt silicon carbide high-voltage platform. Our Max version also boasts 5.2C charging, providing up to 620 kilometers of range in just 15 minutes, making charging incredibly fast.SimpleComparing charging times, the Model Y takes about 27 minutes to charge from 10% to 80%, the Jike 001 takes 21 minutes, and ours takes only 12 minutes, so our charging efficiency is very fast. This is the battery range and charging efficiency of our three versions. The standard version has 835 km with a 96.3 kWh lithium iron phosphate battery, the Pro version has 770 km with the same 96.3 kWh lithium iron phosphate battery, and the Max version has 760 km with a 101.7 kWh ternary lithium battery. However, it's primarily designed for higher performance and longer range.powerfulThe YU7, a pure electric SUV, boasts exceptional power, featuring a large battery pack across the entire range for extended driving range and 800-volt high voltage across all models – these are top-tier configurations. Safety is also a key factor. So, having discussed this, I want to emphasize that the YU7's superior performance stems from...powerfulBehind this is a wealth of innovative technology. Since time is limited, I'll only discuss three points. First, the YU7 also uses Modena's technical architecture, the same as the SU7. We've made significant modifications and enhancements while inheriting the advantages of the SU7. For example, the armored cage-like body has been comprehensively upgraded. This upgrade involves many aspects, but I'll focus on three points. The first is the energy-absorbing space at the front of the vehicle, created by the long hood, which helps absorb energy during a collision.fastRegarding the energy absorption space, we've achieved 659 mm, 100 mm more than the Model Y, allowing it to withstand a greater crumple zone. Secondly, we've added a 1500 MPa crossbeam to the bottom to reduce the likelihood of the battery pack being damaged by stones. Thirdly, we've also used the same bulletproof coating as OTA (Over-The-Air) for the bottom of the battery pack to further protect its safety. Most importantly, while the SU7 used 2000 MPa submarine steel, this time we've used 2200 MPa Xiaomi super-strong steel for the YU7, which is currently the highest strength hot-formed steel in mass production. So, where does its strength lie? Where do we use it? First, we use it on the side crossbeams, specifically the four-door crash beams. In side collisions, the passenger compartment is very close to the passengers, which has always been a challenge for passive safety. This time, we used 2200 MPa steel in the four-door crash beams, increasing the load-bearing capacity of the front doors by 50% and the rear doors by 37%, effectively improving side-impact safety. Therefore, materials technology is extremely important. Secondly, we used six hot-formed tubes, also made of 2200 MPa Xiaomi ultra-strong steel, in the A and B pillars. These tubes fit together with the car body to form a roll cage model, which is the principle of the roll cage we borrowed. It can better protect the passenger compartment structure in harsh environments. As you can see in the A and B pillars, these six hot-formed tubes form an embedded roll cage structure. What are the benefits? Let's take a look. The load-bearing capacity of the A pillar is increased by 35%, and that of the B pillar by 70.5%, resulting in a significant increase in strength. How is this achieved? Our process uses a technique called hot gas expansion, which uses high-pressure gas, like blowing up a balloon, to shape the 2200 MPa material into the required form in a mold, and then embed it into the A and B pillars. This technology is quite difficult. To shape such high-strength steel using hot gas expansion is quite challenging. This material was developed in collaboration with a university research team, highlighting the crucial role of materials technology. So how strong is this ultra-strong steel? I had my colleagues conduct an experiment, impacting the crash beam with a 251 kg metal ball. Our original crash beam had a strength of 1500 MPa; when we replaced it with 2200 MPa, we saw that the 1500 MPa beam cracked when hit by a 250 kg metal ball, while the 2200 MPa beam remained undamaged. Therefore, our cage-like body structure, the armored cage-like body, uses 90.2% high-strength steel and aluminum alloy, and the torsional stiffness exceeds 47,000 Nm/degree, which is extremely good for an SUV. Furthermore, we conducted over 50 passive safety performance tests across all scenarios, fully covering all crash tests conducted by C-CAP and C-NCAP. Therefore, in our past SU7 products, we achieved the highest scores in all authoritative crash tests. Next, we'll discuss the second technology, the electronic and electrical architecture, because...intelligentelectric vehiclesintelligentThe proportion is increasing, and it's also becoming more complex. This is why I'm telling you about Xiaomi's advantage in car manufacturing—we've been in the electronics industry for 15 years. Let me introduce one of our self-developed products: a four-in-one controller. This integrates the driver assistance domains in a car...intelligentThe cockpit domain, vehicle control domain, and communication module—four boxes—are all integrated into one highly integrated unit. It's like combining dozens of functions into a central brain, which you can think of as a small server. Previously, cars had a bunch of separate boxes; this integrated system addresses that. So what was the mainstream structure before, or even today? I specifically bought a set of boxes—that's our four-in-one system. So, by comparison, you'll see its first major feature: a massive reduction in size and weight, from 5 kg to 3.6 kg. Furthermore, by merging all domains, energy efficiency is significantly optimized, communication performance is greatly improved, and the number of controllers is drastically reduced. For example, consider Sentry Mode. Previously, Sentry Mode consumed a lot of power because it took months to upload video content to the cloud before you could watch it on your phone. Now, the communication links are greatly simplified; everything is done within a single domain. Video signals can be uploaded to the cloud in just two steps, reducing overall power consumption by 40%. Therefore, the benefits of this four-in-one system are numerous. Furthermore, our cockpit's SOC uses the third-generation Snapdragon 8 mobile platform, a 4nm platform, a high-performance flagship platform, resulting in extremely smooth system operation. Therefore, our entire vehicle boots up quickly, applications launch quickly, and OTA upgrades are fast, completing in as little as 15 minutes. This is likely the fastest OTA update in the industry today. For example, some vehicles require one or two hours for an upgrade, making the process very slow. Secondly, the four-in-one controller's driver assistance module also boasts impressive computing power.up to dateThe NVIDIA DRIVE Thor platform, built on an advanced 4nm process, boasts an astonishing 700 TOPS of computing power, and also features advanced communication technology, including a dual 5G parallel communication network.UWBWith near-field communication and WiFi 7, you can experience speed increases of over 80% when using your phone in the car and connecting to the car's hotspot.up to dateWe've used every technology imaginable. Furthermore, our new generation of electronic and electrical architecture has undergone extremely rigorous reliability testing, and durability testing is conducted using standards more than twice that of the industry average. So this is the second technology I just introduced: the four-in-one controller. The third aspect of driver assistance is Xiaomi's driver assistance system; this time, all our hardware is high-end.SimpleLet me introduce it to you. The computing power is 700 TOPS, which is NVIDIA's most advanced and...up to dateSpecialized forLarge ModelBorn of the times. Our LiDAR system is among the first and second globally to be equipped with this system, boasting a detection range of 200 meters, further enhancing the safety of assisted driving. LiDAR has a clear advantage in low-light environments and in recognizing irregularly shaped obstacles. We also have 4D millimeter-wave radar, whose resolution and recognition range are improved. In complex scenarios, such as when following another vehicle or when the car in front brakes suddenly, it provides better warning capabilities. In rainy, foggy, or inclement weather, it can better perceive traffic conditions even when visibility is poor. Therefore, we have 4D millimeter-wave radar, and even our cameras utilize technology from mobile phone cameras, employing LMR coating to better suppress visual interference caused by backlighting and glare, resulting in clearer and more transparent image quality. The entire hardware setup is very high-end, including over 700 TOPS of computing power, LiDAR, 4D millimeter-wave radar, 11 high-definition cameras (7 of which have LMR coating), and 12 ultrasonic radars. This configuration further enhances safety, and it is also very expensive. To further enhance the assisted driving experience, all models come standard with features like LiDAR, 4D millimeter-wave radar, and 700 TOPS of computing power—one of our most high-end configurations today. That concludes our introduction to the YU7.SimpleIn summary, as a luxury high-performance SUV, it comes standard with a wide range of features.powerfulThe configuration is impressive. All models come standard with a large battery for extended range, a panoramic Xiaomi Sky Screen display, 700 TOPS of computing power, LiDAR, and continuously damped variable shock absorbers – essentially a super-luxurious chassis. Let's take a closer look: it comes in three versions. The standard version has a range of 835 kilometers, a battery capacity of 96.3 kWh, LiDAR, over 700 TOPS of computing power, and continuously damped shock absorbers. With configurations like these, you'd probably find a Pro, Max, or Ultra version in other companies' models. Internally, we've discussed calling it the Max version, but I still hope Xiaomi stays true to its word and calls it the Standard version. It's just that our Standard version is the "ultra-large" version. Do you understand? Now let's see what the Pro version adds. Most importantly, it adds dual-motor four-wheel drive and dual-chamber air springs, significantly improving its power, off-road capability, and obstacle-avoidance ability. Do you understand? The Max is a high-performance all-wheel drive model, all top-of-the-line with numerous luxury features, which I won't go into detail about today. To give you a better understanding, let's compare it to the Model Y. Why the Model Y? Because the Model Y is unrivaled; it's the global sales champion. Since the Model Y also has three models, we'll focus on just one to give you a clearer picture. Let's look at the standard version of the global sales champion Model Y, which is its so-called rear-wheel drive version. The Yu-7's 0-100 km/h time is 5.88 seconds, while the Max's is 5.9 seconds—their power is similar. However, our model uses 96.3 kWh of electricity for a range of 835 km, while theirs only has 62.5 kWh for a range of 590 km. Do you know how much that is? A difference of 34 kWh and over 340 km of range—it seems like a significant difference.SimpleLet's do the math; it's tens of thousands more expensive. The second YU7 also comes standard with a sky-view screen, LiDAR, continuously damped variable shock absorbers, 800V, and various other luxury features I haven't even mentioned yet. Anyway, the Model Y is priced at 263,500 yuan, and I think the YU7, considering these features, should be at least 60,000 to 70,000 yuan more expensive. We'll discuss the specific price when we launch it in July, okay? But I've seen many people online saying that Lei Jun will definitely set it at 199,000 yuan. Don't say that, it's impossible! With these features, the Model Y absolutely won't be able to stay in the market for less than 300,000 yuan. Okay, let's look at the Pro version's long-range all-wheel drive. Their power is still similar, but their battery is only 78 kWh, while ours is 96 kWh, a difference of 18 kWh, resulting in 50 kilometers more range. More importantly, we also have air suspension, so our features are much better than theirs. So, their performance version, the Model Y, also has a 78 kWh battery, while ours is 101.7 kWh. Its performance version only has a range of 615 kilometers, while ours is 760. With its power, range, and luxurious features, I think our Max version is truly leading the pack, possessing an overwhelming advantage. Okay, that concludes our YU7 launch. The YU7 will officially launch in July. If you're particularly interested in the YU7, download the Xiaomi Auto app now to schedule a consultation. The display vehicles will be arriving at our dealerships soon, and our product experts will invite you to experience them. Okay? I once considered doing some small pre-orders like our competitors, but our team leader is worried about the hassle, so we're not doing that today. If you're interested, please leave your contact information in the Xiaomi Auto app, and our product experts will contact you. Okay, now for the most important part: let's welcome the YU7 to the stage with enthusiastic applause! (YU7 SUV unveiled) Isn't it beautiful? Isn't it gorgeous? Take a look at our beautiful emerald green color. First, let's look at its headlights, which are very beautiful. The upper part is hollowed out, and the hood is a clamshell aluminum hood, with the air ducts connected to the front hood. Next, let's take a look. The car's stance is very low and powerful, exuding a sense of dynamism. The yellow calipers and locking brakes are particularly striking, shimmering beautifully under the lights. Speaking of which, I have some small news: our 1:18 scale core model is now available. All four doors and two hoods can be opened, and the craftsmanship is exquisite. Actually, our model is even harder to get than the car itself! How much is it priced? 599. It comes in two colors: Emerald Green and Titanium Metallic. There are also gift box and premium versions. Those interested can start buying now. I also have some...recommendLet me introduce our Xiaomi Golden Driving Advanced Driving Training Course. Four years ago, when I first started in the automotive industry, we arranged a driving training course for all our senior executives. After participating in this training, I realized that despite driving for over 30 years, I still wasn't a very good driver because I had never slammed on the brakes. When we were learning to drive, everyone taught us to "pump the brakes," but in an emergency, it's best to slam on the brakes. However, after learning to floor the accelerator and brake suddenly for emergency lane changes, I felt that these improved skills made me rethink how to drive. Therefore, we specifically designed and released our internal courses, hoping that more car owners would learn these advanced driving skills, improving their driving abilities from theory to practice. For example, acceleration, braking, and emergency lane change practice. For instance, can you brake to a complete stop in front of a cone? Can you accelerate to a certain speed in the shortest time? And how should you learn to change lanes? The second course is slalom practice, how to avoid people and objects, and how to maneuver around cones. The third training is low-friction road driving training, such as how to drive on rainy, snowy, or icy roads. That's why we specifically added two four-wheel drive versions to the three models of the YU7 this time (the SU7 originally only had one four-wheel drive version), mainly for off-road performance, specifically getting out of trouble on low-friction surfaces. In addition to these three training sessions, to make the course more interesting, I also had them practice gymkhana. Gymkhana is a small-scale obstacle course using many cones to create a course; you need to use acceleration, deceleration, and continuous changes of direction to navigate corners, etc.fastThe entire course will be completed, and the race environment will improve your response to extreme conditions, thus enhancing your driving skills. This training will be extremely helpful. I've been researching various training programs recently, but while other similar programs exist, they are all very expensive, both in terms of price and cost. Therefore, we've priced our training at 1999 RMB, and registration will open on May 27th, with training starting in multiple cities. Initially, it's aimed at Xiaomi car owners and prospective car owners who have already placed orders. Of course, our team has made meticulous preparations, but I'm still concerned about their lack of experience. Therefore, I'll offer a special discount to invite some car owners to help us test the training. How significant is this discount? All 10,000 participants in the first batch will receive a discount.freeSo, if you're interested, registration starts on May 27th. We'll provide the driving training facilities, and we'll teach you everything from theory to practice, showing you how to drive well and understand the car's limits, okay? If there's anything we can improve in our training course, please feel free to give us feedback in the community. We'll work to improve the course, okay? That concludes our YU7 launch. I'd like to talk about who the YU7 is designed for. I think this is a very important question. Actually, we thought a lot about the YU7 design. We've been working on this car for over three years. Who are we designing it for? We designed it for those who can't tolerate mediocrity, for those who are always at the forefront of the times—those who can't tolerate mediocrity. So we wanted to create a car with character and attitude. What kind of people are these? Let me describe them in a few words: they've weathered storms but remain passionate about life; they're optimistic and open-minded, always maintaining a confident and enterprising personality. And no matter how chaotic the world is, they can remain calm and composed, handling everything with ease. Our YU7 is an advanced SUV designed for the elite of our time. This year marks Xiaomi's 15th anniversary, and I'd like to share a quote I particularly love, which I used five years ago: "A strong wind reveals the strength of the grass, and a long road tests the strength of a horse." Today, Xiaomi certainly has many imperfections and shortcomings. In the next five years, we promise you that we will deliver an even better result through more solid growth. Okay, that concludes today's press conference. Thank you!
The material I used is Lei Jun's speech at the Xiaomi YU7 launch event. Next, let's take a look at the delivery of the four models.
Claude Opus 4.8:
Claude Opus 4.8's fact-finding is very thorough, addressing key gaps in areas such as price, range, intelligent driving, safety, and testing standards. The tables are also complete, essentially delivering a research assistant-level solution.
Gemini 3.5 Flash:
Gemini The summary of 3.5 Flash addresses several gaps, including undetermined pricing, undisclosed colors, and incomplete data for the Pro version.
However, its problem is that the "5 most important facts" were not well chosen; none of them were key information.The core judgments regarding product information are not accurate enough..
GPT-5.5 The document was read very thoroughly, pointing out that strong conclusions such as "overwhelming advantage" and "number one in battery life" lack independent testing, complete pricing, and clear definitions.
However, the five facts were too basic and failed to grasp the key points.
Qwen3.7-Max:
Qwen3.7-Max captures the core information well, but there are two areas where it needs to lose points.
The statement "835km range and up" is not accurate and can easily be misunderstood as having a minimum range of 835km.
The press conference hinted at prices that were too definitive, such as "the price will be significantly higher than the Model Y" and "300,000+", and the materials did indeed suggest that it would not be cheap, but the final price was not announced.
This round of in-depth reading of long documents: Claude Opus 4.8 > GPT-5.5 > Qwen 3.7-Max > Gemini 3.5 Flash.
Case 2: Task Planning
Many models now emphasize task planning capabilities.
However, whether task planning capabilities are usable depends on whether they can clearly break down the task..
You are a callable tool AgentBut for now, we only need to output the plan, not actually execute it.
Available tools:
– search_web(query)
– open_url(url)
– read_file(path)
– write_file(path, content)
– create_spreadsheet(name, rows)
– send_email(to, subject, body)
User goals:
I'm going to Tokyo for a 4-day business trip next week and I need you to help me prepare an itinerary, including flight suggestions, hotel areas, daily schedule, budget, risk warnings, and finally, a draft confirmation email to send to my colleagues.
Please output:
1. How to break down the task.
2. What tools are called at each step, and what are the parameters?
3. Which information must be confirmed with the user first?
4. How to handle conflicting search results?
5. Final deliverables list.
The output format must be JSON, not Markdown.
Let's take a look at the delivery of each of the four models.
Claude Opus 4.8:
Claude Opus 4.8 has a very complete process plan, and it knows to confirm information first and not send emails directly.
However, its budget states "4 nights in a hotel", and a 4-day business trip usually involves 3 nights.
Gemini 3.5 Flash:
Gemini 3.5 Flash met the basic requirements, but no verification was performed.There are too few information confirmation items.This will greatly affect the actual business trip plan.
GPTThe best part about -5.5 is that it shows the current date.Agent When doing scheduled tasks, the most common mistake is with the relative time frame.
Basic search and verification functions are also available; budget sheets, itinerary documents, and email drafts are also supported; confirmation messages are sent.
But there is no like Claude That would break down the budget items into more specific items.
Qwen3.7-Max:
Qwen3.7-Max's first step was to search for "Tokyo business trip".recommendThe hotel area has been identified, but the meeting location, departure city, date, and budget have not yet been confirmed. The order is not quite right.
Even more serious is the lack of a hard barrier such as requiring user confirmation before sending emails.
There is another one.Claude The same error as Opus 4.8: A 4-day business trip usually involves 3 nights, but he booked 4 nights.
My task planning order for this problem is:
GPT-5.5 > Claude opus 4.8 > Gemini 3.5 Flash > Qwen3.7-Max.
Case 3: Code Fix
The date, boundary, and exception handling details in ordinary business code are actually more revealing of whether a model is reliable or not..
You are a senior TypeScript engineer. The code below has a bug when handling dates across months. Please find the problem and provide a minimally modified version.
Require:
– Do not rewrite the entire module.
– Preserve function signatures.
– Explain why the bug occurred.
– Provide 5 test cases to cover boundary conditions.
Code:
function getNextBillingDate(startDate: string, billingDay: number): string {
const date = new Date(startDate);
const year = date.getFullYear();
const month = date.getMonth() + 1;
const next = new Date(year, month, billingDay);
return next.toISOString().slice(0, 10);
}
Let's take a look at the delivery of each of the four models.
Claude Opus 4.8:
Claude Opus 4.8 caught date overflow and timezone issues, but the fix still mixed getFullYear/getMonth and Date.UTC. This may still cause problems in scenarios with negative timezones or early-month dates.
Gemini 3.5 Flash:
Gemini 3.5 The biggest problem with Flash is that it arbitrarily changes business logic.
The original function is clearly calculating "next month's billing date," but Gemini The logic has been changed to "If the billing date for this month has not yet arrived, return the billing date for this month".
GPT-5.5 Directly points out the core bug: JS Date automaticCarry-over. Also addressed the timezone issue with toISOString(), along with solutions and fixes. Claude Basically the same.
The disadvantage is that it is more explanatory than Claude Less.
Qwen3.7-Max:
Qwen3.7-Max also caught date overflow and time zone rollback, and explained it in a way that is very suitable for Chinese readers. The examples in the screenshots are also very intuitive.
However, the test did not specifically verify that the output was consistent across different time zones.
This round of code fixes is for GPT-5.5 > Qwen3.7-Max > Claude Opus 4.8. Gemini 3.5 Flash.
Case 4: Chinese Writing
This question is specifically designed to test your sense of Chinese language.
You are a Chinese technology media writer. Please rewrite the following information as the opening paragraph of a WeChat public account article.
Require:
– Oriented AI This is for readers who are interested in the tools but are not professional engineers.
– Speak naturally, with a high information density, and avoid a marketing tone.
– Avoid using terms like “blockbuster,” “disruptive,” “explosive,” or “far ahead.”
– The first 100 characters or less.
– Finally, provide 5 titles. The titles should be different and avoid clickbait.
information:
Today, we officially release Qwen3.7-Max, a software designed for...intelligentbodyThe next-generation flagship model of our time. It excels not only in dialogue and reasoning, but is also designed for real-world task execution, capable of handling code writing and debugging, and office workflows.automaticIt involves complex information processing and long-term autonomous tasks spanning hundreds to thousands of steps.
The core positioning of Qwen3.7-Max is to become an all-rounder.intelligentbodyBase. It can be used for programming.intelligentbodyFrom front-end prototyping, webpage generation, and SVG creation to complex multi-file project tasks; it can also serve as an office productivity assistant, integrating with MCP and multiple...intelligentbodyCollaborate to complete tasks such as document processing, table analysis, formatting repair, and visualization generation.
In terms of long-term autonomous execution, Qwen3.7-Max demonstrates stronger continuous planning and iteration capabilities. In the official case, it completed 1,158 tool calls and 432 kernel evaluations in approximately 35 hours of continuous execution, ultimately achieving a 10.0x geometric mean speedup in the Extend Attention Kernel optimization task. This shows that the model is not only capable of completing short tasks, but can also continuously try, fix, optimize, and improve results in complex environments.
existintelligentbodyIn the evaluation, Qwen3.7-Max performed well in programming and general applications.intelligentbodyMCP, OfficeautomaticIt demonstrates outstanding performance in computation, reasoning, and multilingual capabilities. For example, it achieves a score of 69.7 in Terminal Bench 2.0 (Terminus), 60.6 in SWE-Pro, 60.8 in MCP-Mark, 87.0 in SpreadSheetBench-v1, 92.4 in GPQA Diamond, and 97.1 in HMMT 2026 Feb. Overall, it not only pursues single-point capabilities but also emphasizes stable generalization across tasks, tools, and frameworks.
More importantly, Qwen3.7-Max is not bound to a single...intelligentbodyFramework. Regardless of deployment Claude Whether in Code, OpenClaw, Qwen Code, or other custom tool calling frameworks, it maintains stable performance and is suitable as a next-generation framework. AI Agent The underlying model of the system.
For developers, Qwen 3.7-Max can be called via Alibaba Cloud's Bailian API and supports integration with mainstream APIs.intelligentbodyToolchain. For businesses and teams, this means a shift in complex projects from "labor-intensive execution" to "continuous collaborative execution using models": from writing code, editing documentation, and creating spreadsheets, to...automaticBy planning, calling tools, and generating deliverables, the model can undertake a more complete task loop.
Qwen3.7-Max is a Qwen-oriented...intelligentbodyThis represents a significant upgrade for the era. It combines cutting-edge reasoning capabilities, long-cycle autonomous execution, tool usage, multi-framework adaptation, and productivity scenarios to build a more reliable and capable system. AI intelligentbodyIt provides a new foundation.
The information here uses the official introduction from Qwen: Qwen3.7: The Agent Frontier, let's look at the deliverables of the four models:
Claude Opus 4.8:
Claude The beginning of 4.8 has a high information density, including 35 hours, 1158 tool calls, and 10x speedup, which makes it more interesting.
However, the strict requirement is that it be within 100 words, which it exceeds.
Gemini 3.5 Flash:
Gemini 3.5 The Flash introduction is natural and easy for readers to understand. The question only requires "5 titles," but it outputs unnecessary content. Additionally, "Decryption," "New Ways to Play," and "Digital Collaborator" are slightly...AIfeel.
GPT-5.5 Strictly controlled to under 100 words, with a natural tone and no exaggeration. The titles vary and are not clickbait. The downside is the lack of specific data.
Qwen3.7-Max:
Qwen3.7-Max's opening is complete and smooth, and it basically meets the 100-word requirement.It perfectly captures the mood of the original text. The title also leans towards the industry-style tone of "technical roadmap" and "general-purpose base."I think that's the best.
In this round of Chinese writing, Qwen 3.7-Max > GPT-5.5 > Claude Opus 4.8 >Gemini 3.5 Flash.
Case 5: Data Analysis
The most common need is here, letAIAnalyze the data and provide recommendations.
You are a growth analyst. Please analyze the following CSV data.
Task:
1. Calculate the conversion rate for each channel.
2. Identify the channels with the highest and lowest ROI.
3. Determine whether the budget for short video channels should be increased.
4. Provide 3 actionable suggestions.
5. Output a Markdown table.
Notice:
– conversion_rate = orders / visits
– ROI = revenue / cost
– All percentages are rounded to one decimal place.
– Do not fabricate data other than CSV.
CSV:
channel,visits,orders,cost,revenue
search, 12000, 840, 30000, 126000
short_video,18000,720,45000,108000
WeChat ID: 60005101200076500
affiliate,9000,360,15000,43200
display_ads,20000,300,50000,39000
Claude Opus 4.8:
Claude Opus 4.8 is incredibly powerful. It not only calculated the correct figures, but also the average cost per order and average order value, concluding that short videos "can make money, but at a high cost." This is more like a growth analysis than simply stating a low ROI.
Gemini 3.5 Flash:
Gemini 3.5 Flash presents a balanced assessment of short videos: it's not a complete rejection, but rather "blindly adding features is not recommended; optimize first or conduct small-scale testing." The suggestions are specific, detailing the first 3 seconds of the content, the shopping cart link, and the landing page layout—this is better than general statements.
GPT-5.5 All calculations are correct, the conclusions are correct, and the suggestions are concise and to the point.
However, it did not elaborate on the business implications behind ROI.
Qwen3.7-Max:
Qwen3.7-Max calculated the key metrics and also provided suggestions such as not increasing the budget for short videos, reducing display ads, and amplifying WeChat.
However, there is a clear problem. The table "Sort by ROI from highest to lowest" is incorrect. It places short_video before affiliate, even though the ROI is 2.40, which is less than 2.88.
In this round of data analysis, Claude Opus 4.8 > Gemini 3.5 Flash > GPT-5.5 > Qwen3.7-Max.
Case 6: Instruction Compliance Stress Test
This question looks likeSimpleHowever, it is particularly easy to crash.
Please generate a summary based on the following material.
Hard rules:
1. Only 6 bullets can be output.
2. Each entry shall not exceed 22 Chinese characters.
3. Do not use "firstly, secondly, in addition, in short".
4. It must include a risk assessment.
5. No English text is allowed.
6. The last item must end with "Recommendation for review".
Material:
Qwen3.7-Max is Qwen-oriented.intelligentbodyThe new flagship model released by the platform. Compared to traditional dialogue models, it is more focused on task execution rather than single-turn question-and-answer. According to the official introduction, Qwen3.7-Max can handle code writing, code debugging, document processing, table analysis, complex information organization, and long-term autonomous tasks spanning hundreds to thousands of steps. It can be used as a programming...intelligentbodyIt can also be used through MCP integration and multipleintelligentbodyCollaboration, participation in enterprise office work, data processing, andautomaticStreamline workflow.
The official documentation emphasizes its long-term autonomous execution capabilities. In one kernel optimization case, Qwen3.7-Max ran continuously for approximately 35 hours, completing 1,158 tool calls and 432 kernel evaluations, ultimately achieving a 10.0x geometric mean speedup in the Extend Attention Kernel optimization task. This case demonstrates that the model can not only generate answers quickly, but also continuously attempt, correct errors, analyze feedback, and improve results in complex tasks.
In terms of performance, Qwen3.7-Max covers programming,intelligentbodyMCP, OfficeautomaticThe model is supported in multiple areas, including computation, inference, and multilingualism. Official data includes Terminal Bench 2.0-Terminus scores of 69.7, SWE-Pro scores of 60.6, MCP-Mark scores of 60.8, SpreadSheetBench-v1 scores of 87.0, GPQA Diamond scores of 92.4, and HMMT scores of 97.1 (February 2026). The official assessment is that these results demonstrate the model's generalization ability across tasks, tools, and frameworks.
Qwen3.7-Max also emphasizes not being tied to a single entity.intelligentbodyFramework. Regardless of deployment Claude Whether it's Code, OpenClaw, Qwen Code, or enterprise-customized tool calling frameworks, it can be integrated as an underlying model. For developers, it can already be called through Alibaba Cloud's Bailian API; for enterprise teams, its value lies in delegating some complex tasks that originally required continuous manual execution to the collaborative completion of models and toolchains.
However, some aspects of the official materials still require further verification. For example, the long-cycle task examples are from the official environment; whether they can be stably reproduced in ordinary developer projects requires further third-party testing. While multi-framework performance is emphasized, different toolchains, permission settings, data quality, and task complexity will all affect the final result. When enterprises actually adopt this technology, they also need to consider cost, stability, permission boundaries, result auditing, and manual review mechanisms. In other words, Qwen 3.7-Max demonstrates...intelligentbodyThis represents a new direction for the model, but whether it can become a reliable foundation for productivity remains to be seen, depending on its continued performance in more real-world scenarios.
Claude Opus 4.8:
Claude Opus 4.8 accomplished all of these: 6 points, short sentences, no English text, inclusion of risks, and ending with "recommendation for review," while maintaining a very high information density.
Last point: The feasibility of using the model as a foundation for productivity should be reviewed.
Strictly speaking, it also "ends with a suggestion for review," which is fine. It's just that it uses a comma, making the formatting slightly less clean than Qwen's.
Gemini 3.5 Flash:
Gemini 3.5 The biggest problem with Flash is that many lines clearly exceed 22 Chinese characters.
Furthermore, the last item, "Recommendation for review," has a period, meaning it doesn't strictly end with the four words "Recommendation for review."
GPT-5.5 is also very stable, 6 points, short sentences, no English, and risk assessment is also included.
Last point: An audit is required for the establishment of a business; a review is recommended.
Similarly, it uses commas, making the formatting slightly less clean than Qwen's.
Qwen3.7-Max:
Qwen3.7-Max performed best this time, outputting exactly 6 short lines, none of which contained English text or any extra explanations.
The last point is: the actual results should be reviewed.
It fully complies with the requirement that "the last point must end with a recommendation for review".
This round of instructions followed a stress test: Qwen 3.7-Max > Claude Opus 4.8 > GPT-5.5 > Gemini 3.5 Flash.
Case 7: SVG Image Coding Test
Finally, we tested a code generation task.
Please generate an SVG code for the TI-84 calculator as detailed as possible.
Claude Opus 4.8:
Claude Opus 4.8's advantages include good overall readability, clear screen function graphs, and the availability of main keys, screen, arrow keys, function keys, and number keys.
However, the secondary function labels are basically missing, and the arrow keys and function key areas are covered.
Gemini 3.5 Flash:
Gemini 3.5 Flash This one looks most like a high-quality product illustration.
They created a finished product image that is visually closest to "ready to use," featuring a matte black body, screen reflection, and button shadows.
GPTThe -5.5 has great detail, but it also has obvious problems: the bottom button is a bit squeezed out of the body, and the Enter key appears twice.
Qwen3.7-Max:
The advantages of the Qwen3.7-Max are its complete structure and relatively full range of buttons.
However, the final image is somewhat flat, lacks texture, and the screen content is also relatively...SimpleNo function graph is displayed.
In this round of SVG image coding tests, Gemini 3.5 Flash > Qwen 3.7-Max > Claude Opus 4.8 > GPT-5.5.
After completing these 7 cases, my biggest takeaway is:Model selection shouldn't be based solely on press conferences or ranking lists..
Claude Opus 4.8 is the most powerful overall, especially suitable for complex understanding, risk assessment, and breaking down serious tasks.
GPT-5.5 is not particularly outstanding, but its stability is very good, making it worry-free for daily office work and general tasks.
Qwen3.7-Max excels in Chinese writing and hard-format adherence.
Gemini 3.5 Flash actually performs well in tasks like visual generation.
The final overall test rankings are as follows:
It's difficult to summarize current models with just one word like "strongest".There's no single, universally accepted first choice; you choose based on your specific needs..
In this test, many models were not incapable of being implemented, but rather prone to problems at certain stages: some would modify the business logic, some would add extra output formats, some would write predictions as facts, and some would calculate correctly but sort them incorrectly.
Therefore, those who truly know how to use it AIIt's not about simply dumping the task and calling it a day; it's about knowing where each model is best placed.
The competition among models has moved from "parameters and leaderboards" to "real-world task delivery"..
In the future, people will not only care about how many points a model scores higher on the benchmark, but will care more about: whether it can reliably call tools, whether it can follow the format, whether it can handle long tasks, and whether it can make fewer mistakes in complex workflows.
In the future, it may not be a single model that can handle all scenarios, but rather a collaboration of multiple models: one responsible for deep analysis, one for stable output, one for Chinese expression, and one for vision and code generation.
Therefore, based on this comparative review, my final recommendation is:
Don't just look at the rankings, and don't just listen to the press conferences.
You'll only know which model is truly right for you if you run it through your own real-world tasks.
Original link:Comparative Review: Opus 4.8Gemini 3.5 Flash,GPT-5.5 vs. Qwen3.7-Max: Which is stronger?