AB
AiBoss
Tutorials

Real-world test: Claude Opus 4.6 vs GPT-5.3-Codex – which is the better programmer?

The domestic AI battle, from the "Yuanbao Red Envelope" game to the "Thousand Questions Milk Tea" contest, truly shows that each generation has its own unique approach. Internationally, they're not idle either, with the release of two heavyweight models in the AI programming field: Anthropic's Claude Opus 4.6 O...

实测 Claude Opus 4.6 vs GPT-5.3-Codex ,编程谁更强?

Hi everyone, this is a place to solve problems by using...AIOrange Sister.

Every day when I wake up,AIThe world is full of new sights.

DomesticAIThe battles, from gold ingot red envelopes to thousand-question milk tea, truly show that each generation has its own "egg" strategy.

Foreign countries aren't idle either.AITwo heavyweight models have been released in the programming field:

Anthropic's:Claude Opus 4.6

OpenAI of:GPT-5.3-Codex

Both companies dropped their bombshells almost simultaneously, and we were sipping Qianwen milk tea when...fastLet's take a look at the advantages and highlights of the two models.

Claude Opus 4.6

One of the most significant upgrades is the support for context windows containing 1 million tokens, roughly the size of several long novels. It excels at complex task planning and long-cycle task processing, identifying problems and attempting to correct errors independently.

It excels in professional fields such as finance and law, with a GDPval-AA score leading the industry. (Added)Adaptive ThinkingandFour levels of effort (Effort control)The system also features enhanced integration with office software such as Excel and PowerPoint.

Model Key Performance Indicators

  • Terminal-Bench 2.0 (Evaluation)intelligentbody(Coding ability)The highest score was 65.4%.
  • GDPval-AA (Assessment of Professional Knowledge in the Field)The Elo score is 1606, which is 190 points higher than its predecessor, the Opus 4.5.GPT-5.2 is 144 points higher.
  • Arena.ai Comprehensive ReviewIt ranks first in all three major categories: code, text, and expert.

GPT-5.3-Codex

speedCompared to the previous generationGPT-5.2-Codex 25% fasterThe token usage will be reduced by half when completing the same task.The first to participate in its own creationAIModel.Supports multi-task parallel processing, suitable forfastIterative development.

Users can interrupt and adjust the direction of tasks in real time during execution, and can view progress, ask questions, and suggest corrections. This covers the entire process from writing requirements documents, writing code, debugging and deployment to monitoring and analysis.automaticTransformation. The first to obtain OpenAIA model for "high-capability" cybersecurity ratings.

Model Key Performance Indicators

  • Terminal-Bench 2.0 (Evaluation)intelligentbody(Coding ability)Score: 77.3%.
  • GDPval (Comprehensive Knowledge Job Assessment)Performance andGPT-5.2 unchanged.
  • SWE-Bench Pro (Evaluates the ability to solve real-world GitHub repository problems)Score: 56.8%.

I just bought a limited-time package of two models, so I'm here to try them out!

Purchase package here:https://vipcheap.com/zh

For overseas websites, being able to accept WeChat Pay is truly amazing. I've personally tested it, and it's absolutely safe.

1. Game Remake

Besides the wildly popular WeChat red envelopes, there was also a very popular mini-game called "Jump Jump" that we still remember vividly. Let's recreate it.

accessClaudeOfficial website, enterPrompt wordsSelect model Opus 4.6

Please help me recreate a version of the WeChat mini-game "Jump Jump".

It was generated in less than three minutes.automaticIt provides a 3D perspective and game effects.

After testing, I optimized it several more times:

Please optimize the front-end page, including: 1. The landing area can be objects of different shapes. 2. Adjust the jumping object to a 3D little dragon. 3. Center the game area on the page at the end. 4. Enhance some game effects as needed.

Please change "3D Little Dragon Character" to "Colorful Little Pony"

Let's optimize the "Perfect Landing" effect. In addition to adding points, it will display a different "Year of the Horse Blessing" prompt each time.

The generated effect feels similar to WeChat's "Jump Jump" game, but the UI is not very advanced.

Let's take a look at the gameplay:

To be honest, it's quite addictive to play.

Let's take another look.GPT-5.3-Codex

I used it on VSCode, the same.Prompt wordsEnter it, select GPT-5.3-Codex model.

It only took three minutes to complete.

This effect stunned me. Doesn't it have any artistic talent at all?

I also optimized several versions.Prompt words

To be honest, I didn't even recognize the pony; it's practically fallen off the screen...

2. Generate community service stations

Let's look at Opus 4.6 first. We input...Prompt words

Please design and generate a community service station. The product should primarily provide daily life services to community users, such as dog walking, cat feeding, and package pickup. Users can log in and use it via a mini-program. There are a few things you need to pay attention to:
1. The product should have as many functions as possible, while ensuring practical usability.
2. The design of the service station's functions must ensure the smooth interaction logic between the user and the platform.
3. The coverage area can include cities at or above level four.
4. It should include new types of elderly care services.

I was a little stunned when I saw the results.

That's it.SimpleWith just a few words, the generated result is a complete product that can be directly launched and used, with all interaction links working seamlessly.

It also filled in the service categories I hadn't mentioned.

You know what, Opus 4.6 is actually pretty powerful!

Similarly, let's take another look.GPT-5.3-Codex

Enter the samePrompt wordsIt took 3 minutes to complete, which is comparable to the speed of Opus 4.6.

Let's see how it performs, how durable is it?

What is this? It generated a website introducing a web service? Didn't you understand the actual use case of the product I mentioned? —(Users can log in and use the mini-program.).

3. Mini Program source code review

I previously developed a mini-program called "Are You Crazy?"Only five steps to hand-rubbingAI"CrazyMe APP" – Discover the next product with tens of millions of users. This is the one. There's a minor bug in the user experience; reports cannot be viewed after downloading.

I downloaded all the source code of the mini-program, and I'll let Opus 4.6 test whether it can detect it.

This is a list of all the source code files for this mini-program. Please check the code for bugs, shortcomings, or areas that need optimization.

Opus 4.6 really detected it.

Impact: After successful login, users are unable to redirect to the homepage; after submitting on the homepage, users are unable to redirect to the report page; the "Reanalyze" button on the report page is disabled. All three core processes are broken.

They detected so many bugs and issues and even provided corresponding optimization suggestions.

Based on the code review report, please generate the new, optimized application source code and create the new application.

The new app retains the previous UI design and style, only fixing the different workflow issues. It truly delivers on its promises, like an engineer who really understands you.

GPT-5.3- Codex did not return any results.

Prompt wordsLet's say there are 7 people in our department, named Xiaohong, Xiaolv, Xiaobai, Xiaohei, Xiaohuang, Xiaoqing, and me. We're planning a team-building activity, and I've designed a word chain game. Please simulate everyone and play the game. The rule is: each line begins with the last three words of the previous line, and so on, until you reach a closed loop with the first line. First line: Busy every day, busy every day, busy for a lifetime until you go to heaven.

It looks a bit too plain, and there's nothing funny about it.

Please change the style to an entertaining, dramatic type.

This version is pretty good; it perfectly captures the scenario of working professionals.

andGPT-5.3-Codex: I didn't fully understand the game rules at the beginning; I only remembered "follow the next three words."

Laughing and crying.jpg, covering face.jpg

Overall, Opus 4.6 feels very comprehensive, andGPT-5.3-Codex is a bit confusing to me. It's supposed to be all about programming, so are the remakes of games and the community service site on it for real?

For non-technical people like us,Claude Opus 4.6 seems more suitable, as it can think one step further based on your task, supplement any missing parts, and truly think and execute from a human perspective.

AIToday, programming is no longer about "assistance," but about "autonomy." It's about understanding complex requirements, planning development processes, executing multi-step tasks, and even self-debugging.

For technical personnel, their advantage lies in their ability to determine which option is most suitable based on actual development needs and scenarios, or to combine different options to leverage their respective strengths.

The two models represent different technological approaches:

Opus 4.6 emphasizes "breadth and depth" (large context + multi-agent collaboration), stresses "avoiding errors", and focuses on reliability and stability of long-term tasks.

GPT-5.3- Codex focuses on "speed and autonomy" (real-time interaction + self-iteration), emphasizing "get it running first" and focusing on speed and multi-tasking capabilities.

What do you think of Opus 4.6 and... GPTWhich of the 5.3-Codex versions is best for you?

Original link:Claude Opus 4.6 and GPT-5.3-Codex Real-world Testing: Which is Stronger?