Open AI Models vs. Leading Edge AI Models For Coding Work
Recently I’ve been pushing the local AI server to do more work in an attempt to shift work away from cloud computing. Having a small AI server that runs on at most a few lightbulbs-worth of power while completing tasks seems like a good idea. Sure, the tasks will take longer than they do in the cloud, but there are plenty of things that can be handed off that are not urgent.
With that in mind, I recently set up the local DGX Spark to run some open models. These models are designed to help write the code for one of my online apps. I provide a short description of things that need to be changed in a markdown file. These are small tasks that any AI agent should be able to complete easily. These tasks are easily handled by the latest foundational models provided by OpenAI and can finish the task quickly and in a single pass. With that said, I decided to install the latest qwen3-coder-next:latest model on the Spark. This model is deemed one of the best open source (or Open Models) available today. These are models you can download to an AI-capable server like the DGX Spark or an Apple Mac mini and run with a local LLM processor or agentic application platform such as OpenClaw.
My hope was the latest open model, Qwen 3 Coder in this case, could do the work at a reasonable level of quality despite taking longer to complete the task. The idea is to put multiple short “fix this , patch that, change this” text descriptions into a folder that the “Babbage Sprocket” agent would read and process overnight. When I wake up, multiple small patches are awaiting my review for inclusion in the final product.
I’ve done plenty of test runs with various models and have become fairly adept at writing short succinct requests for the AI bots. The cloud based models since GPT 5-5 from this past January do a great job on the first pass, providing me with ready-to-integrate code changes. I figured the latest Qwen 3 Coder model from this past July should do the trick.
The instruction set markdown, which is embedded in a deep context bootstrap that fully describes the app and standard operating procedures related to the software development lifecycle.
Title: Location List Bulk Actions Revision
Updated: 2026-09-24
Status: Completed
---
In the store locator plus plugin, go to the locations list panel and look for the bulk Actions drop down, apply, and apply to all button.
The first revision is to move this entire Bulk actions interface into the top of the Data Grid Pro component.
- Place Bulk Actions drop down and related buttons on the left of the DataGridPro Header
- Place the Filter and Density DataGridPro interactive elements directly to the left of the add location icon button.
- Place the Quick Search icon and input interface to the left of the Filter and Density UI elements
This image is the UI target.

Next, style the bulk actions drop-down.
It should be using reactionary standard components for the drop-down.
If none exists, follow the SLP power add-on.
Look for the import schedule drop down that is part of the remote file imports interface.
Mimic that drop-down style.
In my instruction set I included a mockup image of the new interface, as shown below.

Qwen 3 Coder Is Not OpenAI Codex Bot
Turns out I was wrong. Qwen 3 Coder did a horrific job of implementing the user interface changes that I was requesting for my app. It made a complete mess of a pre-existing interface that worked well. Instead of augmenting the interface. It stomped all over it, adding extra features and removing some that worked perfectly. Overall, it continued to create a mess of the interface despite numerous iterations trying to correct itself.
As you can see below the UI got notably WORSE as Qwen 3 Coder tried to implement my request.

Not only did the Qwen 3 Coder bot take nearly an hour to craft these updates, it skipped crucial steps such as compiling the code into a usable format or recording a version change for the interface code edits. These are fundamental tasks that many bots are able to handle without an issue, completely outside of the interface implementation.
OpenAI GPT 5.6 Sol Codex
Giving the same exact prompt to a GPT 5.6 cloud-based model yielded a far superior version in less than 15 minutes. It also compiled the code as needed for deployment and set the release version for the updated module properly. It executed perfectly on the first iteration.

Open AI Models vs. Leading Edge AI Models Comparison
As you can see from the results above, the comparison between the latest open AI model from Qwen and the leading edge AI model from OpenAI output, is not even close. The leading edge AI model completed the task quickly and delivered the results in a format that is ready for deployment. It came out perfectly on the first attempt.
The open model on the other hand took six times longer and didn’t even get close on the first attempt. The image above is after the third or fourth turn where I continued to ask the AI to fix this, change that, just do this one last thing and we’ll get there. It never got there. It created bigger and bigger messes that eventually ended up in all of the iterations being scrapped. All told I spent nearly two hours of my time in addition to several hours of AI processing time to eventually end up with ZERO progress.
I would like to say that open models that use less power and take more time are worth exploring, but they simply are not there yet.
The Big Picture : China vs. US AI Models
It also begs the question of where are we in the “AI Arms Race”. I keep hearing things like, “if we don’t keep moving forward, China’s gonna catch up”. From my limited experience, I don’t see it. They are not even close to the current public models OpenAI , Anthropic, and Grok are putting out today.
The open models, such as Qwen 3 used here, are close to the cutting-edge technology that China is able to create. Based on the results here with the coding model and on my creative artwork experiments using open models they are at least two years behind current US models. Think of how far back two years is in the AI race. Two years ago I had much the same experience with GPT 3 and 4o models as I had with Qwen 3 this week. The models created MORE work for me, not less. They were not helpful in day-to-day operations with accuracy mattered.
That said, I think China will continue to close the gap slowly. However that is going to be a challenge for them with current foundational models from OpenAI and Anthropic pushing forward at a faster and faster pace. They have the advantage of using in-house models that are far more advanced than the stuff we have access to today. They are certainly using these newer in-house models to self-train the AI. The AI training AI they keep hinting at…. I am certain it is already happening behind closed doors.
Who knows where this leads in the future, but for today, the “we need to keep ahead of China” mantra is just an excuse to continue to push ahead at breakneck speed with near-zero accountability. That is a topic for another post.