AI MODEL RANKINGS DON’T HELP YOU CHOOSE
On my computer, a model that sits way below the giants in the general rankings beats models five times bigger. Not now and then, always. It’s Qwen 3.5, the one with 122 billion parameters and 10B active, and it runs here in New York, on my own machine.
In August the ranking says something else entirely. Claude Opus 5 on top with 63 points. ChatGPT behind it. Then the Chinese Kimi K3 at 60. Europe shows up in 21st place, with the French Mistral. Real numbers, measured properly.
I built a comparison system that simulates my tasks with the various models. Read files, search, call tools, reorder: what my local agent actually has to do. Sure, you need a computer with at least 256GB of VRAM. Then you find out it does certain tasks better than gigantic models, at zero cost.
The rankings will keep coming out. And they’ll keep answering a question that isn’t yours. What do you think?
#ArtificialDecisions #MCC #LocalAI #ModelRanking #OpenSource
