546
Points
354
Comments
apitman
Author

Top Comments

jjcmAug 6
China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.

What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.

d2pAug 6
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

I have screenshots of both. The description above the chart is the same in boh cases:

> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

What happened? How can the scores change so much in a few seconds?

eliAug 6
I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth.

I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.

onomojoAug 6
Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
embedding-shapeAug 6
Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?
seizethecheeseAug 6
Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence

Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol

Source: http://pellmell.ai/leaderboard.

This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.

theropostAug 6
Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
petercooperAug 6
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
Visit the Original Link

Read the full content on artificialanalysis.ai

Source
artificialanalysis.ai
Author
apitman
Posted
August 6, 2026 at 06:44 PM


More Top Stories

huggingface.co Aug 14
Qwen 3.8 27B
1426790 commentsby erdaltoprak
Details
blog.plover.com Aug 14
Seven books I keep close because I love them
308135 commentsby surprisetalk
Details
mixedbread.com Aug 14
Introducing Toast 1
6820 commentsby mplappert
Details
z.ai Aug 14
GLM-5.3: Frontier coding with emergent cyber capabilities
896460 commentsby pella
Details
rustdesk.com Aug 14
RustDesk now supports true unattended remote access on Wayland
293 commentsby rustdesk
Details
blog.google Aug 14
Google is making private AI practical with homomorphic encryption
497288 commentsby u1hcw9nx
Details
👋 Need help with code?