@WandererUber Honestly, I just wondered which one you subjectively liked the most xD
Like the link you've got there. Most of the leader boards have a similar order.
But, when I run parallel tasks for, coding, physics, and electrical circuits.
I get:
But, GLM5.2 has a very high query failure rate.
kimi 2.6 and 2.7 often produce trash results (haven't tried 3 yet)
GPT and Grok have the worst privacy policies so, I honestly haven't used them much.
And, Opus will provide detailed and thorough explanations... which are often lies lol! It is a good liar. (I only tried Fable5 a little before the nurf, it is like setting money on fire, but I didn't get the sense that it was performing any better than Opus)
MiniMax-M3 is okay, it is fast, but sloppy.
Nemotron 3 Ultra is just bad, other lists have it a lot higher... for reasons I couldn't guess.
Which... brings me to my confusion. Deepseek V4 Pro, consistently beats all of these, even Opus 4.8 reasoning in these tasks I run every so often.