@WandererUber
I've been meaning to ask, which AI model do you think is currently the best?

I sometimes run tasks in parallel to get feedback, but my results so far are... not what I expected
@hazlin hard to tell what's going wrong without any more context.
I've not tried many of the latest because I got Claude this month and cba.
Grok is currently best value for money I think, even cheaper than GLM, Kimi is knockoff Fable.

Since I haven't tried them I don't know which of them don't talk like a serial killer (Claude) so I can't help further.

https://artificialanalysis.ai
@WandererUber Honestly, I just wondered which one you subjectively liked the most xD

Like the link you've got there. Most of the leader boards have a similar order.

But, when I run parallel tasks for, coding, physics, and electrical circuits.

I get:

But, GLM5.2 has a very high query failure rate.

kimi 2.6 and 2.7 often produce trash results (haven't tried 3 yet)

GPT and Grok have the worst privacy policies so, I honestly haven't used them much.

And, Opus will provide detailed and thorough explanations... which are often lies lol! It is a good liar. (I only tried Fable5 a little before the nurf, it is like setting money on fire, but I didn't get the sense that it was performing any better than Opus)

MiniMax-M3 is okay, it is fast, but sloppy.

Nemotron 3 Ultra is just bad, other lists have it a lot higher... for reasons I couldn't guess.

Which... brings me to my confusion. Deepseek V4 Pro, consistently beats all of these, even Opus 4.8 reasoning in these tasks I run every so often.
@hazlin >Grok has the worst privacy
hm okay, idk. See pic. I thought it was quite ok. And the Chinese models also don't really care, except if you pay for ZDR.
>Opus lies
yeah I hate the Claude models. It talks like a serial killer buttering you up. Never a straight answer with this guy.
>Fable not better than Opus
The current cope is that "it can only shine if you do real frontier work". But I largely feel the same. It's better in planning I think (placebo maybe) because that needs a longer horizon. Not too deep into it.

>Deepseek V4 Pro, consistently beats all of these
interesting. I quite like flash because it is fast, but it definitely didn't suggest the same smart fixes that Claude does. Lain's GLM trace is floating around here somewhere also, that thing debugged some insane chain of bugs in an x86 emulator for him. I seriously doubt DSV4 Flash could do that. Pro, maybe. Smaller size than Kimi3 though, by a lot.
Probably best to save some of the convos where you find DS4P the best, and use that challenge for a benchmark suite. Maybe it was luck of the draw, or maybe you can reproduce it.
I haven't done it myself though.
Sign in to participate in the conversation
Game Liberty Mastodon

Mainly gaming/nerd instance for people who value free speech. Everyone is welcome.