Most of the random comments you read on HN and reddit about how nice/bad various LLMs are, is basically based on the commentator's "vibe" about it, and almost nothing is grounded in evidence or actual usage. Don't read too much into it, want to know how good a model is? Run it with your own non-public benchmark, basically the only way to get proper answers you can somewhat rely on, everything else is manipulated, misunderstood or over-relied on.
I thought HN was different. And yeah, wherever I go, my timeline is full of Opus is so bad today and I will switch from Fable to 5.6 Sol, it's 1.5x better and vice versa.
Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.
I am very curious about how the "threat" of local inference and open-weights-models will change the trajectory of the model labs. The air will become pretty thin
I noticed that Claude's reasoning summaries show a phenomenon when I'm using the model in German (on claude.ai) - it mixes English grammar with German vocabulary!
"Evaluating Spülenposition gegen Wasseranschlussabstand" and "Analysierend die Platzierungskonflikte und Rohrleitungszwänge klären" are examples of generated summaries.
I wish someone could explain this - because I see no way that sentences like these would show up in training data. But maybe the RL for thinking has some quirks?
great point, we haven't really spent too much thoughts about that.. changing the logos to companies we know well makes more sense, we'll do that to save us trouble
that gave the best results so far - tweaking the system prompt to make sure the agent respects the project's ui system, is aware of dark mode and responsiveness, etc.