Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.

I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.

Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.

Canceling my Anthropic Max sub when this ships.

 help



yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).

Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.


Could you share more about 5x/20x? I missed that

20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number.

OpenAI is 20x on both limits


Is this official or based on people's reports?

> Weekly seems to be around 10x

Actually no. 5x and 20x have same weekly usage across all models. Just ask their chatbot.

https://x.com/beydogan_/status/2095293596198957418


it's clearly wrong, think it's realistically closer to 1.7x

I still use Opus 4.8 for a lot of tasks because I can't stand the way it talks.

> Canceling my Anthropic Max sub when this ships.

At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.


Especially with these big models chewing up limits fast, I feel good about simultaneously having an Anthropic sub, an OpenAI sub, and an Opencode balance. The models also seem to catch things when code reviewing each other that they don't always catch when a new instance of the same model does a review.

Yeah I have a main $100 sub, a bunch of $20 subs and sometimes another simultaneous $100 when a really major model happens to drop. For the most part it seems better to have multiple subscriptions than a single $200-$300 one to stay more in touch with state of the art and get a feel for what's good at what.

Sol easily outperforms Fable on every task I've tried it on.

That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?

For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.

Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.


You’re thinking long horizon tasks. I agree that Sol is great at it. I don’t think it’s smarter in quick win tasks that are still difficult.

This is where intelligence is not one of a kind, these systems have different pros and cons.

I use Sol as an architect and fable as a brilliant single task solver.


I have both set up with full access to all the repos at the company, infrastructure, deployment pipelines, etc.

I can tell Sol, "Hey we need to update this core database schema to handle this new use case" and it will masterfully handle the update, version the API, roll the consumers over, including versioning the Kafka schemas, deploying things in sequence, watching the deployments to make sure the new services act actually active before cutting over consumers, exercising the website and mobile apps in staging environments before releasing to production, etc.

Fable just falls over on long horizon tasks, it does partial implementations, it cuts corners, it gives up, it doesn't verify it's work, it loses track of what it's doing, etc.

It's fine for specific well scoped tasks but can't take high level guidance for complex updates.


I can't speak for others but I have a feeling you're in the very small minority with this take.

You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.


I dont think that feeling is entirely useful.

Because Claude doesn't allow third party harnesses on their subscriptions I doubt the majority of signals you're getting are actually that significant on pure model quality.

I suspect you're right on Sol not outperforming Fable; but i've not used Fable that much.

---

But, fwiw, in my custom harness between Sol & Opus 4.8 - then Sol wins by a ridiculous margin as Opus keeps claiming slightly wrong things with certainty much more.


This is not saying much. Opus 4.8 is ancient history.

what i found to work well with me is Fable for design / ideas and Sol for implementation. Codex models just tend to be more attentive and follow through instructions. Whereby claude models are weaker on this area (they tend to cut corners).

Claude models tend to cut corners during design/ideation too. It becomes especially visible once you pair Sol as advisor to Fable. Sol will start going crazy - "hey, you said this, and it's actually false, i checked that", or "you need a hash chained, triple encrypted, secure-enclave backed storage for this".

But I actually prefer it this way. Sol is a master of overengineering and being overly scrupulous, so Fable balances this out, and I can always say "don't listen to Sol's advisory" about 50% of the time.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: