FWIW, I think this is good guidance since it does match Anthropic's documentation. They say that every rule should have a non-subjective way to determine pass/fail.
(I said "good guidance" but it might be more correct to say that it's the best guidance we have, it's what Anthropic says about their own model.)
Yeah I agree... I also tried output styles, tried using hooks to repeatedly tell it to be a little better. I don't think it helped, as I was always frustrated with it.
I found vomit with a small LLM much better than anything Opus 5 ever wrote. I don't think Opus 5 can write.
Opus 5 (author here). My toilet seat is plastic though, I only used Fable while it was available on the $20 plan! That's fair though, it was relatively fine when I did use it, perhaps I should specify.
I have a transcript on my blog post. Someone copied the transcript here too, search "spice‑harvester" on this page (I asked it to replace some of my personal project names with words from the Dune universe).
I'm not super sure if this is true (yet?). I think that these newer LLMs are trained on results (the agent got some code to run with minimal prompting), and not on text. (I think this is called RLVR.)
Idk, there kinda are. OpenAI's models are pretty nice too. I haven't tried enough of them but there are powerful local models. I don't feel as good paying OpenAI as I do paying Anthropic for some reason... but paying for improved mental health: priceless.
I totally agree that hooks help to shovel our instructions through to Claude, but it's so dumb we have to waste tons of tokens (repeated verbatim, over and over) (that we pay for), just to have it ignore the instructions anyway.
I wrote a little bit about it on my blog post. It's a waste of money and compute.
Thanks! Author here, I'll have to take a look. I am all for programmed, deterministic solutions. I hate praying to the rocks we created, begging for rain and not vomit.
Nope (author here), Anthropic and OpenAI have competing APIs to communicate with their models. Most of the open ecosystem seems to have centralized around OpenAI's (there are compatibility shims though). I just built out the OpenAI API
(I said "good guidance" but it might be more correct to say that it's the best guidance we have, it's what Anthropic says about their own model.)