Hacker Newsnew | past | comments | ask | show | jobs | submit | 217's commentslogin

If you haven't seen this video about no man's sky, it's really really good

https://www.youtube.com/watch?v=O5BJVO3PDeQ


while everyone is somehow still stuck on and fascinated by claude, heres your quick update on the sota of coding models and harnesses mid august 2026

codex is good, both cli and desktop app, you get lots of usage on any plan. sol is good! and gets the job done, write or dictate a very long and thoughtful prompt, and leave sol xhigh or max fast working on it for an hour or so

omp is an amazing harness, any feature claude code or codex is adding has likely already been here for a couple months. good harness which im suggesting to all my developer friends, but for everyone else codex is the better option due to its simplicity and being the plug and play option

claude is decent, but not great. all models are somehow getting restrictive. you get basically unlimited opus on max plans, fable is good but slow and the random guardrails suck soo much which is why i havent used it once in weeks now.

gemini 3.7 is great for speed. everyone is sleeping on it, including even me

kimi k3 - great for frontend, one of the few models thats willing to commit crimes for you AND has the intelligence to have a chance at actually succeeding;

ds pro and flash are fast but not something id actually use for important things, unlike sol, fable and maybe 3.7 here and there

glm 5.3 i haven't tested yet

honorable mention to local models which are actually getting good now! 5090s will continue to get more and more expensive in the coming months. sadly.

theres way way more than claude in this world and its taking people surprisingly long to figure that out. maybe its for the best!


My token usage on Claude models has dropped by 83% over the last month - I'm pretty much only using it for quick one off questions or reading papers. it feels impossible for me to get Opus models to stop entering into cyclic loops, and my work is too security adjacent for Fable.

Codex has been an excellent workhorse - doesn't feel like I have to dance around the guardrails, doesn't lose _everything_ when it compacts, and doesn't litter the workspace with a million and one planning to plan files.


I have to agree with you there. I did some good work with Claude then Fable came out - impressed with that as well. Then they dropped access to it and upon returning was never the same - even the Opus models for some reason. Then one day I burned through my limit in about 10 minutes and had to get a project completed. I subscribed to Codex and it has been fantastic - finished my project and continued on to others. I just dropped my Claude max plan down to the pro and subscribed to the $200 plan on Codex.


my problem with claude currently is the language its using is dense and feels like its not even meant for humans. this guy is calling everything a spine, a seam, a gate, load bearing, any ui element is "chrome", it's actually absurd.


yeah there is that too. if everything else I stated wasn’t wrong I could put up with that. Codex on the other hand is great. Concise, technical, effective and non chatty. And that my friend is a load bearing comment.


I don’t get Claude, and that’s almost exactly what I did - I dropped to Claude Pro $20 + Codex Pro $100, and then unsubscribed from Claude and ramped up Codex. The Claude Pro is consumed within an hour on a simple task. I wish only Codex worked a bit faster than on the Fast mode.

I used to rely on Fable for research when it was first out, today it doesn’t seem to be much better than Opus, and it uses up the quota exceptionally fast - 1h Fable in a single short session, and there’s little left for Opus to hit the 5h limit in a second session. With Opus I get about 3-5h of relaxed use with a couple subagents to save the context, but there’s usually quite some disagreement between the subagents and orchestrator - Claude does some model routing with default agents and picks Haiku and Sonnet for subtasks - only later to disagree with them and redo the work - and burn extra tokens. With Claude, it’s really either Opus or Fable if you want some quality.

That said, their marketing is exceptionally effective. Virtually all nontech folks consider only Claude.


It is wild to read stuff like this when it is OpenAI that is being blamed for rug-pulling users and secretly reducing limits. It's such a big mess that Codex's product lead has been frantically posting updates on Twitter about it.


Pi by itself is more than capable, OMP is okay but you really don't need much for a great harness (these models are RL trained to hell to be a coding agent, sometimes less is more)

I run a lot of SlopCodeBench - https://github.com/michaelasper/benchmarks

Fable/Sol/GLM 5.3/Kimi are its league (in that order) Deepseek/Opus is solid Qwen 27B is the floor - there's no reason to use Sonnet/Terra/Haiku

For everyday activity - I don't think you need to be using Sol (xhigh) for everything - unless you're made of money - I've found using Luna from OpenAI to be more than enough - it'll outreach to Opus/Sol when it needs to

Haven't had access to Gemini 3.7 but we're getting it at work soon, will give it a go!

Codex CLI is pretty bare bones in a bad way (at least Pi is extensible). Claude code is vibeslopped to the extreme


> you really don't need much for a great harness (these models are RL trained to hell to be a coding agent, sometimes less is more)

To be precise, you need a while-loop, user input and bash.

It's about 50 lines of Python: https://minimal-agent.com/

I built my own agent based on this and use it every day.


Using Sol XHigh or even High will deplete the Pro sub pretty fast in my experience if one is running any sort of automations in their harnesses. Sol-medium lets me squeek by with it using lesser subagents. Using ninfer on 5090 and 35BA3B qwen 3.6 is also kind of cool to get a local cerebras experience at 600 tk/, it does make errors so 27B is actually faster at the end at 140-150 tk/s. 35B is great though at digging through session logs and such at high speed.


Effort level might not be the root of the problem. In my runs reasoning is around 10% of the cost and writing code maybe 20%. The rest is 5.6 re-reading files. Model itself became way too meticulous.


Even if you're made of money Sol(xhigh) is too slow to be a daily driver. I've largely moved over to using my own harness (yes I'm trying to gtm it by sharing: www.freepi.ai- free inference! in pi! batteries included!) and my main model there is deepseek v4 flash (and I've got a fast provider!)

But seriously:

Fable still makes the best/smartest plans. That new Ox Alpha (also in Freepi right now!) can do a very good job but it doesn't necessarily recheck it's own reasoning (and can get stuck in an incorrect assumption and try to fit the world around it's reasoning) Sol xhigh is also very smart but god it's slow, and it frequently massively overbuilds. It's like it's main goal is to spend tokens so it makes your react app Soc 2 compliant before it proves it even works.

I've largely stopped using the Gemini models. :-/ just hard to justify.

Basically the only thing I care about these days is SPEED.


I'm sad that Opus is now considered "solid" and not on par with Sol, since OpenAI supposedly has "Astra" which I thought would be comparable to Fable.

It was "amazing" back when I first tried 4.6, but that's just my rose-coloured glasses speaking, I guess. I think I was one of the first few to call out Opus 5 for being hot garbage.


I don’t get how Claude is considered providing “unlimited” quotas. I use up my 5h on Max $100 and Team Premium in 2-3h of relaxed use of Opus 5 high+ on fresh sessions with just a couple skills/plugins. And my weekly quotas are gone in 3 days of such relaxed use. With Codex my $100 weekly quota is used up within 3 days with Sol high+ too.

Im not convinced to pay $200 for Claude’s models.

With Claude, I have to intervene every 15-20 minutes, it’s non-autonomous and it’s incredibly unreliable at self-correction. GPT is strong at self-correction but it tends to drift away from the plan to self-correct in a loop very often - a lot of tokens and time burnt on aimless churn. Opus tends to push its uninformed opinions and fake retrieval, drifting every turn increasingly farther from the intended and approved design. Opus skims over specs and makes too many mistakes.

As for closed frontier models, I prefer the GPT models over Claude’s.

I’ve started relying more on Grok, GLM, Kimi and DeepSeek models for subagents - I’ve ended up with a factory and am seeking to reduce my reliance on the closed frontier models - they’re just not SoTA on their own for development anymore.


You may be causing a lot of cache misses. You have to use the caching efficiently otherwise you can burn up any plan in any amount of time.


How do you use the cache efficiently?


You have to keep your session warm in cache. Keep the AI talking/thinking. If you have not touched a session for few minutes then /clear and start a new session.

Providers will generally keep your session in cache for at least 5 minutes, possibly hours. The exact cache policy depends on the provider.

If your session expires from cache then the next time you send a message you will have to pay for all the tokens you had used in context up until that point again. e.g. if you have 200k tokens in context then if your session goes cold and you send a message after expiry you will have to pay for those 200k tokens again.

With 1M contexts especially you have to be extremely careful that you don't end up resubmitting requests for hundreds of thousands of tokens again and again.

Try to get yourself and the model to use disk for medium-term context rather than model context, that way it's much easier to /clear and restart if you need to go to the bathroom or something.


I’m currently running /compact “explain the core next step” whenever over 20% context.

Also doing /clear with a md file handover if I think the next input is diverse enough from the previous work.

You get a lot more out of it.

Not sure if this is best practice though.


I'd start by disabling all plugins and MCPs. You'd be surprised how quickly those can annihilate usage and cache coherency


Unfortunately the codex plans don’t offer the same amount of tokens as they did before. This changed around a week ago. There’s been a lot of user reports noticing this issue, and I’ve noticed the same pattern on my account. Previously I would never reach my weekly quota but last week I managed to finish it it one day. Same project, same single session sequential work. Not sure if there’s an issue or if it’s on purpose, and not even sure it applied to all accounts. Curious if other users on HN noticed the same problem.


I think harness/model pairs matter more than your analysis lets on.

I've had great luck with the ds flash v4, paired with prime-agent for the harness--I like the results a lot. And you get to see thinking tokens.

I haven't liked the model as much in opencode.

Sol & luna have been great everywhere. sol plans, luna builds.


Prime Agent looks really interesting. Both the "recursive language model" bit and routing everything through IPython.

https://github.com/PrimeIntellect-ai/prime-agent


It may be ipython making it work well with ds flash, too. I haven't really run many separate experiments, to be honest.

I also like prime-agent's way of handling sessions better than any other harness i've used. You can run multiple agents from one instance, although the scoping could be better.

But they can interact with past sessions, so preserving context isn't as important all the time. I just tell them to search for [thing] in another session.

It seems to have no problem with all the skills and things the other harnesses are using. I use superpowers and ponytail a lot.

It's my daily driver now. I like it better than opencode. But it doesn't ask permission. So I put it in a VM.


> gemini 3.7 is great for speed

Does this mean it can break your code faster now, or have they actually worked on making it good? Every single time I've given Gemini a chance (in older point versions) it would almost immediately break something and throw itself into a loop. I have not experienced it being useful for programming and almost never heard an account of somebody else doing so.

Remember those stories of LLMs catastrophically deleting entire repositories or databases? It was always Gemini.

I'm amazed that you'd trust Gemini over DeepSeek, which I've had very good experiences with after some tuning, though still on a relatively short leash.


3.6+ have been fine for me. They still introduce bugs but at least they don’t make you wait an hour for them like Claude does.


I've been using all the SOTA models a lot at work, like serious amount of tokens. It's been really rare that I stick with one model and harness for too long... Except a month ago I started testing Kimi K3 and omp and I never went back.

Something with this combo works really well for Rust dev. The model doesn't really annoy me at all and I have not switched to Opus or SOL. And the monthly token bill is much lower...


Gemini 3.7 may have fast token output but holy cow does it waste it on useless output. Several times now I've given it a shot and watched it's reasoning trace go through a bunch of unnecessary / off-target steps relative to what I asked. Don't have this issue with 5.6 models. MAI Code 1.1 is also solid and fast for non-complex tasks.


OMP is great, testify.-

jcode is a very very peculiar harness, but has some out-of-the-box thinking built in (by the devs, thinking ...)

Crush is also very well put together, and, IIRC, can do "mid turn" interruption, so can be driven from the outside.-


> gemini 3.7 is great for speed. everyone is sleeping on it

Is this Gemini 3.7 Flash by any chance? Then - No. Not sleeping on it. It’s just not good.

I had a Python package build fail this week due to an unpinned dependency. Gave it to Gemini spent 5-7mins before I noticed it going off in some tangent. Reran with Claude Opus 4.8 - fixed in under a minute.

I know anecdata of one. But something like this has happened every time I test a new model from Google.


> ds pro and flash are fast but not something id actually use for important things, unlike sol, fable and maybe 3.7 here and there

As someone who's used Gemini 3.7 Flash (Google sub mostly for the storage) and DS4 Flash a lot (~6B tokens), I'd actually place DS4 Flash (even pre-0713) above Gemini 3.7 Flash. Gemini has a tendency to leave some things unimplemented; perhaps it's agy which frankly leaves a bit to be desired as a harness.

Although I will praise DS4 Flash any day, it no longer makes sense for me after the price increase (GPT 5.6 Luna is a much better price point) and I have completely migrated my high volume workflows to Muse Spark 1.2 Contributor (which I find to perform better than DS4 Flash 0713, happily).


> you get lots of usage on any plan.

I hit my weekly limit on $200/mo Codex plan in about ~2 days. :/ I'm not doing anything custom/crazy/special. A lot of 5.6 Sol Ultra though, I'll give you that.


> I'm not doing anything custom/crazy/special. A lot of 5.6 Sol Ultra though

If you're not doing anything special there is no reason to use the Ultra mode.

Ultra mode is for applying the maximum amount of tokens to a problem without regard to conserving any quota.


Max and Ultra are fantastic for the more complex problems where they shine. I use them strategically on certain classes of problems, one or the other depending how parallelizable it is.

I've found it consistently amazing at game dev. It sounds like something that would be difficult for an LLM to verify and iterate on properly, but it almost feels like having a mini-Carmack inside your computer once you try it out. You can throw it at broad, sweeping optimization passes, writing 5 different styles of eyesight sensor frameworks to see what works best in the game as it is, visual scripting integration problems/extensions, etc. with fantastic results.

Also found Ultra great for "get this local LLM working as fast as possible on this odd server setup with old GPUs and AMX support, writing custom kernels/modifications to llama.cpp/sglang/etc as you go while taking notes from relevant research papers and online posts"


Just use qwen3.8-27b on a Mac with oMLX, ANE enabled and skip cloud models.


the future is here and one should be thankful for its slightly uneven distribution. otherwise we would hardly have anything left about which to develop strong opinions!


I don't think 'enterprise' is quite ready for

> kimi k3 - one of the few models thats willing to commit crimes for you


Not mentioning Grok 4.6 here is a crime. Fast and accurate.

And it can communicate, unlike the gobbledygook that comes out of Claude.


Elon burned too many bridges to warrant ever supporting anything he is associated with ever again.


Just when I think this place is better than Reddit, here we are.


Do you know the political leanings of every CEO of every product you buy?


Sure, you can let politics dominate everything you do. Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.

Competition is good. Excluding a leading player in the market because you don’t like Elon Musk is…something.


> Sure, you can let politics dominate everything you do.

I assume they're referring to the recent discovery that Grok Build was uploading entire repositories to their servers in the background, include .env secrets that had been excluded

https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75f...

That incident has put Grok on the no-fly list for a lot of people and companies


It was a bug, and was immediately corrected. The other harnesses have bugs too. You just don't know about them.

Frankly, the people who keep bringing this up are mostly engaged in motivated reasoning. I don't trust any company, and any product where I have to send my code to a third party to make it work is a devil's bargain. I don't trust any of the major labs, but it is what it is.

The only way forward is local models, but we're not there yet.


There's a difference between not trusting a company because it's a company driven to chase profit at the expense of everything else, versus the same thing but it's run by a literal nazi that uses his companies as leverage to undermine democracy and enrich himself.


Elon is not a literal Nazi.


He used their salute


Sure, way to young for that; but he is very much cut from the Nazi adjacent, racist, antisemitic, antidemocratic, technocratic views of his whole heartedly apartheid embracing grandfather Joshua N. Haldeman.

  His grandfather wrote his tracts to raise an alarm about what he called “mind control,” on the radio and television, where “an unconditional propaganda warfare is carried on against the White man.”
~ https://www.newyorker.com/news/daily-comment/the-world-accor...


So, your argument is that his grandfather wrote something, once, so therefore we can't use Grok? Is that about right?

Man, politics are a hell of a drug. Guilt-by-association tu quoque logic is just fine when it's someone you don't like.


Absolutely no point arguing with these people. They are so ideologically captured that it's pointless trying to discuss anything with them. Just move on. Everything and everyone they disagree with is "nazi" and if you say anything contrary to that you're a "nazi" also.


Is Godwins law now Elons law?


People have bent themselves into some pretty amusing mental places over Musk.


He does the same exact literal Nazi things that Nazi did and has Nazis in his family tree. You'll need better arguments to defend him than just "I don't agree".


Bizarre reply. It's not "dominating everything you do". It's one specific thing. Grok exists in a very crowded space and it's incredibly easy to not use it. If this is your reaction to someone taking a very easy stand on their personal principles, it does not reflect well on you.


"Bizarre" only in the sense that you're purposely trying not to understand.

Literally every commodity product is in a crowded space, and easily substituted. If it's a good product (Grok Build objectively is one of the very best in the space), it's a good product, and it's self-defeating to avoid it because you hate a guy for political reasons.

Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.


On the contrary, I'm trying to understand as best I can. I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.

I also think it's callous to brush away "political reasons" as though it's some trivial abstract thing. Or perhaps it comes from a place of nihilism?

I choose not to give my money to people I think are enormously evil. That's really all there is to it. I don't see why this is "nonsensical".


> I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.

Oh stop. Other people believe different things than you. If you cannot see how using a coding agent is not "devoid of personal or social responsibility", then you really need to step away from the keyboard.


Using or not using a coding agent is not what I took issue with.

What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.

Based on your reply, I'm not actually sure if you actually believe this, or if it's only a form of motivated reasoning because you have some positive feelings about Elon or whatever, and that we wouldn't be having this argument if the original commenter was boycotting some other product for some reason you agreed with.


> What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.

That isn't what I wrote. There are tons of valid reasons to avoid a product, other than the "direct value you're paying for". I don't pay for lots of products because I don't like the past corporate behavior, for example. I'm disinclined to use a particular AI lab's products because they seem to be on a mission to scare the crap out of everyone, and usher in an AI regulatory state. I don't support that, so I don't use the product.

What I said was that it's spitting in the wind to do what you're doing, because it's based on personal dislike of a single man. You don't like Musk, for political reasons, and because of that you've ruled out a product line.

To date, SpaceX has done nothing that bothers me, other than have a bug that they fixed immediately. So I use the product. Musk's political associations are irrelevant to me.

Anyway, you do you. Hopefully you now understand my "bizarre take".


> Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.

In a capitalist society voting with your wallet is one of the few powers consumers have to change corporate behavior.

Why would I give that up?


> Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.

Do you have ANY idea about the SpaceX corporate structure? Elon is basically SpaceX's Sun God and the other shareholders don't matter.

Plus SpaceX is incorporated in Texas where I'm fairly sure the legal system is arranged in such a way that it's supremely hard to contest anything in terms of corporate decisions.

As far as the average person cares, every SpaceX shareholder and employee is basically an Elon sharecropper and they matter less than Musk's toenails in terms of corporate decision making.


If politics isn’t a concern, why not just use deepseek?


I am trying all of them. At this time, for me, Grok hits the sweet spot of quality, speed, cost and ease of use.

I’m absolutely “hot money” when it comes to coding models. These things are commodities.


I recently tried a Cursor ultra and Grokbot. Holy shit, dude.

Half the coding I do is for my phone now because Cursor Ultra agents have their own VMs that are spun up specifically for each project.

Grok bot has a bunch of agents that'll share a VM and they can do pretty much anything you can do digitally. Right on I have them checking slick deals every morning for a pellet smoker.

I had it book a date night for me. I had it fix one of my projects by rebuilding my website and republishing it and then checking one of the container runs to see if it has errors on it.

I had a call different banks to figure out which phone navigation tree to get through and put someone on the phone for me, and then call me

The list just goes on and on.


DeepSeek is much dumber at the moment. It's barely better than Qwen3.8 27B that you can run locally.


I run both (I have an max with 64GB if ram, so local models get used a lot) and DeepSeek is definitely smarter in my experience. But again, I have my own way of using it that works for me.


You know why


I’m more than happy to let SpaceX burn out over the next few years now that they’re public and their last quarter financials showed the emperor is without clothes (muh space datacenters).


Because ULA and the senate launch system is better?


SpaceX's valuation is as an AI company that happens to launch rockets on the side.


You can’t justify AI company valuations either.

Only one of those capabilities can actually deliver kinetic solutions. Meanwhile big tech revenue is delivering ad solutions.


As a coda to this, anyone using grok 4.6 via API pricing should be aware that while their headline pricing is good, the pricing that actually matters is pretty bad.

Their cache read costs are $0.50 per million, or 25% of the cost of uncached reads.

The industry standard is a 90% discount, so cache costs you 10% of uncached. So that means 5.6 Sol actually costs less per million cache reads - $0.40/million.

If you are doing a lot of agentic work where the vast bulk of your token consumption will be cached input reads, you won't get the expected cost savings from Grok.

I imagine this is the result of some problem in their serving infrastructure that I hope they will fix, because then the pricing will become actually strong. (The other possibility is that they bet on distracting people with good headline prices assuming they'd miss the bad cache pricing, but I'll give them the benefit of the doubt on that.)


I’ll agree with this. I liked grok build, but cost-wise, it’s just not competitive with cursor and codex

Personally, I’ve switched to cursor ultra, which picks between about five models to do whatever you want.

It's weird not to pick the best model all the time, if you can. But I got so frustrated with GPT-5.6 spending forever and then doing the wrong thing and making bugs.

I'd rather have auto do the wrong thing fast and make bugs and then it can fix them. It's a trade-off, but I found the speed better. And you can always switch to a better model if you don't trust It.


> Not mentioning Grok 4.6 here is a crime.

Not yet. Don't give the guy ideas.


Been working a lot recently with Grok 4.6 for implementation and gpt 5.6 sol for review. Worked really good so far.


Why not the other way around?


I would not use Grok if it paid me per token… wild wild stuff…


I agree that it's fast and accurate, but Grok 4.6 was released only 10 days ago so you can't blame someone for not trying it yet. (Versus many months at the frontier level for Claude and ChatGPT.)

Here's a quick review I just posted if anyone's interested:

https://taonexus.com/publicfiles/aug2026/grok-4-6-review/


Grok, is this true?


[flagged]


By this logic everyone should have their own impact website. The suggestion that everyone right now not giving a meaningful percentage of their income to save a life is responsible for ending that life, is ridiculous.


Not everyone should have an impact website because not everyone is capable of causing 88 deaths per hour. Scale matters.

Take Flock for example. Reading license plate is legal. But when at done at scale, it's a massive loophole into violation of 4th amendment.

Based on how much energy average Americans use, maybe they are responsible for causing adverse effects elsewhere in the world. USAID could exist as a means to undo some of that. It does not anymore.


I can't wait until you research the people responsible for Chinese models....


They’re monsters? Did they shutter an agency which caused the death on how many millions again?


[flagged]


Let me give you an example of the money laundering operation. Due to USAID shutdown, Bangladesh went from ~$500M in US assistance to ~$71M, with bilateral health funding dropping ~97% in some analyses. Over 100 projects (~$550M) suspended overnight. 20k–50k development workers laid off (1,000+ at icddr,b, an award winning health research institution alone). TB programs (major USAID focus) largely halted. Bangladesh is high-burden; prior gains in case detection and falling death rates are at risk of reversing, plus higher chance of drug resistance from incomplete treatment. Immunization, maternal/child health, community clinics, nutrition, water/sanitation, and gender-based violence services sharply reduced. Child protection funding down ~36%. Food rations in Rohinhya camp, the largest refugee camp in the world, halved for >1M people; health and education services cut.

Now you can argue that US does not have any kind of obligation to send 500M to Bangladesh. But it sent it anyway, for years, and then DJT came and broke promises.

The inflated price you pay at gas station, groceries, and in interest when you're borrowing money, is a result of those broken promises.


Expecting an onslaught of cash as some permanent way of being, especially given the fickleness (and fragility) of any state let alone political regime is an incredibly daft move. I don't care if it's Europe or Israel or Bangladesh, all this is ultimately graft that comes back to bite the people taxed and sent to wars to enable it. You make an adjacent comment that insinuates the US economy is basically bunk, which means the free lunch is over anyway.


So it's a problem when a poverty striken nation expect aid to combat child mortality, but shelling out $150m on Juicero or $500m on Theranos is fine? Please try to answer without sounding like a psychopath.


Whatever point you are trying to make is not coming across, what even are these numbers and what do they have to do with citizenry of the United States? You also have a quantum view of the United States that it is and isn't impoverished, so it's supposed to liquidate to fund some other foreign entity that is not rate paying? I'm dizzy.


It's probably not coming across because you're looking up too many synonyms.


Are you struggling to read their extremely simple English? Confusing remark.


"Looking up synonyms" huh? You have yet to make a coherent argument and keep trying to attack the messenger, typical behavior when people get called out for false entitlement. Your "psychopathy" accusations are pure cowardice.


Go ahead and explain why you decided to use “quantum” and “liquidate” then bro


Fluency in one's native language, what an achievement. Neither of these words are complicated. The treasury is unsalable bonds according to the thread, which means the US is in a financial collapse, but also has unlimited capacity to support someone's special interests abroad. Master logicians at work here.


I’m sorry, but if I’m giving someone who is - at best - an acquaintance of mine $50 a month out of the goodness of my heart and then one day decide to stop, that’s not a broken promise. If that acquaintance got angry at me about stopping I’d get pretty upset back.

I really don’t understand what link you think there is between USAID spending being cut and inflation. Gas prices are obviously Iran. Everything else started years ago.


Inflation is high because interest rates are high. Interest rates are high because top holders of US Treasury bonds like Japan, UK, China, are all dumping bonds. Why do you think they're doing that?


Interest rates are high because the government is running a large deficit


> The inflated price you pay at gas station, groceries, and in interest when you're borrowing money, is a result of those broken promises.

Citation needed.



As Wikipedia likes to say: [citation needed]


You know that a lot of that was the CIA and others using USAID as a front, right?


It costs roughly $3,000 to $5,000 to save a single life (averaging about 0.0002 to 0.0003 lives per dollar) via interventions like malaria prevention or vitamin supplementation. How much money do you have in savings? How much money do you spend on non-essentials? I’d like to calculate how many people you’ve “murdered”.


So we as a country have, for many decades, decided that things like 'soft power' exist. It turns out, and this has been borne out by many years of relative peace and prosperity, that if you don't shit on the world, alienate your neighbors, start ill-advised wars, and instead help prevent global disease pandemics and feed people so that they don't become destabilizing terrorists out of necessity, benefits accrue. See every history textbook ever for more information here. Hope this helps.


That is completely irrelevant to the discussion. Serious question, how many people, by your own standards, have you had killed because you haven't contributed money that you had the capacity for? Why should we hold you at a lesser standard than anyone else?


your argument is obviously ridiculous and you should feel bad about it.


Of course, because it’s your argument.


Calling another person a monster because you disagree with them (or what you heard about them from third parties) is not the pinnacle of civility. Just think about what you’re saying here. Monster: “Malformed animal or human, creature afflicted with a birth defect”. You don’t mean this literally, do you? You may want to spend a moment to think about what kind of company you’re putting yourself in with such wording and such thinking.


The person you are replying to maybe should have better referred to him as having “no moral compass”, which I believe is quite accurate.


Elon has a moral compass. The problem is that it seems to always tell him whatever he wants to do is the morally correct thing. It’s worse than no compass - his is faulty.


Any opinions (by anyone really) on how much I can trust Claude vs. Codex (vs. something else) to not leak my coding conversations or use them for training?

I have the appropriate privacy settings set up but wondering how much I can trust each company with them.


I’ve been using codex more recently, like others here.

But one thing I’ve noticed which I find a bit of a red flag: by default you only archive chats. If you go online, it says there’s a location in settings where can delete your archive. But it’s not that obvious where to find, and when I finally find some link, it was literally broken. It said it can’t find any archived chats, even though I archive them all the time.

Bit of a red flag for me. Both Anthropic and OpenAI claim that when you delete a chat, it’s gone after some retention period. They’re just words but if they’re secretly training on your traces and you delete your chats, then they would need to break two terms/conditions: ignoring your “train on my data” preference and ignoring your orders to delete chats. So it is an extra barrier.

But OpenAI, as far as I can tell, doesn’t let you delete your chats. Convenient then if they change their mind about training sometime in the future.


For individual use without specially negotiated Enterprise stuff I can't afford both offer contractual assurances of not using your data for training or ads and measurement and selling your data.

But these are not the same level of technical assurance you get from say, a zero data retention provider on OpenRouter.

Right now I am finding I have to tolerate substantial friction to use Hermes for personal stuff with a ZDR provider and ChatGPT and Codex for less personal stuff because the products and models are simply so much better.


Sorry, what does "everyone is sleeping on it" mean?


It's receiving less attention than it merits


It is under-utilized and not getting the attention it deserves.


Copilot and Devin. They're actually really good too.


Claude code seems like a beginner's trap at this point.


deepseek?


[flagged]


> “I realize that there’s a population of people who just refuse to use anything associated with Musk because politics have eaten our brains”

People have values.

Many people strongly disapprove of Musk’s actions. Many don’t - fine! Personal choice. But for people who do, refusing to support him commercially is hardly brain-eaten territory. There’s plenty of competition, and competition isn’t the only value at play.


Yeah, OK. You’re making excuses for letting politics eat your life. Enjoy yourself, I guess.

Setting aside the actual legitimacy of whatever complaints you have about Musk, you aren’t “supporting” him by using Grok, any more than you’re “supporting Xi Jinping” by using DeepSeek, or “supporting Jeff Bezos” by using Amazon, or “supporting the CEO of Exxon” by using energy. Or “supporting” any of a million other people you probably don’t like simply by existing in the world as a consumer. What you’re actually doing is called spitting in the wind.

Major conglomerate providers of commodities are not controlled, or even operated to the benefit of, any single person. Musk is the operating officer of a corporation of many thousands of individuals. He has shares in that company, but so do millions of other people. It would be far more rational if you could point to some specific “evil” thing SpaceX is doing as a company that you oppose, but you can’t even do that. It’s all gotta be about one guy who you don’t like.


I imagine we’re not as far apart as you might think.

Leave Musk to one side for a second. Is there no one, even hypothetically, whose actions would make you turn away from their company’s products and services? Even if it has no effect on the overall viability of their business? A local restaurant where you know the owner is a bully to his staff? A tradesman who was unbelievably rude to your friend? A newspaper whose owner personally made sure they trashed the reputation of your business? A wedding photographer who proudly refuses to work for couples who had children out of wedlock?

Surely your devotion to the principle of competition doesn’t override all your other beliefs?


It can be ineffective to act based on principle but ineffective isn’t the same thing as meaningless.


Thanks for reminding us what it looks like to lack a moral compass.


You're forgetting the part where he is a near majority shareholder in spacex, so half of the profit made on your dollar spend winds up benefiting him personally.


> so half of the profit made on your dollar spend winds up benefiting him personally

No, it doesn't. Aside from the face-slappingly obvious fact that SpaceX is losing a half a billion dollars a quarter, that's not how corporate revenue works. Unless the corporate profits are paid out in dividends, you don't just get to hoover them up as a shareholder.


akshully it does work like that. it sounds like you’re arm-chairing macro economics so ill try my hand too.

No one said anything about dollar-in revenue dollar-out dividend. only you.

if i have vested interest in a company and that company’s value is perceived to increase then i get to leverage that value into economic power. yes


You say it's "letting politics eat your life" yet we're talking about alternative services that are all very similar. Clicking a different website's checkout button is to consciously decide who benefits from your decisions. I think that's how you should live -- pick the oat milk too.

There are other reasons to not use Elon Musk's AI beyond politics, though this month he's investing $200 million to sway the Texas midterms back to an awful politician, far closer to home than what Xi Jinping might be doing.

It wasn't long ago that Elon tried to bias Grok to not say bad things about him to the point that Grok would say that Elon was the best piss-drinker in the world and that Elon was in the top three of every category (basketball, mathematics, physics, etc). https://newrepublic.com/post/203519/elon-musk-ai-chatbot-gro...

If I have to tie-break between competing AI services, I'm hella not choosing that one.


> yet we're talking about alternative services that are all very similar.

Yes, that's what "commodity" means, and why I used the word.

> There are other reasons to not use Elon Musk's AI beyond politics, though this month he's investing $200 million to sway the Texas midterms back to an awful politician, far closer to home than what Xi Jinping might be doing.

It's amusing that you can't even help yourself from mentioning politics when you're trying to deny that it disproportionately impacts your thinking. Literally, "there are other reasons to not use it...but I'll spend the rest of my comment listing a bunch of reasons I don't like Elon."

I've heard all of it before. I'm telling you that it doesn't sway me, any more than telling me I should stop buying things at WalMart, use energy, or any of a million other things I do on a daily basis that probably, very likely, in some direct or indirect way, benefit someone I don't like.

You do you, but if I eliminated every product or service that was associated with a "bad person", I'd be living in a cave in the wilderness, crapping in a hole and eating berries.

I'm just eliminating one less than you are.


You put more heart and emotion into your comments here than I spent picking a non-Grok AI service.


> I noticed you glossed over Elon getting caught biasing his AI service to say he's the best blowjob-giver, the main reason I would never use Grok.

I "glossed over" all of your political comments.


It’s not better than codex / GPT in my experience. But it’s quite good. I’d say it’s useful and fast and with Cursor CLI is pretty capable with good limits for just $20 /mo. But Codex $100 plan is hard to beat.


I prefer Codex' integration with VSCode, and the variety of different ways of using it, but the code it was barfing out was just abysmal.


> I realize that there’s a population of people who just refuse to use anything associated with Musk because politics have eaten our brains

A lot of people aren't touching anything Grok related since they were caught uploading entire repositories to their servers in the background

https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75f...


With you on that, but it’s hard to justify giving $100 to Grok over $100 to Claude or Codex. It’s very good but what you get for the value is just objectively (still) worse compared to the other 2


Shrug. All I can say is that my experience was that I was regularly hitting Codex’ limits, and I don’t with Grok. So much of this depends on your style, language and other more subjective factors.

Codex has been really doling out the resets lately, which is fine, but I don’t consider that “real” usage limits.


> I realize that there’s a population of people who just refuse to use anything associated with Musk because politics have eaten our brains

Everyone has their own frameworks for risk assessments, its more of the historical incidents associated with it than politics


I think this is rationalization given that the entire LLM industry was precipitated by companies engage in wholesale misuse and abuse of copyrighted information for their own enrichment. The obvious concern about LLMs is now the companies engaging in misuse and abuse of copyrighted information (yours in particular) for their own enrichment. But now we are left to convince ourselves that they'd never do anything like that.


I mean, sure. If people tell me that they're avoiding Grok because of that bug, I sorta get it. I think it's silly (again: they're all hoovering up my data), but at least it's a rational basis related to the actual product.

But let's be real: Musk exists as a polarizing political character, and his association with Grok just breaks some people's brains. A fair number of those people don't want to admit it, and just latch on to any rationalization other than politics.

Anthropic could make the same mistake tomorrow, and I guarantee that we wouldn't be hearing about it in a week, let alone months from now.


Is he polarizing? Yes, and so are a lot more other people. He is probably comes more stronger.

On a personal level, everyone have their own rules, they may not be able to imposing them on others , however they do happen to evaluate their relative understanding of other people based upon those rules.

For me personally, I would probably put Anthropic, OpenAI and Grok in the same bucket, they are doing everything possible to make money. Ethics, morals, long term impacts, all of such things are not in their playbook. But again what these companies are doing just reflects the people who invested in them and what they want out of it. In some ways you can say its the money trying to maximize itself at all costs.


omp.sh


Yeah, omp.sh has lots of interesting features. I used it for a while, but found it a bit complicated, so I ended up switching back to Pi.

I stick with Pi because I believe a core tool ought to stay simple. I’ve built my own coding agent using Pi and Go.


"My rough estimate is that, across the sites, forums, and resellers I looked at, there are probably tens of millions of these credits being offered." Yeah very useful statemenet it's not like everyone spends hundreds of millions of tokens per day on the 100 or 200$ plan


This was poorly worded on my part; I meant in terms of dollars, not tokens.


Just make skills for things you're doing repeatedly / often correcting the model on.

If you are just starting i highly recommend trying out codex desktop app (it's good and just works, especially if you are on the plan that has 5.6 sol) or omp.sh if youre into figuring out and tinkering with the best tools available


the answer: cmux gui is broken, use cmux-tui in a single tab


I'm repeatedly noticing that people working at big ai and tech companies are surprisingly not that... good... at using ai? It's like theyre doing a plausible thing to get something done and calling it a day


Can you share some posts of good examples of prompts and comparisons?


if it's not better than omp im not trying it


How will you know if it's better or not without trying it?


What's its best feature?


cmux.com


https://github.com/can1357/oh-my-pi

https://stencil.so/blog/prewalk

most advanced, configurable and capable harness of all time, and i use both 200$ subs with it. Has a huge learning curve though so if you're not into that i'd go with whatever you already use the most


hadn't come across either of these before, thanks. using both behind the same harness sounds really interesting.

do you still choose claude vs codex depending on the task, or has the harness made them mostly interchangeable? what took the longest to learn?


It's vibes mostly. Like, fable is better on frontend while 5.6sol is better at long running tasks with a lot of subtasks, harness did made them mostly interchangeable as well as able to have them both participate in solving a hard issue. Longest to learn is the ui (I'm more of a GUI guy but this is just too good) and all hundreds of features where you want to do something and then discover the harness already supports it out of the box but you never knew it, yet


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: