Under the same computing power, the improvement in LLM intelligence and the improvement in the upper limit of LLM intelligence are equally astonishing. At least in coding, the best local models that can run smoothly on a DGX Spark are now less than one year behind the strongest SOTA models in capability (I measured around 60 tok/s).
I tried it yesterday. The experience seems smoother than on Windows. I hope Computer Use will be added soon.
Also, when will OpenAI fix the issue where Xhigh and Ultra are both translated as “极高” in Chinese? It’s been this way for quite some time. As far as I know, the proportion of Chinese employees at these Silicon Valley AI companies is quite high.
You have "send feedback" button in the app. Something tells me that it's much better place to report issues with the app rather than comment section on unrelated website.
I think their backend data probably shows that very few Chinese users are using Codex through the official way, and the number of those who have switched to the Chinese interface is even smaller, so there’s not much incentive to fix this issue.
> On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.
That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.
It’s always been the case, it’s more the anomaly that LLMs work at comparable speeds on M series because almost all other ML runs way faster on Nvidia cards.
Right? Do you also type on a Mac Studio without a keyboard plugged in? Like tap the ethernet port 3 times in a row then this sequence of sticking your fingers into TB5 ports? I mean, it’s clear you stick your RTX into a computer.
reply