Hacker Newsnew | past | comments | ask | show | jobs | submit | gbnwl's commentslogin

This is a nit but his name is actually Erik Satie not Eric Satre.


Thanks! I was confused, thinking "Is it Satie, or is this a case of Muphry's law?". I wanted to learn more about this Satre composer I never heard of.


Yes. Made a type and after I noticed it was too late to edit my post.


It's a nit. They made a typo.


It seems like the entire thought process you’re trying to sell hinges on the idea that OpenAI reported the attack first. Did you forget that it was actually HuggingFace that reported it first, and OpenAI only stepped forward latter?


That's incorrect. My thought process hinges upon the fact that Hugging face did not report that uncontrolled models escaped confinement autonomously by coordinating with other models via a series of zero days during training. OP is concerned about the autonomous nature of the incident, not whether or not the incident happened. Nor am I contesting that the incident happened.


Every day I wake up and open HN.

“LLM has made legitimate mathematical discoveries” —> Wow the rate of progress is amazing. Highly upvoted.

“LLM does something not good” -> Does everyone else not realize LLMs are just dumb next token predictors? Highly upvoted.

So tired of this discourse and this site.


The rate of progress can be high and they can also be dumb next token predictors. Not sure why that is hard to understand.

These models can do a lot of things but they also can't do a lot of things. In order to use these models effectively you have to understand that they are next token predictors and how that allows it to do what they do.


Are they useful or not? Will they continue changing the world or not? People who choose one way or the other for describing them typically fall on one side or the other in these questions imo. What do you think? Will these next token predictors change the world or not?


They are useful. They will continue to change the world. They are still next token predictors with all the problems that comes with that.

For them to change the world you have to work with them as next token predictors. Ensure that the next token predictor has enough prediction paths to solve the problems you want and so on. Since when they don't they fail spectacularly. These big companies will continue to add new skills to them, so they will continue to get more useful.


In all fairness humans can also be considered next token predictors. It could be said that’s how we communicate with one another today. Presently LLMs lack other things, like physical presence in the world and continuity of input sensory data.


Humans learn to be a next token predictor as a kid when they learn to speak, an LLM cannot learn to be a next token predictor or anything of the sort, we have no clue how you could have an LLM learn human language just based on a thousands conversations with a human.

You don't see how that is very different? For an LLM to be as smart as a human it has to be able to learn like a human. Like you don't evaluate how smart a human is based on how much he knows, you evaluate it based on how fast he learns. And LLM are so bad at learning its ridiculous, they lack that part of the brain that lets humans be smart and learn so fast and easily.


> For an LLM to be as smart as a human it has to be able to learn like a human.

"For a plane to fly as well as a bird it has to be able to flap its wings".

"For a submarine to swim as well as a fish it has to be as light as fish".


> "For a plane to fly as well as a bird it has to be able to flap its wings".

> "For a submarine to swim as well as a fish it has to be as light as fish".

These are false equivalences. The post you're responding to defined intelligence as learning rate. LLMs unequivocally do not learn. You can disagree with OP or agree, but what you have done is out of bounds. You're implicitly claiming that LLMs learn, albeit differently from humans. This is categorically false, unless you count training as some kind of "learning".

That's of course absurd, almost no user of an LLM also trains it. Instead they rely on queries submitted to pre-trained LLMs ("inference"). And don't bother yapping about context windows, it's just not anything like learning.


> The post you're responding to defined intelligence as learning rate.

How is the learning method or rate related to intelligence? LLMs learn during the training, much faster than any human. Yes, they drastically slow down their learning afterwards, but they still can learn a bit from an uploaded document or a web site.

And in the end they may be better at intelligence than any human, depending on the task and time given. I though intelligence is not about history but current performance.

How is this different than my examples, in which a possibility to fly is judged by (learning to) flapping the wings and not by the actual result?

> And don't bother yapping about context windows, it's just not anything like learning.

That's debatable, but I don't see how it's relevant here. You can ignore that part of my reply above, and the argument will remain unchanged.

Nevertheless, I don't understand how this not learning. Without it, no meaningful intelligent task can really be performed. LLM/human must learn the relevant bits from current situation in order to answer meaningfully. Often it requires to actually acquire new knowledge like reading a new piece of text unknown before. Feel free to link to a relevant discussion for me if you find this boring and settled.


Somewhat.

Yes, but not because they are useful.


It can be a token predictor and still tell me exactly how my life will proceed from now until the indefinite future, or be the most intelligent conversational entity you have ever witnessed.

The issue is of course with using the word "dumb": they are next token predictors, no doubt about it, but whether LLms as a class of system are smart or dumb is entirely unknown and entirely variable in time.

To interact with them effectively you must know how they behave, just like you have to know how humans behave to interact with them effectively. If you disagree, find someone with autism and have a conversation with them.


The "dumb" part comes from how it behaves in contexts where it lacks a lot of data, or where the data is skewed. Since they are tuned to give a prediction anyway and just make something up since sometimes those made up things are useful they will produce dumb results.

So people call them dumb since like dumb people they make strong statements about things they don't understand. And it doesn't matter how much smart things you encode them with, they will keep making strong statements about things they don't understand until they are fundamentally changed.

But since LLM are very smart about things where they have extensive data they can still be used to reliable solve many problems and probably in the future where we understand that better almost completely replace most lawyer and doctors work etc, because a lot of what a frontline doctor or basis lawyer work is very repetitive and can be encoded with billions of examples and decision paths into an expert system framework the LLM will follow.

So people say LLM are dumb since LLM will always keep making dumb statements. This is the same way we call Elon Musk dumb for making a lot of dumb statements, he is a smart guy but he makes dumb statements so her is dumb.


> they will keep making strong statements about things they don't understand until they are fundamentally changed

If ever there was a human quality.

Also, your explanation of "dumb" is really favoring the anti-llm side, and its a very generous interpretation. I suspect what is much more likely meant, is that token predictors cannot be smart, not now nor in the future after improvements, because they are token predictors and predicting tokens is not how intelligence works.

All of this is of course unfounded, and hidden behind the word "dumb".


> I suspect what is much more likely meant, is that token predictors cannot be smart, not now nor in the future after improvements, because they are token predictors and predicting tokens is not how intelligence works.

Why do you think that? LLM are used as expert systems today, in order to quickly navigate problems by breaking them down and iterating between different well known possible solutions and paths to check etc. That is how they work, they do that by using their next token predictions, and for things they aren't well trained on they will produce dumb results.

LLM has solved enough problems that almost nobody has the view you ridicule here, but there are still many who think LLM are thinking just like humans and that you can trust them just like humans. So its important to remind people these are just token predictors and lack many things humans do.

> If ever there was a human quality.

Humans can avoid doing that by using introspection, LLM can't. That some humans do it by not using introspection doesn't mean humans are incapable of it, we know humans are capable of it, which is why we can point out when the LLM is wrong with certainty, humans as a group make extremely good predictions.


> Why do you think that?

From (highly upvoted) parent:

> LLMs are text-prediction engines. They are not Artificial Intelligence, and shouldn’t not be treated in any form or fashion as if they possess intelligence > LLM-based “AI” ....

Hardly neutral statements, hardly accurate statements, yet highly upvoted. LLMs are AI, there is nothing to gain by pretending it isn't because of some secondary motive or opinion someone has.

> LLM has solved enough problems that almost nobody has the view you ridicule here

I think I have just shown you the parent literally claims LLMs are not AI, and do not possess intelligence.

> but there are still many who think LLM are thinking just like humans and that you can trust them just like humans.

I don't think many people think that, especially on HN, for two reasons:

1. Most high profile AI tools have immediately visible disclaimers saying, more or less, "AI makes mistakes". 2. People don't trust humans either, if anything people trust the AI more because it is not human.

People may be incorrectly worshipping LLMs, but they are doing it precisely because it is not human. If you put chatgpt behind a believable chatbox so people would actually think it was human, they would be much more skeptical.


Much of an LLM's capability comes from the structure encoded in its learned representations. The probabilistic outputs are primarily a way of expressing uncertainty and generating fluent text, while compression during training is what forces the model to discover that underlying structure.


> Much of an LLM's capability comes from the structure encoded in its learned representations

And thats encoded as a set of next token predictions. So the way to see how reliably it solves a problem is to look at the chain of predictions, and see where it is unreliable at finding the next spot, or where it always fails and you need to add that link to the dataset to train it.

This isn't magic, today we understand pretty well how to add new skills to LLM, and the better this is understood the faster progress will be.

This also means that if a context doesn't have any good predictions, it will produce a dumb prediction for that context. This results in these bad outcomes, because currently LLM doesn't have a map for where predictions are good or bad.


I agree. To me it seems blindingly obvious that a huge swath of the HN audience is gripped by fear and a loss of identity as a result of what LLMs have demonstrated over the last couple years, and they're lashing out as a result. Shaking their fist at the sky because they don't like the weather. Very understandable, but it's getting old to read month after month.


Would be nice to get high karma commenter votes count only ..


Not a dichotomy actually. Highly depends on the task.


Opinions differ. This is not news.


It's almost as if there were many people using this site, and there is no clear consensus on LLMs, so people from various camps upvote interesting stores to support their cause. And people who are still somewhat undecided upvote both, if they present good evidence.

I mean even perennially contentious topics will get this behavior.... some thing about emacs makes the front page, within a day or two there will be a vim post up there. Same with Rust is (good|bad), or if systemd creates an even more awesome tool, the haters will come along and recycle stories about bugs from over a decade ago.

There's a lot of people here. Not all of them read it every hour, and discussions like this among large groups often take a very long time with lots of repetition. Human group dynamics (aka politics) is slow.

> So tired of this discourse and this site.

You're welcome to leave if you don't like it. The site was like this long before you joined, and will like it long after you leave I'm sure.

It's also worth noting, that an awful lot of math discoveries are perfectly in line with dumb next token generators - they are finding a way to formally construct an argument and being surprised when it doesn't work, or surprised at the outcome of the grind. Not all of them are made by brilliant leaps of intuition.


Not sure what your point is? Those things can both be true.

Or should the discourse in a diverse community like HN only reflect the positions you personally hold?


Tell me how a 'nExT toKeN prEdIcTor' can make breakthroughs in math or play a game of chess. These activities aren't pure symbol manipulation, they require actual understanding at some level.


By predicting next tokens


Eh, I'm not going to litigate your claims.

My point is it's silly to whine that HN is a place where multiple points of view on the topic are aired out and discussed.

If you want a personal echo chamber where only your own beliefs are affirmed and anything else is flagged off or downvoted, I'm sure you can go find one or, worst case, vibe code one into existence.


Fair so let me be clear. I’m whining because the “next token predictor” reductionist point of view has been wrong and is only growing more wrong with time. Clearly these things can do things that actually matter. Do you disagree?


Even now you're engaging in this discussion as though I'm trying to litigate your point and that somehow forcing me to concede is, what, winning? I don't know.

I get the impression you want me to concede that the particular points of view you disagree with aren't worthy of representation here on HN.

I'm not going to do that.


I just prefer HN comments to be better reflections of reality. There is an unspoken expectation here that people here know what they’re talking about especially when it comes to technical matters. The rise of LLMs has given way to a HN branded populism that willingly denies reality as well. “Next token predictor” truthism is just so dumb and completely ignores the reality of what these tool are able to do. Smash that upvote button every time it feels good if you want but it’s just a meaningless take at this point. It won’t help you predict anything that’s coming.

Since we disagree on the present let’s informally do a “remind me 2 years” to this discussion and see what’s happened then.


> I just prefer HN comments to be better reflections of reality.

You mean your particular version of it.

It's interesting to see you consistently missing this point.

You've decided LLMs are clearly more than just complex but mindless statistical models.

You've decided that based on, it seems, the very impressive things these tools are capable of.

Therefore if anyone claims they're just mindless stastical models--with or without any attached judgement as to their actual utility or usefulness--then they are ipso facto wrong.

(And yes I just used endashes, damnit!)

That's on you.

It is in fact possible to simultaneously believe that LLMs are mindless token predictors and that they're enormously powerful.

These are entirely orthogonal beliefs.

Heck you could equally believe that LLMs represent true emerging AGI and that they still remain deeply flawed and are only an incremental step along the path of automation.

Or somewhere in between.

And discussing that space of possibilities is, I'd hope, precisely what HN is for.


It’s just a boring and unhelpful complaint that afaict largely serves to soothe the commenters ego rather than point at anything insightful that’s useful or predictive. Point me to your favorite “next token predictor” comment that was actually insightful or predictive. You have years of material to draw from.


https://news.ycombinator.com/item?id=49155075

"These models are probabilistic, you shouldn't blindly trust them in spaces where accuracy is really important" seems like pretty sound advice to me.


More that the “discourse” around anything “AI” is just another culture war at this point - dominated by facile taking points that are shallow, ill-considered, uninformed-and-uninformative, and re-re-recycled into worthless pulp by their respective tribal bubble.

And if someone’s first response to this comment is to try to peg which “side” I’m calling an idiot as if that’s the most important bit of info to dictate their response - that’s exactly what I’m talking about.


There are articles with far fewer upvotes and comments ranking higher on the front page right now, despite being the same age or older than this one. HNs opaque ranking system at it again.


Everyday? Which 10 problems were solved by mathematics grad students in the past 10 days?

OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?


Experienced similar between 5.4-mini vs 5.6-luna in our own pipelines but after spending some time on prompt optimization and testing out various reasoning effort levels 5.6-luna was well worth it. Did you just replace model selection while keeping everything else in place or spend some time on evaling with newer prompts etc?


No we kept prompts as is, just swapped model. The prompt is already quite optimized for the task. How would updating it possibly make a more intelligent model spend less tokens than a less intelligent model? Care to elaborate?


Most of the time when upgrading models we have needed to change prompts to get the same performance (let alone better performance). Usually, your prompt is overfit to the specific model doing the specific task. For example often your previous prompt is overspecifying and creating contradictions that a dumber model would just gloss over whereas a smarter model will try even harder to follow.


I think the fundamental difference between our assumptions is you believe prompts to be optimized for tasks rather than model-task pairs. The only elaboration I can give you is empirical observations and model providers own guidance (as someone has already linked here). I'm pretty sure you probably have specific parts of your prompts that came about due to specific failure modes observed in your evals of running the task against first model. These vary across models in my experience, and it's always worth redoing this calibration process.


Yikes, you can't really expect prompts to just be model agnostic


Here is an example of a guide from OpenAI on how you should prompt 5.6 differently than their previous models.

https://developers.openai.com/api/docs/guides/latest-model#p...


We can assume (outside of Ollama) that they meant the strongest model from each lab. If you limit yourself to just looking at the literal strings in the list, literally none of these are models. What model is "Deepseek" or "GPT"?


Well, GPT-1 was originally called GPT, and it's certainly a model.


How is this any different than what we have already? We've had this ability for ages (6+ months, decades in the AI world), you can literally today easily prompt CC or Codex to use subagents to accomplish tasks and they'll do it well. My entire workflow is one top level orchestrator chat creating tickets to dispatch to subagents to implement, and other subagents to verify. Why is this being sold as a new thing? Have HN users never tried tried asking CC or Codex to use subagents?


The top comment on the thread explains this will involve subagent to subagent comms.

To what effect I don’t know… I thought subagents were useful because they were explicitly single purpose and bound to a narrow context


I'd love to read more about how to use this workflow. What kind of top level instructions does this actually work with? Is there an article out there with some concrete examples of how to do this effectively?


in opencode, if you directly referencee a subagent keyword or the name of a defined agent, it'll often spawn the agent.

If you don't mention it directly, it's 50/50 whether any given request will invoke a subagent.

The same with tools, skills, etc. No matter how smart these LLMs appear, they rarely do thinking as you expect.

So basically: learn how the harnesses operate, and know the names of the tools they have.


I assume this is ~equivalent to ultracode in Claude Code, which can deploy a tree of hundreds of nested subagents and was just released experimentally 5 weeks ago IIRC.


Because most people need complexity to be wrapped in a simple UI/UX. Most people just want the one-two button press and be on their way.


As usual HN posters are hyper aware of other's credentials while ignoring that their BS in CS (if that) doesn't magically qualify them to assess everything in every domain.

"I'm a software engineer, I'm sure if I had the time to study Neuroscience, I'd figure out what all of these researchers failed to realize all these decades! I (alone) have the magic of critical and logical thinking"


A lot of us here have masters, PhDs, have published in academia, worked in the hard sciences or different engineerinf disciplines.

But I agree, when youre on the internet no ones knows you're a dog.



Seeing Jonathan blow invoked and his response in the comments tickled me pink


Was it Casey Muratori that spoke about an AI educing allocations from 10k per frame to 200 per frame, but the manual programming work got it to 0 instead?


The stupidest thing is that according to that logic, ncurses is a game engine too.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: