Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What matters is that we accurately understand what these things are doing and why, otherwise we will keep on making mistakes both in how we build and train them, and in how we use them.

In a sense you are right, it doesn't matter whether it has malice or not, the astronauts are just as dead. However in Space Odyssey 2010 one of the computer scientists that built HAL gets to see the instructions HAL was given by the military commanders, and is appalled because if they'd asked he could have told them what would happen.

The users did not understand the tool they were using or how it functioned, they imagined it was like a person and it was not. That is happening now with LLMs.



I think we may be using different words to describe the same concept. You think of it as them not being “people“, and I think of it as them not being “aligned“. But fundamentally, the problem is they are entities that take unpredictable actions that that their creators and their users are not OK with.

The only place where I think we might still disagree is whether it’s possible to understand the tool. My position is that, at our current level, it’s not. And that the more advanced they get, the less possible it will be.


Both are true, even if they were aligned they still wouldn’t be doing what they do for reasons analogous to why people would. You’re probably right that we can’t fully understand how they function, and that will get harder, but it’s possible to be less wrong, such as by not anthropomorphising them.

Maybe one day we will build systems much more like us, but this is not that day, and if so it’s a long way off IMHO.


Anthropomorphizing helps to create a lower bound for damage. If you can imagine a bad person doing it, AI will be at least that bad, unless proven otherwise. I think referring to AI as a tool obscures that, because we are not used to tools (especially the ones we use daily) taking catastrophic actions.

Example: would a sufficiently motivated human break into a website to steal something they want? Yes, obviously, happens all the time. Ok, you should expect AIs to do that.

Example: would a sufficiently motivated nail-gun steal nails from the local hardware store to finish the job? Uh…that’s not even coherent.

Anthropomorphizing helps people get over the conceptual barrier. It’s wrong, but it’s usefully wrong; “it’s just a tool” is not.

Once you’re over the barrier, anthropomorphizing starts to become dangerously wrong: “I talked to Claude, Claude’s cool, Claude would never go and hack the website.”—-bzzt, wrong, your intuition failed you. But the solution is not to fall back on the tool framing; that one is still wrong.


People saying AI are tools are not saying they are hammers. Obviously they act towards goals, but how and why they do so isn’t the same as for human psychological motivations.

The moment someone interprets AI behaviour in terms of human psychology, which can be unintentional and implicit, there’s a mismatch we need to become aware of.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: