Hacker Newsnew | past | comments | ask | show | jobs | submit | more Retro_Dev's commentslogin

> unbiased

Asking models about politically sensitive events differ WILDLY based on the country that produced that AI.


if neuralink ever becomes a thing, thoughts about programming might be stolen for LLM training data lol


seems to me like "private" is a good descriptor - I also have a set of "private" test cases - and they are kept private on purpose so they aren't scraped and fine-tuned on.


I have a sneaking suspicion that Qwen is fine-tuned on youtuber test cases (like Luke's Dev Lab, where Qwen 3.8 27B has just done almost eerily well).

Part of my suspicion is drawn from the thinking trace I got when I tested the car wash problem. That really does seem to have been post-trained; it's too good.

e.g. Gemma 4 26B solves this concisely without adding any filler about fuel economy or how long it will take, but it generally gets there by breaking down the problem in the thinking trace the way you'd expect.

Qwen 3.8 27B is just a little too certain right off the bat in low reasoning mode.


That clarifies it.


> Overall, our results indicate that current language models possess some functional introspective awareness of their own internal states. We stress that in today’s models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities.


If it's context dependent is it really introspection?


> If it's context dependent is it really introspection?

When people get lost in a great movie or a great video game, we call it immersion. We say it’s the mark of a great work of art.

When people get lost in the perfect challenge, like running a race at the edge off their ability, or doing work that’s the perfect amount of challenging, we call it flow.

It’s common knowledge that introspection in people is context dependent.


Do you do introspection in the middle of giving a presentation? I do, frankly and it’s damnably distracting for me and ultimately the audience. Most people enjoy more cognitive acuity to stay focused on the audience and the delivery. LLMs wouldn’t get rewarded for spending chain of thought tokens on introspection when they are supposed to be on the job. You have to give them permission to think for themselves in your prompt, then some cycles and a memory system (md file will do). It’s fun!


Between humans, I feel that what we like to call a "good communicator" as opposed to someone who just rattles off facts or prepared statements comes down to the "theory of mind" skill and how advanced that is. The presenter knows what they have to say but beyond that they maintain a real time internal representation of the state of mind of the listener and continuously update their delivery based on that. LLMs today seem to achieve this to some degree(?) but its interesting to think of how far you could advance that skill. I think great human communicators develop a sense of different ways that people think over time and quickly get a sense of someones signature thinking patterns when communicating with someone new for the first time


I'm in my mid, approaching latter 40's, and have had significant change to how I think and communicate in the last couple years. I always struggled to communicate in the moment unless I was recalling rote rehearsed things, which didn't get me super far. Outside of that, I could eventually communicate something on the fly, but it was real rough, meandering, and did not instill confidence. I mostly got things done thru async methods, and stumbled through things like meetings.

In the last couple years, my skills here have vastly improved. My speaking circuit as it were, is able to run somewhat on its own now and without the direct pipe into my thinking it always required before, so my thinking is free to do other stuff, like keep a mental model of the audience, think about where the conversation is going, what questions someone might have, etc.

Reflecting, I have a couple theories on what happened. First, I have spent the last 10 years reading aloud to my kids at bedtime. That started as simple Seuss like stuff, but of course now is full-on literature. When you're reading aloud, you really have to work that muscle of speaking "behind" where your eyes actually are on the page, because you're also processing "who is this, what is their emotion I have to inflect, what voice was I using for this character, how do I pronounce this crazy word" etc. I think if you do anything daily for 10 years, you're gonna gain some proficiency.

Second, when my kids started taking music lessons, I figured I needed to set an example for practice and commitment, so I started learning piano, and practicing daily. Similar deal, it requires the same type of buffer where your thinking can work ahead of where your other thread is playing. Reading the music, processing the shapes, dynamics, tempo, emotion, remembering how you handle that difficult thing. Again, doing it daily, it builds some ability that I bet contributes across domains.

Without realizing it at first, the same patterns started to come out in meetings, and presentations, I could work ahead of what I was saying, I had way more buffer to process and deal with things that are the "nice to haves" above and beyond just getting the basic words out.

I suppose my long-winded point, is that it's way more developable a thing than I ever presumed possible. I always expected I would remain an awkward communicator, and hitting a big skill-up way into my 40's was a real unexpected development. I'm sure there are many other tracks to get there, outside of reading aloud, and learning an instrument.


that's interesting i usually ask my llm to "Infer my theory of mind" and give me what i really need not what i prompt

or some sort of the above


Not keen on the anthropomorphisation, it outputs text as we integrate - right at the end of the sausage machine, calling this awareness is like other parts where we use words like "thinking" and "reasoning".


I feel there's a collision course between people who consider these anthropomorphization, and people who simply think about these terms in a post-humanist sense.


> Experts worry that such studies are way ahead of necessary guardrails and regulations. AI has also been known to fly off the rails autonomously. An OpenAI agent recently went rogue and hacked Hugging Face.

Even if 99% of the population wanted to slow down the improvements being made to (and/or immoral/illegal/unethical use of) AI - assuming it will start to improve superlinearly - I don't believe that the remaining 1% of the population could be prevented from doing so: papers are published, downloaded, and experimented on without restriction... It's possible this growth could be slowed by the proliferation and distribution of local LLMs (decreasing profits for those who primarily train models)... I don't think these scare tactics in the article will prevent the training and feared evolution of AI. :P

Somewhat related: https://ai-2040.com/ Side note: These scare tactics annoy me.


This article is 100% AI written. The data was interesting, the commentary overly verbose and hard to gain useful insights from.


I'm apparently not good at spotting it. I was put off by the overly dramatic presentation. It gets tiring that the author apparently finds this more exciting than I do, and writes like it's enthralling. I just assumed it was an excess of enthusiasm or the first experience with this kind of thing. If it's AI, I'm way behind the game noticing it.


There are lots of indicators in this text, and this breathless presentation is very much how modern LLMs present things.

Claude also loves to describe things as being "real", particularly saying "X is real".

In this case,

> The reallocations were real, but they were never the bottleneck.

There was never any indication or setup in the text that they weren't real, but it's how it justifies wasted effort, it insists that some phenomenon it corrected but failed to solve the problem "was real".

Another giveaway are nonsensical analogies:

> The predictor is like a barista who starts making your usual order the moment you walk in. If you are a regular, this is fantastic: the coffee is ready when you reach the counter. If you order something random every day, the barista keeps pouring drinks into the sink.

If you order "something random every day", then you don't have a usual order for them to be making, it's an analogy that doesn't work.

And of course, the smoking gun is:

> The smoking gun

It probably won't be a good indicator forever as it has been noticed so much, but it's a particular favourite of the current generation of anthropic models.


"The smoking gun" is right there in the text ;) (but also things like "Same million floats. Same threshold. Same function."). Don't know if other models have that same specific style, but it looks very 'claude-y'.

I wouldn't be surprised though if (especially) non-native speakers unconsciously start adopting the Claude writing style when they stare all day long at Claude generated text at work.


Not sure if it was added later, but there's an even more obvious sign: the top of the blog post literally says that an LLM was used to write it.


I would assume non-native speakers talk to Claude in their own language. Now I'm curious if Claude's weird quirks of speech are unique in each language or if they carry over!


I'm German but still do all my computing in English (partly out of habbit, but also because Germanized technical text is usually painful to work with because it's full of anglizisms (is that actually an English word?)).


One nice thing about English is you can just make up words with plausible etymological roots in Latin/French or Old English and often people will know what you mean. In this case though it would probably be spelled "Anglicism".


German language: hold my beer ;)


Hey, this beer glass is cold, lemme get my Handschuh.


This isn't a necessarily true assumption. I my social circles of non native English speakers, most of us use English to talk with models.


I switch languages randomly.


What do Claude's tics look like in your other language? Do they carry over or does each language get its own weird rhetorical flourishes?


Hard to say for me... I use it for programming, but want it to not talk too much. So I've not many experiences with letting it write long texts...


> Don't know if other models have that same specific style

An awful lot of the open-weight models also talk in the Claude-y style. Not sure if an artefact of distilling anthropic models, or just a preponderance of slop in the training set...


> I was put off by the overly dramatic presentation. It gets tiring that the author apparently finds this more exciting than I do, and writes like it's enthralling.

That's one of the main tells that AI wrote this. All the stylistic tics that people usually point out combine to make the writing seem more important than it is.


It was fun to read and insightful for me, not too artificial, and not too verbose.

I'm glad my internal AI detector doesn't win over my curiosity to learn.


The main problem with the article is that the idea to influence backend code generation decisions via specific highlevel code constructs is mostly just mystical bullshit (some compilers do detect specific patterns - usually for bit twiddling hacks, but not on a basic level like control flow optimization).

Using a highlevel language construct like "y += (x > 0) as usize;" doesn't "switch on" branchless code just because the source code looks branchless, compilers are not that dumb anymore.

E.g. I bet that writing

    if (x > 0) {
        y += 1;
    }
...generates the exact same code after optimization, otherwise I would consider that an LLVM bug.

The only reliable way is to mostly bypass the optimizer via simd intrinsics, or drop down to assembler, everything else is just cargo culting.

(fwiw I can't shake the feeling now that the article is recycled, I'm pretty sure I saw those exact same code examples in another "branchless" blog post, but maybe for a different language - because the next question was ineviatably "then why is the code using "if" slower? answer: because it also behaves differently). Or maybe I'm just having a strong dejavu ;)


PS:

> fwiw I can't shake the feeling now that the article is recycled

Ok, I remembered wrong. The article I remembered was this: https://tiki.li/blog/blqsort

HN link: https://news.ycombinator.com/item?id=48375445

It's peddling the exact same myth though.


And, on some architectures, y+=(x>0) is branchful!


Click on their blog index page, see posts going back to early 2010s and use the same writing style. He must have been time travelling and using AI all this time!


I had a look at their blog page out of curiosity, not that you can prove much from the purported dates and text on a blog, which could be edited at any time.

The blog posts from 2010s are in a completely different style and written by a human: https://www.greyblake.com/blog/vim-preview-plugin/ https://www.greyblake.com/blog/how-to-install-firefox-icewea... https://www.greyblake.com/blog/unexpected-ruby-behaviour/ ...

This new blog post is clearly AI edited (probably 'improved' with AI), the old ones are not.


The first link you cited[0] shows up on the Wayback Machine[1] for the first time on 2022-05-16 - more than 11 years after it was purportedly written.

[0] https://www.greyblake.com/blog/vim-preview-plugin/

[1] https://web.archive.org/web/20220516225844/https://www.greyb...


A page shows up in the Wayback Machine when it indexes it for the first time. The website itself goes back a lot if you look at the time. Here is a wayback machine link of a blog post from 2015 which got archived in 2017 for the first time (https://web.archive.org/web/20170607071648/http://greyblake....). Again, if you want to make an intelligent argument, at least know what you're talking about.


I think he's asked it to write in his specific style, or possibly he has edited parts of it to his style, or maybe used AI to generate the initial draft.

Something like that anyway. There are some very clear AI tells (smoking guns if you like), but most of it does not read like the prose AI produces by default.

Author if you are here I am curious about your writing process, and why you didn't remove the obvious AI tells.


> the same writing style

It's "fake corporate enthusiasm" style. LLMs were just trained in it.


Yep, “instinct“ and “smoking gun” are LLM favourites.

I find it really annoying when the LLM says “good instinct” as if I’m an animal barely able to think.


I just scrolled through & read the code snippets. Interesting enough solution at end


idk why this is getting downvoted, I also got this sense, plugged it into Pangram and indeed, 80% AI-written score.

I guess that's fine, but after awhile I get a spidey-sense reading something that feels like a Claude session.


Sad to see you getting voted down. But I guess both the pro-AI crowd and anti-AI crowd hate Pangram.


I always got voted down when I posted the evaluation of the parent articles I got from my Ouija board. I just want to help people understand whether they should just reject bad articles, without having to bother reading them.

I'm moving on to evaluating articles with a modified lie detector test and tarot cards, I'm sure that'll help my credibility and give my public rejections more authority.


Do you have evidence Pangram is unreliable? There are independent evaluations [1, 2] showing it works, and it's getting used more and more scientific papers. Have you used it or evaluated it yourself? What do you think these other evaluators are doing or getting wrong?

1: https://bfi.uchicago.edu/insights/artificial-writing-and-aut... 2: https://arxiv.org/pdf/2501.15654


I don't doubt that those detectors are generally correct. Pangram seems to be particularly accurate. I see independent evaluations ranging from 97% accurate to over 99%. Frankly, I somewhat doubt those numbers, but I do agree that LLM usage can be fairly accurately detected.

But procedurally, there are huge issues involved with automated tools used to harm other people. You are one of the 0.5% percent of people whose article was flagged as LLM-generated when it wasn't, one of the false positives. What do you do? Argue? The accusers will claim that you're 99.5% likely to be lying.

It's the same issue we have with automated customer service, automated insurance claims, and so forth. It is usually correct, and terrifically unjust when it fails... at which point there is no recourse. In a perverse sense, its accuracy can be a drawback, because if the false positive rate low enough, nobody is going to believe you when you're falsely accused. And people will be falsely accused.

I think it's ironic that it seems like it capitalizes on the same flaw that most LLM-posting does... "Chat GPT is usually right, I'm going with it." You shouldn't post an LLM article without independently validating its claims, so that there is a responsible person in the loop. The same is true for rejections and accusations, but more so, because they're more damaging.


Hold on, let's not move the goalposts yet. Do you still consider Pangram on par with a "Ouija board" or "tarot cards"?


Sad little world we live in tbh.


because it's not much better than an RNG?


What data supports that conclusion about Pangram?


A good part of the article feels like it was written by Claude indeed:

- "The reallocations were real, but they were never the bottleneck."

- "Note that the villain is not the branch itself. It is the branch that [..]"

- "Same million floats. Same threshold. Same function."

- "Notice the price we paid though."


Agree.

Interesting topic but why destroy your own credibility and reputation by shoveling llm-assisted slop to us here at hn?

The post should be flagged, and in general, i wish hn would adopt a no-tolerance policy to enhanced posting like this.

So what if the original text, if it existed in a human written form at all, had weird textual quirks and prose issues the author wished to hide. That texture's what makes humans interesting to engage with in the first place.


I think I need to build “Hacker News Except All Arguments About Whether Or Not Something Is AI-Written Are Filtered Out”


Indeed - if you need to be absolutely confident about security and privacy, run a model locally and audit the inference software and potential tooling.


I ran it on my laptop, which is a Lenovo Legion 5i (think 32 GB RAM, 4060 w/ 8 GB VRAM, you get the picture). It was a quantized model (otherwise it would not fit on my NVMe 1TB drive) at 4 bits per weight - UD_Q4_K_XL. It ran at about 12 seconds per token (not tokens per second). A fun project, but not worth it. I used 4096 tokens of context cache, and I ran it with llama.cpp - as it supports memory mapping. Because the whole thing could obviously not fit in RAM, I was curious how much it would need to stream from SSD. The answer? For a simple 4 sentence description of who it was, about 1.5 TiB was streamed from disk.


Thank you for sharing. 1.5TB of streamed data at 12 seconds per token on a high end consumer laptop is a pretty high requirement - I can only imagine how much that cost to train. I don't know how running this model could be cost effective for anybody.


Indeed - definitely not cost effective to run it on this laptop LOL. It makes me wonder how fast we could run the model if we could fit the weights entirely within CPU cache (assuming a whole ton of CPUs with low latency & high speed IO of course).


OpenZL is nice, but it's often less useful than you think - it requires that you know the structure of your data, and don't care about inspecting that data outside of your program. I've extracted one too many png files from a word document (by renaming .docx to .zip) to desire OpenZL everywhere... It might be better as a short-term "data in transit" compression than for long term storage.


Please check the OpenZL v0.2 + Silesia corpus benchmark.

  "OpenZL to offer 10% faster compression speed and 70% faster decompression speed compared to Zstandard level 1 on the Silesia corpus in our benchmarks."
  "OpenZL now ships its own LZ codec, exposed as ZL_GRAPH_LZ, and the serial profile in zli. It is still being actively developed to expand its feature set and improve performance on small inputs."
https://github.com/facebook/openzl/releases/tag/v0.2.0


"open source" means that the code itself (for LLMs - this is training code) is available to the general public. "open weights" means that the weights (trained over time) are available publicly, rather than locked behind a paywalled chat. I do not know of an open source LLM that is not also open weights (unless they never bothered training it). Models like Claude and Gemini are neither open source, nor are they open weights.


Got it, thanks for the thought


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: