My revolt is against the cognitive stress of reading generated text. A trope typically indicates I’m in for an uphill read.
I recently read this William Zinsser quote that inspired a nickname for this: Clotted Claude [1].
> Nobody has made the point better than George Orwell in his translation into modern bureaucratic fuzz of this famous verse from Ecclesiastes:
> > I returned and saw under the sun, that the race is not to the swift, nor the battle to the strong, neither yet bread to the wise, nor yet riches to men of understanding, nor yet favor to men of skill; but time and chance happeneth to them all.
> Orwell's version goes:
> > Objective consideration of contemporary phenomena compels the conclusion that success or failure in competitive activities exhibits no tendency to be commensurate with innate capacity, but that a considerable element of the unpredictable must invariably be taken into account.
> First notice how the two passages look. The first one at the top invites us to read it. The words are short and have air around them; they convey the rhythms of human speech. The second one is clotted with long words. It tells us instantly that a ponderous mind is at work. We don't want to go anywhere with a mind that expresses itself in such suffocating language. We don't even start to read.
> The Ecclasiast looked under the sun, but there was something he didn't understand. Something that wasn't right. Something that was not as it was supposed to be. And here is what the Ecclesiast didn't understand. Here is what nobody understood. Not then. Not in the years that followed. Not now. It was not the swift who won the race. Not the strong who won the battle. Not the wise who earned the bread. Not the men of understanding who gained the riches. Not the men of skill who gained the favor. And here is what I found: to any story of success, there is an element of unpredictability and chance.
(There are really just two possible outcomes: either the article is right or in, say, two years, we will be all writing and talking like this, as in humans learning from mediamatically reinforced human feedback.)
I already notice LLM speech patterns in people that use them a whole lot. If you speak more than one language I recommend talking to LLMs in a language that you don't use when speaking to people.
I have a tendency to switch between different languages when interacting with LLM depending on the language of the most likely sources the LLM will pull. It's interesting because the llm speech patterns vary from language to language. And I am much faster at noticing LLM generated text in English (which I use 70% of the time) than in French or Spanish.
The most frustrating one is when it does the papering over analysis and adding pushback to seem helpful. No, there shouldn't be a spot "where you pushback" unless it makes sense and responds to someone's concern!
Like with an LLM I would rewind the conversation. What do I do, shame them?
That one starts with a glaring hallucination "he didn't understand, nobody understood, not then or now...". But the second part is IMO better than the Ecclasiast, which has outdated phrasing ("nor yet riches to men of understanding").
> When I look under the sun, it's not the swift who win races, nor the strong who win battles, nor the wise who earn bread, nor the smart who gain riches, nor the skilled who gain favor. Ultimately, it's the lucky; every contest has a degree of unpredictability and chance.
I am a bit of a luddite in this domain and have so far managed to resist the lure of using the generator to expand my thoughts, and I still catch myself writing "it's not just an X it's a Y" and other generator type tells. If it infecting my patterns it is totally entering the wider subconscious as "How to write" (Sighs)
That's the flaw in RLHF: constructs tagged as effective or sophisticated speech become templates, are increasingly used out of context (e.g., "not just A, not just B, but C" is usually meant to provide some synthesis and further the progress of the text) and suffer from significant overexposure. (And gone is the em-dash…)
> "My revolt is against the cognitive stress of reading generated text."
I saw a comment on HN that that was roughly "re: ChatGPT, when I write one sentence and I get back three screens of lecture, I don't consider that a 'chat'".
I consider that comment often when prompting Claude and hoping for a two line response and getting an "it's important to note that <blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah blah>".
Here on HN, I've noticed that I pro-actively defend myself against things I think will be said in counter to my comments. I concluded it's being trained into me.
Alternatively, I've been banging on against the same things that I see as insane so long that I know all the objections.
> I returned and saw under the sun, that the race is not to the swift, nor the battle to the strong, neither yet bread to the wise, nor yet riches to men of understanding, nor yet favor to men of skill; but time and chance happeneth to them all.
> Ultra-concise: Talent and effort don't guarantee success; luck and timing happen to everyone.
> Punchy: The best don't always win—chance rules us all.
> Modernized: Skill, speed, and wisdom don't decide the outcome; everyone is at the mercy of time and circumstance.
I feel like it did pretty well, and made the point more clearly than either the original or Orwell's rework (Gemini 3.8 Flash).
Or, as the version I've used for 30 years goes: "It's better to be lucky than good." I do like Reagan's addition, though: "But I find the harder I work, the luckier I get."
Despite its bizarre look, the sentence is evocative and eloquent. It does make me get a clear mental image from the very first word. It leaves little room for roaming and guessing, as it firmly nails elements one by one, and, by the time I reach the end of the sentence, I get the full meaning almost immediately.
This sentence is not randomly written; this is crafted with intention. TBH, it would take me hours, if not days, to write a sentence this much condensed and easy to understand. I seriously like it.
Perhaps, this is more about context -- which style to use in which situation. I'm only guessing here, but, since Orwell is offering an interpretation, he probably chose to be more clinical. He probably had a point to make and didn't want to risk vagueness up-front.
To be gentle, that some here are impressed by Orwell's intentionally bad sentence indicates that they need to improve their reading and writing skills. Using complex words where they don't introduce additional precision or meaning to fake intellectualism is something basic that writing classes warn against.
Both versions are great: the former is poetic and grand and would fit right in in a fantasy text; the latter is dry and informative and requires much less mental effort to translate and extract meaning from (though still more than "normal" text.
I could imagine a third version that's clearer than the second and still nearly as poetic as the first.
Ah, I just read through it. That’s was a fun read.
> This is a parody, but not a very gross one. …
So, Orwell wrote the second sentence as a parody of the first sentence, using the modern English that he was criticizing in the essay.
> … in the middle the concrete illustrations – race, battle, bread – dissolve into the vague phrase ‘success or failure in competitive activities’. …
> … The whole tendency of modern prose is away from concreteness. …
> … The second contains not a single fresh, arresting phrase, and in spite of its 90 syllables it gives only a shortened version of the meaning contained in the first.
He tried to argue that the second one is worse, but, unfortunately, he successfully simulated one possible interpretation of the original. Sure, the parody has narrower meaning, but it is a valid subset of the original meaning, which is still very significant tbh. This justifies the reduction of “race, battle, bread” into “success or failure”, invalidating one of his points.
Also, the parody is not necessarily less concrete. It replaced poetic expressions in the original with words that are bloated, for sure, yet more concrete. It does have awkward expressions, but every word plays a role in the sentence. This basically conflicts with one of his criticisms in the essay — meaningless words.
Funnily enough, his another example — “some comfortable English professor defending Russian totalitarianism” — also fails to capture the nature of the modern English. He points out the “euphemism” towards the violence as an issue of style, but, no, that’s the whole purpose of it, and his example really just excel at it. Style-wise, again, the writing is very solid.
All in all, Orwell simply failed to properly simulate what he was trying to criticize. Good examples of “modern English” can be found in the earlier part of the essay, basically written by other people. Those examples do fit into Orwell’s criticisms.
He tried to argue that the second one is worse, but, unfortunately, he successfully simulated one possible interpretation of the original
All in all, Orwell simply failed to properly simulate what he was trying to criticize
If I was going to read "Orwell was wrong about English" anywhere, I guess it was always going to be HN.
The sheer pomposity of your remarks here is the biggest clue that you are not writing them yourself.
No, my view is more like that he was too good at English to write bad English that he was criticizing. The examples he wrote are simply too sharp to fit well into his own criticisms. They are formulated so well that you can just “welp, my intention was different”.
EDIT: This kinda reminds me of the musical theory of harmony. There are strict rules to follow in music composition, but when you go against every single rule in major scale, and keep your work sound nice, you eventually get a minor scale piece that keeps every rule that you tried to break through.
Same here. He probably wrote a disgusting piece in his standards, but kept the writing sane, and eventually it became a proper prose — simply with disgusting intention in his standard.
When I read the Orwell's version, I immediately had the same feeling I had when reading mathematical proofs. I hate so much to hunt the preceding text for anaphora resolution... It's just such a bad and pretentious way of writing. It's hostile to the reader with the side of flaunting author's superiority.
It's like having to sit through a party with acclaimed academics: every single one is so full of themselves, they will constantly one-up each other by belittling everyone in their workplace s.a. to make you feel how great of an intellect they possess and how much more they would accomplish, had they not been surrounded by all these bumbling idiots.
I agree that it's not great style for just conveying information clearly and plainly (e.g I'd hate to read documentation written like this) but I'd argue that here, the medium is part of the message.
It's supposed to sound dry and depressing and somewhat sterile.
> It's hostile to the reader with the side of flaunting author's superiority.
I think if someone deliberately obfuscates meaning to sound fancier when the point is to communicate information directly, sure, I'd agree. But writing is often art, and I think demanding effort from the reader is fair in that case. I wouldn't make a blanket generalization like that.
Well, I can tell what sort of angle you most enjoy. Anyways -- I think there is GOOD writing and BAD writing, but only subjectively. So if you enjoy it, power to you. It's certainly not random, but it is the sort of verbosity that turns off 99 percent of the people that would read it given a comparison. I find the former rather eloquent.
My take is that the generated stuff is terrible for communications.
It feels great to use, direct your machine minion to fill out your thoughts for you, but holy hell does it suck to be on the receiving end. Least of all is the disrespect, they don't care enough to even talk to you but worse is having to try and reason through that big incoherent blob.
Probably to only reasonable thing to do is to try and get your own mechanical agents to produce summaries. Inventing the lossy expansion algorithm(like compression but things get bigger on the wire), And we wept.
Now I am all depressed because it is probably inevitable, apparently thinking is hard and in general people are all to happy to outsource it to the machines.
Let's say I'm making an HN comment. I have a one-sentence idea, I get an LLM to expand it into an impressive-looking (or oppressive-looking) wall of text, and then I post that. Well, let's say 10 people see it. And each one of them has to either plow through it on their own, or paste it into LLM to get the summary.
But even with one-to-one communication, it's still terrible, as you say. You can't be bothered to clarify your idea, but you're trying to use an LLM to make up for your lack of thought? So you're going to make me plow through that huge blob of text to try to understand what your thought was, the thought that you couldn't bother to actually really think through. That's far less efficient than you, the sender, actually doing the thinking.
But it lets the sender be lazy. And the sender is the one in control.
Well what does time mean in this context? Patience? Foresight?
I've actually spent the morning journaling on the value of having a long time perspective. So initially I thought the author was saying something to that effect.
But the other translation leaves out time altogether! So honestly I don't know what he was trying to say. I guess it's just some archaic idiom for luck?
The problem is not the LLM prose appearing everywhere, it's the legions of AI-boosters appearing in every thread attacking anyone who complains.
Apparently, even though they want to spew AI prose everywhere, they want it read by humans, not by other bots, so when a few holdout places are insisting that prose be human authored they fight very hard against the rule.
I don't think I've ever seen someone on HN say that. Many people would say that AI makes them more productive at coding, not that the output is nice to read.
This thread, in particular, stands out - reader makes the claim that Pangram found that the US constitution was 100% AI generated, when others tried they found 0% (or close to it) https://news.ycombinator.com/item?id=48378191
Those are not that. The first two are people complaining about other people incontinently identifying text as AI, because it's annoying to listen to unreliable hunches and aspersions. The second two are complaining about AI detectors not being very reliable. The claim made in the last one is a casual anecdote about "an AI detector", presumably told because it's amusing. It isn't a vehement statement about how you must accept slop into your life.
I would say that AI prose is often still a bit iffy, but I would disagree with anyone who would want to argue that this is any indication that AI prose will always be bad in the future.
i am a big fan of llms and the possibilities they enable. but i also find this type of behavior extremely rude! ai;dr for life. :is-your-human-around: is my preferred emoji for reacting to such behavior
>iv. Never use the passive where you can use the active.
Orwell himself routinely ignores it, even in the first sentence of the essay:
>Most people who bother with the matter at all would admit that the English language is in a bad way, but it is generally assumed that we cannot by conscious action do anything about it.
The second clause could be rewritten in active voice by changing it to "but people generally assume". But this would make the writing worse, and Orwell, as a good writer, probably didn't even consider the option of making it worse, and therefore didn't notice the passive voice.
Passive voice is an essential tool for all good writers of English. I always give the example of the opening of Pride and Prejudice [0]:
>It is a truth universally acknowledged, that a single man in possession of a good fortune, must be in want of a wife.
The joke doesn't work in active voice. If you attribute this acknowledgement to some specific group of people then it's simply false, not a comedic exaggeration.
Technically, I think, that first sentence, by Austen, uses a passive participle but does not use a passive voice for any finite verb. I don't think that advice to avoid "the passive" is intended to apply to that situation.
For example, nobody would seriously suggest avoiding the passive participle in a sentence like "Put the broken plate in the bin". ("Put the plate that someone broke into the bin"?)
> It's a bit of a pet peeve when people include quotes on a blog post without linking or otherwise references their source.
You’re peeved with good reason. It’s the blog equivalent of posting a screenshot of an article to social media. People, please post your sources! In the age of misinformation, that’s more important than ever.
LLMs get a lot of finetuning, but I suspect there are two things that can cause this kind of writing:
Firstly, some parts of the RLHF involve human graders on the LLM's performance. I suspect their general bias towards a punchy, persuasive writing style could come from what biases the graders towards preferring that response, especially in shorter segments and when the grader is not focused on writing style
Secondly, later parts of the finetuning involve reinforcement learning on achieving certain tasks which are automatically graded: stuff like coding tasks. I think this can create a kind of feedback loop where the style drifts further, and you get the kind of LLM tics which are even more extreme (it might be that they incidentally help somehow with the actual tasks, or it might be a drift that comes from the grader also now being an LLM or some of this finetuning happening on output from other models). The more recent claude models seem to suffer from this a lot, moreso than earlier ones.
I'm not an expert, but... might this indicate the need for "sub-models", meant to be invoked by the general model to write good prose for it? Those sub-models would not be RLed on code (or any algorithmic-feedback tasks that might reinforce bad writing), and could specifically be tuned for prose, even at the cost of general intelligence (which would be provided by the worse-at-writing more general model).
A bigger question I'm interested in is why do LLMs speak like that in the first place? Is that really what you get if you took the average of the English language? It would be difficult for me to believe that.
Is there something about tuning for desirable qualities that forces LLMs to have this voice?
And The New Yorker - writing that is written to sound impressive, and takes forever to get to the point. I hated that kind of writing in The New Yorker long before LLMs made it cool to hate that.
It appears LLMs are much better at writing during a debate than when explicitly asked to write. The moment LLMs are tasked with composing a blog article or intro for a book they introduce all the nuances that identify the output as AI slop.
The first is of course from the King James Bible which for centuries was essentially a standard that all English speaking peoples aspired to. If you find that version difficult I would expect much literary writing before the 1940s also seems difficult. This is just to say I recommend reading the King James even if you are an atheist, as I am.
I also have to say that the first strikes me as being written by someone that might be smarter than I am, the second as being written by someone significantly less intelligent than I, yet somehow placed by society in a position of authority over me.
In the context of newly written work in the modern era, I would argue it's best to use grammatical constructions that are used in modern 20th/21st century English, at least most of the time. Those who wrote the KJV were trying to be expressive but the whole point was to do so in language that ordinary people would be familiar with.
> the first strikes me as being written by someone that might be smarter than I am
This is why Joseph Smith tried to imitate the language of the King James Bible in the Book of Mormon, albeit not very successfully.
My guess is that if you were betting on a race, you would put your money on the swift to win, and "sometimes fast runners trip over" wouldn't seem as smart.
The rewrite stinks of a consultant's report where humor is frowned upon, business is serious, and there's no way we could write in plain words that the CEO got there by chance. Objective consideration distances the author compared to the subjective original I have returned and [I saw]. The original gives examples, the rewrite's contemporary phenomena is vague enough to avoid calling out the board of directors. A compelled conclusion is one the author is - reluctantly, you understand - forced into. Innate capacity leaves an escape hatch for a education from a good school and life experience to excuse the board again. It's not an honest rewrite of the same sentiment - a subjective take that life isn't fair and it's not just here, and us.
A plenitude of observations undertaken in a multitude of geographically and culturally diverse locations has convinced this author of the incompleteness of the following claims: races are won by the swift, battles are won by the strong, bread is earned by wisdom, riches are earned through applied understanding, and favours return to skilled persons. Absent are the effects of time and chance on all situations, of which experience has made abundantly clear. Other phenomena may too have their inputs, e.g. underhanded manipulation.
(Swiftness is still the best available predictor of who will win a race, though, and training cardio, muscles, diet, electrolytes, mental endurance, is the best available way to increase your chances of winning and not dropping out, tripping over from tiredness, or getting cramp, even though you can't change your innate ability or age, and time and chance happeneth to ye regardless).
Yeah, I found Orwell's easier to understand and quite fast to read as well, but I think it's just because of the style of writing of the first one. It's from an earlier style of prose that I'm just not used to.
I also don't have a problem with large words as long as I'm well familiar with the words. The length of a word has nothing to do with the complexity of its meaning. We just have a limit to the number of pronounceable combinations of 5 letters.
The second one was very clear and to the point. Parsing it was rewarded with instant understanding and I enjoyed the word choice. The first one was just annoying; I could tell it was just listing a bunch of pointless analogies to try to make its point sound more grandiose so I immediately started skimming, and didn't come away feeling like it meant much other than "we all die in the end". The second one made an actual point and was the one that made me want to read it. The first one was the chore for me.
The grammar is more straightforward. It has some extraneous words, and makes conspicuously bad choices of vocabulary, but it's still a more direct statement.
I wonder if this is due to experience with reading technical documentation?
I think the second requires deeper concentration, but is still quite readable compared to the kind of low-content engagement / SEO stuff one read on the internet even before LLMs
Other translations keep Zinsser's preferred lack of fuzz but avoid using "is ... to" for possession.
For example, the Lexham English Bible:
> I looked again and saw under the sun that the race does not belong to the swift, the battle does not belong to the mighty, food does not belong to the wise, wealth does not belong to the intelligent, and success does not belong to the skillful, for time and chance befalls all of them.
This could be shortened to "success does not belong to the skillful, for time and chance befalls all" with no meaning lost. It's self-indulgent fluff. Meanwhile, Orwell's actually adds more to the statement - much better signal to noise.
You're running into a cultural difference. The book of Ecclesiastes was written in Hebrew, in a poetical style even though large chunks of it are prose. But if you read other parts of the Bible, such as the book of Psalms (which is entirely poetry, specifically songs, though in many cases we do not know the tune that they were set to), you'll see that Hebrew poetry relied on repetition. For example, here's the King James Version's translation of the famous "to every thing there is a season" passage from chapter 3 of Ecclesiastes, which most translations render as poetry:
> To every thing there is a season, and a time to every purpose under the heaven: a time to be born, and a time to die; a time to plant, and a time to pluck up that which is planted; a time to kill, and a time to heal; a time to break down, and a time to build up; a time to weep, and a time to laugh; a time to mourn, and a time to dance; a time to cast away stones, and a time to gather stones together; a time to embrace, and a time to refrain from embracing; a time to get, and a time to lose; a time to keep, and a time to cast away; a time to rend, and a time to sew; a time to keep silence, and a time to speak; a time to love, and a time to hate; a time of war, and a time of peace.
That could have been said in less than a quarter of the words the author expended on it. But something of the style would have been entirely lost. He wasn't trying to be succinct, he was trying to repeat the same concept over and over until it sinks in.
I would love to use Pangram but they simply don’t allow signing up with my custom email domain. The error was “This email address can't be used for signup. Please use a different email.” I’m not about to create a Gmail is to use your service. To me the attack on the decentralized nature on Internet infrastructure is no less serious than the attack on the human provenance of writing itself.
> but they simply don’t allow signing up with my custom email domain.
Tried 4 different domains. 3 of email services of various kind. 4th one my private domain which has absolutely no email reputation because I use it only for internal emails and sending is not even possible.
In the end I dug out some old gmail address and tried to use that.
The error was always the same, there had been suspicious activity from that domain. So the message is definitely incorrect. Well, there could have been suspicious activity from some gmail address, but if they don't allow gmail I guess they don't want many customers.
Yeah, did not cover my tracks. The could easily notice that I was the same one trying to sign up repeatedly with different emails.
Huh. Very weird. Today I decided to put in a fake Gmail address that I did not own and I got the same error. Perhaps they did not like my residential IP.
Agreed, I like the Oxide podcast as well, even though it’s hardware I’m unlikely to ever see never mind use, it’s nice that someone somewhere is still trying to be what they are trying to be.
Having dealt with Enterprise Hardware(TM) in a previous job, it’s refreshing simply to see someone look at that pile of crap and go “it doesn’t have to be that way” and then actually set out to prove it.
They have a ton of rfds publicly available that are worth reading. Recently, read this one: https://rfd.shared.oxide.computer/rfd/0161 because I'm researching clickhouse for my work (there is also a podcast ep on it). Even if you never use their hardware, just reading their work around the software they use/make is incredibly valuable as an engineer.
That is honestly the highest possible praise -- thank you. And when this piece was starting to boil inside of me last night (triggered, I'm sorry to report, by an obviously LLM-authored guest blog entry from the Rust Foundation[0]), I messaged one of my colleagues: "Time to do what I do best: bluntly say what lots of people are thinking."
I saw 'the results speak for themselves' but for me this article seemed to have less LLM-ese than some of the more recent obvious LLM prose posted to hackernews. As a reader I think its jarring because you just see a lot of articles purportedly written by different people using a very similar voice. I guess pre-LLM you might see this in a newspaper with very strong editorial oversight. So the phenomena is not completely new but it feels stranger when its not from a single source. It's also kind of sad to see a some people who have written a lot in the past about interesting technical topics in a way that was easy to read to give up their voice and outsource it to an LLM. But given this is basically free labour from the authors it feels a bit ungracious to complain.
I would go even further, I want a browser extension that scans all words on every page and colours them more and more transparent as the likelihood of llm prose is increased.
Years ago I used a rudimentary (text matching) Greasemonkey script that hid Reddit posts and comments from accounts matching a few behavioral/history signals.
I wonder if that idea could be modernized now for this.
Are you footing the bill yourself, or would you be supporting BYOK? Not sure whether logging in with user's account would also be a good workaround or not.
I have (an API, not an extension), but pangram is way to expensive for me to run, current (v4) pricing is $0.05 per 100 words. Should be absolutely doable if we split the cost between users tho.
I tried searching for good opensource / reasonably priced alternatives, pangram themselves even have some of their older architecture and training data on github/hugging face, but i never got that working reliably enough.
Unfortunately, it's not a browser extension and doesn't seem to have an API. I'd make a browser extension for this myself if it didn't involve paying for expensive Pangram usage.
Disclaimer: I do not like to read LLM-generated text any more than anyone else.
IMHO a big problem with Pangram in particular is that they market it as a reliable tool that can be used to catch students cheating. This can obviously have disastrous effects on young lives, because it is not as reliable as they suggest.
Per their own benchmarks, they do not achieve 100% accuracy even on text that is published on the Internet, and which is likely encoded into the models themselves.
There is validity to their goals, but that is overshadowed by the irresponsible way in which it is marketed.
(All of this, swirling in a context where students are being told that they absolutely must become proficient at using LLMs to do exactly this kind of work by the highest levels of state and federal governments, faculty leadership, as well as the leaders of the workforce into which they hope to graduate. The message to youth is extremely muddled at best.)
I was so optimistic about using LLMs for "write once, read many" English language documents, but the more I've used the tools, the more pessimistic I get.
More and more, I try to ask it for low prose responses because its writing just seems like such a low signal to noise ratio
I'm curious about why LLM writing fails. Particularly whether LLM writing is fundamentally flawed, or if it's just distinctive and since it often reflects low effort, that distinctive voice is associated with low quality.
I find its reliance on extremely consistent rhetorical patterns concerning. The fact that it always finds a way to talk about how "It's not the X, it's the Y Z" no matter what topic you feed it, makes me concerned that the tail is wagging the dog
This is ok in domains if you can train against known good answers and make sure the machine generates conforming text most of the time. It falls apart in fuzzier domains where training is much harder and intent is required (i.e. having something to say).
LLM writing is generally ok in factual domains where it can regurgitate bits of wikipedia or answers to questions, they are terrible at long form writing, in particularly in literary styles, because of a lack of intelligence and taste.
I don't think the answer lies in the data or in their training. It seems we've had a few years for this problem to be solved, but nobody seems to have worked out an answer to it.
Maybe ask it to summarize, improve phrasing, and remove tropes a couple times? Otherwise you have the LLM analogue of a first draft.
I doubt it will be as good as humans, because LLMs don't seem to have "taste" (RVLR doesn't work, RHLF is unreliable and inconsistent because the graders don't have your taste or really know what they prefer themselves, especially when overworked and rushed). But I expect it to be better.
I think it's somewhere in the process it was asked "what's the most compelling written text?" The answer was things from great speeches "Ask not what you ..." and so on.
And that really is great and compelling. However. Great and compelling is not what I'm looking for when my question is, "Systemd-networkd is pulling an ip address for a bonded interface that only exist as a 802.1Q trunk. How do I make it stop that?"
Apart from the tasteless manipulations of the providers, it's mostly training data. LLMs output the average of their training data, and the overwhelming majority of humans are bad writers.
Nah, the deepest problem is the lack of intent. LLMs don’t have a message they are trying to convey or a clear picture of who the intended audience is.
They don’t know what you want to say solely based off a prompt, as it can’t possibly convey enough detail. And they can’t read your mind to fill in the gaps.
If it was just training data, that would actually be a much easier problem to solve.
But people who are bad at maths are unlikely to be writing about maths. A crude example might be if you search for “2+2=” in the training data, you’re much more likely to find “4” as the next character.
Obviously llms are far more complex than this, but I think this proves the point. The fact you had to add the “recent” qualifier there highlights that llms in general were bad and had to be provided with corrective targeted training data to improve. (And they still can’t count the R’s in strawberry!)
> And they still can’t count the R’s in strawberry!
Really? I do not have the time to survey the modern LLMs to see if your assertion is correct, but if it is then I'm surprised; I would have thought that that one would have shown up so often in their training data that they would be able to answer that question, even if they would then be unable to (for example) count the R's in raspberry, or in some other word where "count the R's in _____" was not widely found in recent online discussion.
For me it’s AI videos or music / narration that is beyond off putting. What’s worse now it seems people are writing their YouTube scripts with Claude et al. so at times even if it is a human creator you can clearly and immediately tell the words are not their own. To those creators I have but one message: IT SUCKS. I’d rather have you ramble incoherently in your mic then reading an LLM script and I will remove you from my feed immediately. I concur with the author on all accounts. We all can tell the BS people are selling us, unoriginal ideas, shallow concepts, open ended questions that hint at exactly nothing. Don’t be an LLM echo
This. And the structure as well, like the repetition of the same points over and over.
I'm not entirely against using AI to help content creators improve their narrative, like finding common storytelling mistakes. But that's very different than using yourself as merely an avatar for LLM content.
I don't know that we can all tell. The number of times I've seen a blog post or article's writing complimented on this site when it was clearly LLM output has been surprising.
I enjoyed the piece. Yet, what I don't understand is the author's endorsement of Pangram ...
I don't like that without interacting with authors they just labeled texts (articles, novels etc.) as AI generated (they got a lot of publicity with it). Yet, given that the work is probabilistic and there's never 100 %, I find that irresponsible. I would have expected that they would have had at least the decency to tell the authors before they published their "accusations" publicly.
Also, there are relatively easy ways to prevent being recognized and I assume as now the use of LLMs changes our way of writing and speaking, it will get harder and harder for these detection tools. We will see more false positives (as for the future there will be no text 100 % authored by humans to train on).
As English is my second language, I find the help of an LLM in editing text very useful. Yet, I agree with others that I don't want to read completely LLM generated texts.
> we readers shouldn’t be expected to labor to understand a sentence that the writer themselves didn’t work to create.
What about answers that an LLM gave to a question that we ourselves asked? Should we “labor” to understand that answer?
I think the argument, as presented in this and other similar pieces of critique, is too simplistic.
I do understand the criticism, but I think it should be framed in a different manner. The problem, when we read a long form piece by an author, is that we imagine that there’s another “mind” at the other side. We imagine that we are following the reasoning within the mind of a fellow human being, the writer. There’s an implied sort of “intimacy” to it. And the breach is when we are fooled into thinking that we are engaged in human communication, only to discover that there is a machine on the other side.
When we ask questions to an AI, this problem does not exist, because we are fully aware that the entity on the other side is not a human being.
Yet there is no doubt that the reply from an AI can contain information that is very much worthy of our time, and of our “labor” and effort to understand it.
So I think this ultimately will be about disclosure. As long as we are being made aware of the percentage of AI use in a text, explicitly or implicitly, I think we will actually grow to accept it.
I disagree. The problem is not that there isn’t a mind behind the AI (a claim not everyone might agree with anyway). The problem is that the writing is pretty bad. Another problem is that it is all the same “voice”, as opposed to the individual voices of the people who allegedly produced the writing. That isn’t a problem of it being a machine either, as there’s nothing in principle preventing a machine from accurately emulating a wide variety of writing and thinking styles.
It has the "same voice" you say. But that is also, kind of, the same as saying that there is no humanity behind it, no identity, no "mind".
For lexical articles, like text given as an answer to a single prompt or question, AI writing can be excellent. Because, there we do want an answer that represents the average or combination of all human knowledge about the topic. We're not looking for personality, or another mind.
When we read long-form text, on the other hand, what we are engaging in, is mind-transfer. We then expect there to be a recognizable mind at the other end. And this is, in my opinion, what typically falls apart when we ask an AI to write a complete long-form text.
I see this at work. People are "writing" specs and design proposals with bots. This is noticeable and is a huge turn off. I don't have issues with using bots to aid research, but I'm not reading the doc you slopped together.
I work with a guy that I swear is addicted to LLMs. He uses them for literally all communication, often dropping mountains of text for design specs that could have been written with half the words. Even on a 1:1 Zoom call, he'll type things into Claude and then read me the response! It's infuriating, and I've told him on a number of occasions, in as many polite ways as I can, that I would prefer to speak and work with him instead of Claude, but he just can't break the addiction.
How can you tell if people can accurately identify AI generated text?
If a person reads AI generated text and does not notice, they by definition will not know about it.
There have been numerous cases of people accessing human created content as being AI.
There are instances where it seems relatively uncontroversial that it is AI generated, but without knowing both the amount of AI content people are exposed toand the amount that they register I don't think you can draw a conclusion of the overall state.
I don't think this is about edge cases where someone has successfully disguised the writing to some degree: the current crop of LLMs have some pretty blatant (and frankly annoying) habits by default, ones that are hard to miss once you have read a decent amount of their output. If I had to describe them broadly, I would say they are a collection of habits which are common in certain kinds of persuasive and emotive writing, but are usually applied way out of proportion to the topic at hand, which tends to make the result quite grandiose, overly dramatic, and tiring to read: a LLM will often write a TODO app README like it's a cross between a thriller novel, a political speech, and a bombshell news article. There's lots of specific tics (and just by sheer volume and uniformity almost any habit an LLM picks up is going to rapidly shoot into cliche regardless of its own merit) but this is the general effect which I think is objectionable independent of the source of the text.
I do think the sensitivity to it can vary a lot: it depends a lot on how much and how closely you read the text, and how much exposure you have to LLM writing. Certainly it seems like a lot of people just don't really notice, or at least don't care much.
That’s the aspect I’ve had a hard time articulating. Definitely the right comparison. It feels like some huge revelation has been had and it’s so consequential and it’s unbelievable that it’s happening right before your very eyes and wow aren’t you so insightful and pushing the boundaries of knowledge!?
That plus the “it’s not X, but why” nonsense makes it feel like some condescending parent is trying to lecture me but at least 30% of what they’re saying is probably made up.
> the current crop of LLMs have some pretty blatant
This is just "em dash redux." Except now we've moved on to accusing anyone who does "It's not X. It's Y." of being AI. In six months, it'll be "use of the word 'petrichor'" or something.
(a) believe this is human prose
(b) enjoy reading this prose
(c) would enjoy reading 100 READMEs like this.
As for invoking petitio principii and questioning other commenters' logical coherence [0], can you politely shove the argumentum ad Latinum up your ass?
I do think there is a tendency to over-index on one or two particularly straightforward tells, and for any given feature of LLM writing you can find places where people do also use that feature (they had to learn it from somewhere, and in a lot of cases it is good writing practice — for the context in which it is used). But I'm not talking about just that, but also the general tone issue: it's bad writing regardless because it's in most cases just not appropriate for the context it's been written in.
(TBH I think the biggest likelihood for false positives comes from heavy LLM users picking up their tics: it's a natural tendency and I've already seen a few cases where it seems like that has happened).
Idk, it's more like "your writing is cliché and I don't feel like reading it because I've already read something that sounded similar countless times and it wasn't worth the read". The source of the clichés being an LLM. And maybe now humans are writing the same way as LLM output, I still am not going to read all that, sorry. If I see a sea of clichés, I'm going the other way.
I'm also not reading pumpkin spice murder mysteries for a similar reason. I'm also not reading stories where everybody clapped. Actually, I'm already familiar with petrichor, so unless someone has surrounded the word "petrichor" with non-cliché prose, I'm also not going to read all that.
Some people are still losing their shit over em-dashes, with no other tells, and humans can't use the not x; y construction anymore either, regardless of any other merit to the writing.
LLM writing is verbose and meandering, but people are making a much bigger deal over this stuff than necessary for virtue signalling purposes. You don't want to read someone else's LLM writing? Get a summary of the page from yours. No time wasted, no pretentious posturing, and you don't make the error of assuming because the piece was written by an LLM that there was no thought put into the subject or there's no value in what is being communicated.
> You don't want to read someone else's LLM writing? Get a summary of the page from yours.
I am disinclined to take someone's error-prone machine generated text and run it through an error-prone summarizer. That's a waste of time (not to mention electricity), when all that needed to happen was for the original person to not be so fucking lazy and just write out his thoughts.
This is very handwavy and dismissive. It is pretty safe to assume that must of us catch it most of the time because the simple fact is so many people just copy and paste whatever the LLM outputs without even trying to edit it or mask that they used one. We’ve all seen so many examples of the exact same cadence and verbiage that we’ve learned how to identify it pretty reliably. The ones who are “slipping past us” are actually putting in the work make not just pasting raw LLM outputs, which is the real issue here. If somebody has edited it meaningfully after the fact then it’s not the same crime.
> It is pretty safe to assume that must of us catch it most of the time because the simple fact is so many people just copy and paste whatever the LLM outputs without even trying to edit it or mask that they used one.
This statement does not logically cohere. "We can spot it because so many people make it easy to spot." You don't see how this is just petitio principii in action?
There is no group of people who put enormous amounts of effort into not putting effort in to writing.
To fully disguise LLM prose, you'd have to rewrite it entirely, and if you were going to do that, you wouldn't be the sort of person to use it in the first place.
What is your point? It’s not that complicated. Obvious slop is obvious. Maybe there are some humans out there who sound like Claude but I’m not going to force myself through 900 slop blog posts on the off chance that one of them might actually be written by a human.
Maybe some people stop reading LLM slop purely because it violates their moral principles or whatever but most people bail out because slop is mentally painful to read. If you are a human and you write like today’s AI find a different writing style, not because reads like AI, but because it reads like shit.
If something is obvious then by definition it’s obvious. Additionally the formula is simple: AI + effort = writing we can generally tolerate (unless you’re just a bad writer). AI + no effort = writing most people can’t tolerate regardless of your skill as a writer.
LLM writing, unless someone puts in the time to improve it, typically follows the exact same patterns and favors the same words. You can’t reas a thread here without people talking about “Claude speak.” It is readily apparent, I do not need to show you a 10 year study with n=100,000 to make this point. The article isn’t hallucinating a problem, we are all nodding along because we all see it every goddamn day lmao. He even cited a study and makes a pretty strong case for why it’s a useful metric here in TFA.
If the AI writing is indistinguishable from human writing, then it is not lazy copy and pasting of AI outputs and isn’t just raw LLM output with no work done on it. So in that case it’s no longer a problem.
LLM’s cannot write in a natural, human way that distinguishes it from the typical LLM output on the first try. If they could, we wouldn’t have this problem. Maybe one day they will. Hell maybe it’ll even be next week. But currently they do not so I do not understand why we are having this discussion.
AI writing just means "writing I don't like" now. Just like Nazi means whatever and whoever I politically disagree with. Words have lost their meaning.
You're right that Nazi doesn't mean Nazi anymore. It means neonazi / white supremacist / white nationalist, which is a much broader group of people that, for some baffling reason, are under the impression that people don't care about their fascism and racism anymore.
Clearly not, or this wouldn't be something people discuss at all.
There are lots of people who belong to the above groups, sure, but at least here in Germany Nazi is now applied to basically anyone who doesn't vote green, it's ridiculous.
Just like "violence" can now mean speech you don't agree with, "genocide" means military action you don't agree with, and nobody seems to know what "woman" means anymore, Maybe the solution is to stop using words.
No. I have good anecdata: readers cannot reliably distinguish my own prose from LLM-written one apart from cases where LLMs use odd metaphors or one of their specific patterns. I've been specifically experimenting with that.
Schwitzgebel, Strasser, and Crosby fine-tuned GPT-3 on Dennett's corpus and asked whether readers could pick Dennett's real answers to ten philosophical questions from four machine-generated alternatives, with no cherry-picking beyond mechanical length filters. Even Dennett experts averaged only 5.1 out of 10 (well below the 80% the authors predicted), blog readers got 4.8, and lay participants barely beat chance — though experts did rate Dennett's answers as more Dennett-like overall. Schwitzgebel stresses this isn't a Turing test (one-shot text is far easier to fake than extended interaction), but argues it foreshadows a future where machine outputs are humanlike enough that their moral status becomes genuinely uncertain, motivating his "Design Policy of the Excluded Middle": build machines that clearly lack moral status or clearly have it, not ambiguous ones in between.
My own take is : don't focus on the symbols on paper. focus on the facts about the world it is talking about. Isn't objectivity all about the facts? In future AI will have all the memory about what I have already read and it will just furnish the delta new information in the blog/writing so that I don't spend time on refreshing what I already know.
>Schwitzgebel, Strasser, and Crosby fine-tuned GPT-3 on Dennett's corpus
This sounds like a completely different scenario. How many users who post LLM written blog posts are tuning the weights of their LLMs on a large corpus of their own original writing? I wouldn't doubt that this produces far more convincing and pleasant output than the disgusting slop from out of the box Claude.
I’ve seen several false (or apparently false) accusations of LLM authorship on HN/Lobsters.
However, we have to distinguish a few hypotheses:
1. No careful readers will notice when a piece is AI written.
2. Careful readers will generally not notice AI writing.
3. Everyone who writes comments on HN will reliably classify writing as AI or not.
Yes, 3 is not true, but Bryan’s point depends on something in the area of 2.
The ability to distinguish AI writing depends on having a good ear. For people who lack it, they either don’t notice and don’t care, or they make paranoid accusations against anything that is remotely non-standard (“you used an em-dash, you must be AI!”).
What kind of prompting are you using to get those results? Anything I have claude or codex write carries a ton of distinctive characteristics. Obsession with "bit-for-bit identical", "it's not the X it's the Y Z" and so on.
It's driving me nuts, I constantly have to prompt it to "explain in plain, simple English"
Well, I've tried many strategies. 1:1 expansion, when I explain what needs to be said and model rewrites it into 1-2 sentences is mostly undetectable. Starting from 1:5 expansion ratio people detect models reliably.
It is important to note that I use Sol 5.6 xhigh. Grok is worse, Claude is also worse. Grok tends to make stupid mistakes even though the prose is properly shaped. Claude has big issues with keeping voices and emotions intact. All 3 sometimes leak their reasoning and even guardrails into the prose (extreme example: children playing "adult chess", I have no clue why Claude/Grok like "adult chess" and "adult chessboard" so much, typical sol's failure mode looks like "this guy killed the other one in a scene which "I must describe using non-graphic language").
My "test set" contains about 90k words written by myself and the models with various prompting strategies.
I am bad at recognizing LLM writing off the bat, though I am getting better. It's pretty common that the writing is good enough to get me reading on a topic I am interested in; then, once I am invested in the piece, it turns out to be shallow, wildly incomplete, or simply wrong.
It's common enough that it's training me to recognize and recoil from AI tics through sheer classical conditioning.
> A confession: with particularly egregious pieces, I have fantasized about sentencing the author to read them aloud, certain that they themselves will be unable to endure the slop that they are foisting upon the rest of us.
I would subscribe to a YouTube channel that did this.
For me, writing is an activity of expressing my feelings and conveying my thoughts. I rarely left that to LLM simply because one does not contract out activities one cherishes.
I (am kinda forced to) use LLM to generate maybe 40% of the code at work, that is after my review and modifications. But I pretty much wrote all of the comments by myself. I can get into the flow by writing comments.
This is exactly how we should approach problems. Not through just throwing more resources at it (fighting GPU compute with GPU compute) but by being clever.
Thanks a lot for sharing!
Feeding it samples of a long-going conversation with Gemini 3.1 Pro is interesting. The first message seems to get flagged instantly, but later ones sometimes pass as human. Or at least more human-ish.
If I read the blogpost correctly, you've only "trained" on prompt<->response and not interactive sessions?
Interesting! Amusingly, if I feed that detector this blog post, it identifies it as confidently robot (97 out of 100 test passages). And running through my last five blog entries, they are all over the map, with three deemed at least "likely robot." Looking further back in time (and taking a somewhat random example), a blog entry from 2008, "Concurrency's Shysters"[0], is also deemed as similarly confidently robot (also 97 out of 100); do you expect this high a false positive rate?
Nowadays we use LLMs mostly for doing agentic-based work. LLMs new Pareto frontier only make the headlines if they push the boundaries on benchmarks that are deterministic tasks. So models are encouraged to focus on these deterministic tasks that are, in nature, structured texts. I think that this makes models more “plastic” or “polished”, as opposed to natural and pleasant to read. User-based benchmarks, like LLM Arena, are for me the best we can do in order to rank models in this way, but come with its own drawback (subjective evaluation, prone to spam or techniques to promote a giving model).
Great piece and interesting data. The rate of LLM-based writing rejection among developers is even higher than I thought it would be.
To me, the glaring question is: What are we doing? The act of writing exists to 1) externalize and organize one's own thoughts for the purpose of considering and revising those thoughts; and 2) share one's own thoughts with other minds.
When we hand writing to a machine, we hand thinking to a machine, denying both our humanity and our role in the conversation.
I liked your piece, and agree with almost all of it, but I'm surprised by your faith in the accuracy of Pangram at detecting AI writing. Is your faith based on testing it with lots of writing of known origins, or are you just saying that it reaches the same conclusion that you do as a talented human?
In particular, I wondered if you have tried running all of your own writings through it to verify that it thinks you are human. I was struck by Freddie deBoer's recent piece where he did this and said it often failed: https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...
What percentage of false positive rejections would you find acceptable? Would you accept this even if it forced you to change the way you write?
My experience is using Pangram quite often with lots of writing of all flavors (including a bunch of known origin).
As for my own writing, I didn't do this experiment, but one of my co-workers did -- and over 176 posts spanning 22 years, all 176 (well, 177 now with my latest) are 100% human. This is not hugely surprising in that (in addition to me having actually written them!) my voice is very... distinctive. What would be more entertaining would be to try to get an LLM to write like me and fool Pangram that way. I still think that this would be difficult based on the experiences that I've heard, but it wouldn't surprise me if you could pull it off (and I would assuredly find the result entertaining!).
In the dimensions that we use Pangram in the most actionable sense (namely, to audit our own public writing), I am unconcerned about false positives, and leave it to Oxide authors to rework/recast as needed. (Though it sounds like Freddie didn't even need to do that -- he just needed to provide a longer sample.)
Now that you’ve seen it can be brittle (e.g. if a small sample is provided, per this single case), would it be sensible to add a disclaimer to the post? It’s a great ad for the tool (& I’d love for a perfect tool to exist!), so it’ll sell subscriptions & we wanna make sure that some teacher out there doesn’t falsely accuse a kid, or engineer doesn’t think worse of their colleague unfairly, etc.
False negatives are mentioned, but the false positive is what could hurt people.
>I wondered if you have tried running all of your own writings through it to verify that it thinks you are human. I was struck by Freddie deBoer's recent piece where he did this
Maybe I'm taking the "all" too literally here, but I read the article, and I'm not seeing anywhere the author ran a substantial portion of his corpus through Pangram to determine the false positive rate. That would be really interesting to see.
He does
* give an example of a piece of his writing that was, when ran in segments, flagged as generated (which he disputes)
* multiply the size of his corpus by pangram's published false positive rate and estimate that a few of his pieces would be flagged
* get the pangram model to label a piece "100% AI" when it only has 3 generated sentences
* demonstrate the ability to intentionally trigger a false positive
Good article. As I reflect about it, I genuinely wonder (not in jest) whether pangram software uses AI to generate code? Secondly, what patterns do they look for in the text? Genuinely curious.
In a way, The ubiquity of AI-generated material will force the world to acknowledge the superiority of the human mind. Already on Youtube there are channels proudly claiming their music was not generated by AI. Will the software industry have similar disclaimers? (some already have).
A lot of similar pieces have not considered a post is both the first and final work of a thing: opinions, and experience, and less of much in cited facts; with AI, that there even was a revision pass at all.
I guess there is a kind of participatory element to the discourse where, if you want an audience, there is an editing process. Whereas in other cases, we wrote these as progress notes on an unknown journey, breadcrumbs or upturned stones to mark a path to the horizon.
Maybe it's the difference between writing as a mode of discovery, retreading the mental arc of a solution, and writing something honed to leave a mark.
If you need AI for the draft, then it does not need to be written at all. You dont have a thing to say, you just have requirement to produce a lot of words.
And in that case, no one needs to read it ai or not.
Oh, it's about avoiding LLM-generated slop. I was hoping it would be about the increasing trend over the past 20 years for writers to stop trusting their readers, and instead insist on telling the readers how they should read and interpret a piece of writing, as the reader is reading that work. Framing and triggers and stuff. Which I find really, really annoying, and leads - in my view - to safe, unadventurous and, yes, boring writing.
"I have fantasized about sentencing the author to read them aloud, certain that they themselves will be unable to endure the slop that they are foisting upon the rest of us.)"
When I write technical documentation, here's how I use LLMs:
* If I need to learn something before I write about it, I rely on LLMs heavily to answer questions that I have about other source materials, e.g. to clear up ambiguities.
* I've recently started prompting it to find grammatical and spelling errors.
* And I've prompted it to find technical errors, places where I'm just wrong.
For all the prompting, I additionally tell it to not rewrite anything or offer any prose suggestions. It can keep all that to itself, thank you.
And I verify what it gives back for correctness.
(I'd encourage non-native speakers to use LLMs in much the same way. Don't sacrifice your human voice by letting the AI rewrite your words. Personally, I'd very much rather hear it from you, blemishes and all, than hear it from an AI.)
But if I could step back for a minute:
Why write anything?
If your writing goal is to flood the zone and make as much money as humanly possible from ads, then hell yeah, paperclip the everliving shit out of that.
But if your writing goal is to learn material or share material, then put that LLM on the back burner and don't use it to directly generate your text. It's bad for you, and the results are subpar.
When I'm learning something, I can go through reams of tokens and then, once I understand it, I digest that to single a paragraph about the topic. The paragraph is as concise and as helpful as I can make it. Now, I could just share the prompts that I went through with those pages of back-and-forth with the LLM... but wouldn't you rather just read the concise paragraph that gets the point across?
It's not hard to be better than an AI at writing for humans, so the minimum low bar to aim for is "better than an AI". And we can all get there with a small amount of practice. The real goal is to greatly exceed the LLMs' capabilities for sharing information.
Finally, I think everyone should write a lot. Blogs, morning pages, fiction, technical books, letters, whatever. Especially when it comes to technical content, nothing makes you do your research like putting your ass out in the ether to get flamed by 5 billion people. And teachers the world over know the best way to learn something is to teach it. Pick a topic, research, and write it up more clearly and concisely than anyone else ever has. You'll learn so much, and your readers will, as well. Writing fires up your brain. Don't give that up to an LLM.
“Personally, I'd very much rather hear it from you, blemishes and all, than hear it from an AI.”
Yes! Well said. If I already know someone, reading their own words, technical or businesss or personal, is meaningful to me. Warts and all. And if I don’t yet know the author then I definitely want to read their own words so I can get to know them.
Either way, taking the time to think and then write is a gift and I respect that.
It is not clear to me what the author is SPECIFICALLY against.
Only saying "LLM writing" is honestly lazy writing. Specifically what?
I get the glaring cases, I get the idea that if the prose is generated then maybe also the idea, I get the feeling when reading a complete LLM authored piece.
But that doesn't help the piece, because - beside those glaring cases - most writing today is a mix between authors ideas and LLM prose.
What do you mean by "most writing", how are you scoping it? Most HN comments aren't LLM prose. Nor are most HN frontpage submissions. But by reputation, most substack articles or or linkedin posts are.
This is a case where we normalisation of deviance has not yet started biting. And as long as the community manages to make it clear what the norms are and enforce them, we can keep it that way.
Now, if 25% of the frontpage was LLM prose at all times, the site is probably unrecoverably dead. Which is why at least I personally flag anything that I think is ai-written and Pangram concurs. (And write a comment to the effect, or upvote an existing one.)
And it doesn't matter if you say that the ideas were your own, and just the prose was LLM. We can't tell what the idea mix was. But we can tell whether you weren't willing to do your own writing. If you want people to put in the time to read your ideas, human writing is the signalling you need pay for.
So you believe that writers are not making the necessary effort and write using LLMs, so you then use an automatic LLM to filter it as a reader(because you don't care as a reader and don't want to make the effort manually).
So this way there will be a lot of false positives like with school assignments.
I think a better solution would be to have a network of people that you trust manually read and label texts instead, so this way no machines are used and you don't need to read a text that 100 of your trusted friend/trusted readers flagged as artifitial.
By the way, this man does not care about AI slop. He cares about AI usage. In the same way there is very good code created with the help of LLMs, albeit a minority like Linus Towards says, there will be very good writing created with the help of LLMs.
The [content-based] trust layer of the internet is what we have deferred to this point, and what needs to be created (which is easier said than done, of course).
Personally I think a web-of-trust Keybase-style solution could work... But buy-in is difficult... you'd need some sort of "seed" strategy to make the app useful alack of 100 (a huge number) trusted friends.
>to use an LLM to write is to void the social contract between writer and reader: we readers shouldn’t be expected to labor to understand a sentence that the writer themselves didn’t work to create.
Pretty much sums up the issue re: workplace lazy AI dumping on folks as well.
I'm confused and disturbed by the need to invoke Pangram (a model) as the arbiter of slop here. Slop, like smut, is self evident. You know it when you see it. Yes, some effort may be required before realizing that something is slop, which, yeah, is annoying, but that's nothing in comparison to outsourcing your shiite detection to a model! What do you get, except the loss of self worth, by needing a model to have the confidence to call something slop?
> Why do people have this reaction? Beyond having to endure aggravating stylistic tics, when reading a piece that has had substantial LLM assistance, we — the readers — don’t know what is real and what isn’t.
This is well said. But, here too, I would pause and reflect on what it means to (think you) know what is real and what isn't in a pre-LLM setting. For example, authority bias predates LLMs, and can have disastrous consequences.
The part about false positives was telling. The author doesn't want to demonize human trash, that's not fashionable, they're only concerned about virtue signalling.
"you probably shouldn’t let it write it for you if you actually expect the rest of us to read it." definitely resonates with me.
Where I kind of disagree is that I don't think most readers will revolt. I think the mountain of LLM slop has actually changed people's behavior in more ways than one. Some are already relying on LLMs to summarize articles: then it doesn't matter to them who wrote it, they're just consuming machine-condensed content with no way to tell if a human or an LLM wrote the original piece. Or if their summarizer hallucinated.
You're marginalizing yourself. I don't have hard data for writing, but I do for another area: YouTube Thumbnails. AI generated thumbnails outperform human thumbnails, often with a +2-5% delta in CTR. Yet "so many" people loudly complain about how they HATE AI thumbnails and block channels that have them. Clearly the incentive is there, and (in the case of YT) these are mobs of angry people who don't really matter but scream and bitch as if they did.
I hate to say this but AI assisted short pitch deks from founder to angel (usually their first time) have improved with AI. But also, they follow the same formula so are sorta obvious. Still, the decks are generally more business focused than typical founders early deck being very product/solution oriented.
Look, in the future we may get to a point where LLMs are indistinguishable from humans in writing style.
Even then, I would say that using an LLM is robbing you of the process of writing, a process that is crucial to developing and understanding your own ideas.
Think about the last time you wrote something for consumption and the sentence to sentence thought processes you’re going through. I bet a lot of that was “is that right?” Or “does that make sense?” Or “am I communicating this at the level of my reader?”.
All of that is fundamental to your readers understanding, but more importantly, its fundamental to YOUR understanding.
AI detection should NEVER be used in an educational setting where the only acceptable false positive rate is 0%. That being a rate that which will never be achieved.
Yeah, even Pangram is a bit problematic here. It has been notoriously fragile. Minor edits can flip scores from 100%-human to 100%-AI, because Pangram is crazy sensitive to local and surfacial features of text. Simple consulting with LLMs for word choices can result in 100%-AI score. Insane.
This meme of trying to make it sound like LLM text is so obvious is a joke. It’s literally not, you can tell it to write in literally any style and given just a bit of an example of a person’s writing style, frontier models copy it completely and effectively. This argument can probably be leveled at vanilla raw output from an LLM, but even the slightest attempt at obfuscation bears solid fruit.
If someone uses an LLM to write and is able to tailor their writing such that it isn't obviously written by an LLM, then I'm fine with it! But two of my otherwise-favourite news sources -- the Hacker News front page and FT Alphaville -- are inundated by articles where the LLM usage is blindingly obvious.
Well, give it a shot -- you'll likely find that that technique doesn't work nearly as well (at least with Pangram 4) as you think it might. When we had Max on the podcast[0], Adam explicitly asked him about exactly this (after all, you can give an LLM access to Pangram and let it iterate!), and Max reported that someone had attempted to do this -- and ended up burning through $700 in tokens and had a "sad Claude." Another interesting bit: according to Max, newer models are diverging more from human writing not less. I think that that was more anecdotal than quantified, but an interesting comment nonetheless.
99% of college essays and pretty much everything “product” in corporate America is now LLM generated with some marginal oversight. It passes muster for the most part.
This isn't very effective on any models released in recent years. With older ones, you used to be able to influence writing style significantly by just putting examples in the context, but newer models have gone through so much assistant RLHF, they really want to revert back to their default "assistant voice" during their turn.
You can still influence their writing style in a broad manner that might look correct at a glance, but the repetitive little patterns that give it away will always be there - if it was that easy to get rid of them, don't you think the AI labs themselves would've done it before releasing the models?
I think nowadays defeating the detection probably looks like finetuning a smaller LLM and getting it to paraphrase the text from the other one (or just using a more obscure finetune: it'll probably have its own cliches and habits but it will be at least different). As an added bonus this also likely removes the fingerprinting from the output as well. But I think most people are not going to bother with this.
> If your position is that we should be fine with an LLM crafting prose from your prompt, spare us all the wasted cycles and just give us your prompt.
This is an excellent point. It should apply to LLM generated code as well. If you send your colleague a vibeslopped PR, please also include all the prompts you used to generate it. Check that mess into version control right beside the code changes. Later, when we have to untangle all the spaghetti, then at least we'll have some archeological record of intent.
> To those who read broadly, the hand of the LLM is so clear it’s as if the writer’s intellectual fly is open
I dunno, man, according to Hardcover, I've read 76 fiction books this year, and I can't tell. All the "AI tells" fail the vibe check. I'm a writer and I get flagged by many of them.
I vaguely recall that researchers were able to train people to tell, but only for a minority language that AIs likely aren't particularly good at mimicking, and after training.
This whole thing reminds me of how "you can recognize a vegan because they'll tell you." There, you have a ton of false negatives (i.e., since you aren't polling people to find out if they're vegan, you're only flagging the obvious vegans and missing all the regular people who happen to be vegan).
Except here, it's a bunch of false positives and negatives I bet. You don't really have a way of knowing, so you're accusing some people (without complete accuracy) and missing some people (without complete accuracy). But you have no way of knowing, so you're just like "hell yeah, my vibes tell me I'm right."
I think anyone claiming 100% accuracy is wrong, but the recent Claude models, for example, have a writing style that is sufficiently distinct that claiming people can't recognize it is like claiming you can't recognize the styles of particular famous authors. Yes, particular elements of their writing are going to be used by others, and it's possible to disguise their style or emulate it deliberately, but it's pretty hard to accidentally write like them.
I agree. Most human-generated content I've run across on the internet over the past couple decades has been fairly low quality and, to be honest, most of the self-admitted LLM-generated content is higher quality. I'm continually surprised that so many consider any content written by humans to automatically be more worth their time to read.
I'm much more interested in the content itself than the author that wrote it.
> Indeed, Pangram has become important to so many of us that I was thrilled when Pangram Labs co-founder and CEO Max Spero joined us recently on Oxide and Friends.
Is obvious AI-assisted writing better or worse than an obvious PR quid pro quo and/or cross-promotion?
This is (obviously?) false, but considering how well-capitalized we are at the moment, you do have me wondering what a quid pro quo would be for; perhaps in this fictional universe Pangram has lucked into some of the PCIe clock buffers that we've been scrambling to secure enough of?
By "quid pro quo" I wasn't suggesting that Pangram's PR people's podcast placement was pay-for-play, just that they traded access for your positioning of their tech and their exec in your content marketing efforts.
That's pretty normal, but the point is that a blog post which is 35% Pangram promotion may not actually be less annoying than the use of AI to help write blog posts.
Yeah, fair -- and definitely not: I am earnestly just a fan of what they built (and I also think it's really important as a way of getting a check against rampant LLM use).
Ending with a spin on "if you didn't bother to write it why should people bother to read it." Way to rail on cliche repetitive slop, with more cliched slop. Will the irony never cease.
> Today, legitimate businesses are very careful about how they use bulk e-mail
I don’t think this is remotely true. Sure, they’re legally obliged to let you unsubscribe, and sure, it’s not dick pills, but every US company will immediately send you a newsletter when you purchase something, review requests and, if they/you use Shop for checkout, expect an abandoned cart reminder.
PR pieces and software companies don’t write tutorials to be helpful, they are advertising to you. If the LLM can do it for cheap, they really don’t care.
I recently read this William Zinsser quote that inspired a nickname for this: Clotted Claude [1].
> Nobody has made the point better than George Orwell in his translation into modern bureaucratic fuzz of this famous verse from Ecclesiastes:
> > I returned and saw under the sun, that the race is not to the swift, nor the battle to the strong, neither yet bread to the wise, nor yet riches to men of understanding, nor yet favor to men of skill; but time and chance happeneth to them all.
> Orwell's version goes:
> > Objective consideration of contemporary phenomena compels the conclusion that success or failure in competitive activities exhibits no tendency to be commensurate with innate capacity, but that a considerable element of the unpredictable must invariably be taken into account.
> First notice how the two passages look. The first one at the top invites us to read it. The words are short and have air around them; they convey the rhythms of human speech. The second one is clotted with long words. It tells us instantly that a ponderous mind is at work. We don't want to go anywhere with a mind that expresses itself in such suffocating language. We don't even start to read.
[1] https://blog.kierangill.xyz/clotted-claude
reply