Hacker Newsnew | past | comments | ask | show | jobs | submit | atomic128's commentslogin

There is a large community of people that poison scrapers.

The poison gets better every day, and the community is continuously growing. Poison Fountain, alone, transmits hundreds of gigabytes of poison per day, which goes into scrapers, git repositories on every hosting platform, social media, etc.

Part of the poisoning community on Reddit, for example: https://www.reddit.com/r/PoisonFountain/comments/1uocaii/a_n...


I've banned this account because we don't allow single-purpose accounts on HN, and your account has been doing that for quite some time now.

We ban such accounts regardless of what the single purpose happens to be. Pre-existing agendas are not what HN is for and destroy the curious conversation that it is supposed to be for.

https://news.ycombinator.com/newsguidelines.html

Edit: If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.


Just curious dang, did you warn them before banning?

Im not against the ban perse (single purpose accounts are bad), just curious if they had a chance to change their contribution style.


No, but if they have a change of heart and genuinely want to use HN as intended, they're welcome to let us know. I've added that in an edit now.


Seriously, dang?

10 comments (excluding subsequent in-thread replies) over four months, always in contexts in which either the topic of LLM scraping or Poison Fountain itself has already been mentioned.

This strikes me as contextually informational, and is no different from other project representatives appearing in threads discussing their own subjects or posts. Such as, say, Jon Corbet (@corbet), of LWN, whose activity on HN shows a similar pattern and roughly equivalent frequency.

I hope it goes without saying I'm not suggesting corbet's handle be banned, anything but.

atomic128's comments are predictable, but apposite, informative, non-disruptive, and address an increasingly urgent issue. Whether or not it's an effective mitigation is of course another discussion, but it seems plausible at first blush.

As dang should well know but others may not, I often contact mods directly for HN issues, including numerous "one-note flute" alerts. atomic128's account should be un-banned, though perhaps they might communicate with HN's mods over what would be a more acceptable mode of interaction.


The most recent 60 (!) comments plus every submission of the last 6 months were all about the same thing. That's extreme. The posts didn't all mention that specific project, but there was only one topic and they were extremely repetitive. This is not a close call.

I made it all the way back to https://news.ycombinator.com/posts?id=atomic128&next=4628060... (6 months ago) before seeing posts about anything else, only to find that there was a different agenda before that. Not cool.

Edit: and before all that, there was this: https://news.ycombinator.com/posts?id=atomic128&next=4164795.... This is obviously not using HN as intended.


Exceptions?

<https://news.ycombinator.com/item?id=47116093> (LLM but not PF).

<https://news.ycombinator.com/item?id=47095664> (a16h)

<https://news.ycombinator.com/item?id=46695693> (vuln exploits) 2026-1-20

<https://news.ycombinator.com/item?id=46280602> (???, but not PF)

<https://news.ycombinator.com/item?id=46195234> (Monero / Dark Web) 2025-12-8

<https://news.ycombinator.com/item?id=45894305> and <https://news.ycombinator.com/item?id=45826273> (Tor hidden service) Nov 2025

That's from the past 20 comments.

Submissions: 8 most recent on PF, 9+ cover nuclear power, Tor dark web, robotaxis, and other topics.

Again: Not a one-note flute, though fairly focused of late on AI and poisoning.

Again: I think the ban is unwarranted. I'm not sure what's driving your thinking here, but a no-warnings ban seems excessive. And given YC's current preponderance of AI/agentic launches (<https://news.ycombinator.com/launches>), self-serving and contrary to the "we moderate YC stories less" guideline.

(Yes, I'm aware "less" isn't "none", and this is an account/user rather than story, I hope my point stands and is clear.)

The Yann LeCun posts you link are ... a bit OTT. That's also a couple of years ago.

I've said my bit. I'm hoping you and atomic128 can come to an understanding in email.


> Such as, say, Jon Corbet (@corbet), of LWN, whose activity on HN shows a similar pattern and roughly equivalent frequency.

I took a look at the most recent comments from both accounts and they don't look similar to me in this respect.

I think there are two questions here though:

1. Was the violation egregious?

2. Did it deserve an immediate ban, or did they deserve a warning etc.?

Seems to me the answer to (1) is yes, but the answer to (2) I'm less sure about.


Jon's been around a while and some of the piss and vinegar of youth may have subsided. He does tend to show up with LWN comes up, whether as a topic of discussion (or more often) from submitted articles. That's his baliwick, and again, I don't fault him at all for it.

Our other friend here is a more recent participant to HN, at least under this handle. (I don't know that there are others, only what I can see from this one.)


I think the reasoning is about having alt accounts for different purposes. He intention is to map one human to one account and have all of their thoughts from that one account, instead of one human having one account to discuss scraping on, and a different account to discuss crypto on.


I'm pretty confident it's not that.

HN's prime directive is "anything that gratifies one's intellectual curiosity": <https://news.ycombinator.com/newsguidelines.html> and many, many, many dang comments.

I'm pretty sure that the specific gripe is posting excessively (not even necessarily exclusively) on a single topic or theme. See <https://news.ycombinator.com/item?id=19392902> for a more detailed comment from dang.

Occasional alts are explicitly permitted, though not to engage in abuse (e.g., mutual admiration societies, sock-puppetry). See: <https://news.ycombinator.com/item?id=9963551> <https://news.ycombinator.com/item?id=9823379> (both against sock puppetry) and <https://news.ycombinator.com/item?id=9122086> and <https://news.ycombinator.com/item?id=7504621> (on where throwaways are/aren't permitted).

Where HN does favour persistent accounts the stated claim is to foster community, rather than for nefarious tracking purposes: <https://news.ycombinator.com/item?id=18082346> and <https://news.ycombinator.com/newsguidelines.html>. From that last:

Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.


I'm confused why everyone is pretending that a ban holds any meaning here. he probably already has a new account, it might have made sense to just silently ban him in the hope of imposing a minor cost on what you believe he is doing, but wasting your time addressing it imposed a far greater cost on you than him.


I understand how it can be confusing. The key factors in doing it this way are (1) the community regards older accounts, especially ones that have significant posting history, as more credible; and (2) doing it publicly rather than silently has transparency value.


This is a strong positive sign that poison fountain works.

I wasn't aware of this project. Thanks for the heads up.


It's not a sign about that project in any way. I had never heard of it and have no opinion about it one way or the other.

It's just a sign that single-agenda accounts aren't allowed here—no more, no less. That's why I said "We ban such accounts regardless of what the single purpose happens to be".


People think this is causing issues for data collection for LLMs, but in reality it's not and there are several very trivial mechanisms to employ in data collection to bypass the "poison data" issue. The internet landscape was already poisoned with fake data, fringe conspiracies, and text before this Poison Fountain initiative.


Yeah. A fun thing to do is to try and actually read common crawl!

Really makes you think, what we're feeding them...


exactly i took a look at that subreddit and doesnt look like theres any professionals just bunch of anti-AI users who thinks they are smarter

its very easy to detect and bypass poison type of tools largely because of the fact that there are far more outlets for truthful info so unless you can get everyone to buy in (with real legal liabilities) its not effective

also its possible to poison the poisoners with a certain pill that would have very real consequences for those maintaining whatever github repo/communities


Poison Fountain on Reddit: https://www.reddit.com/r/PoisonFountain/


Do you feel poison fountain is actually effective? To me it seems like free chaos engineering for ingestion platforms. Wouldn't it paradoxically harden data ingestion?


Not sure about flawed code logic, but it is embarassingly easy to plant false information into models. Make a few static sites with random info, crosslink them, reference them on reddit a few times, then plant the payload there.

I know it because i tried ...



I have found my people. Thank you for sharing.


I would advise against using it. The code it returns comes from real public repos, so including it in your work could lead to copyright issues. You'd probably be better off asking an LLM to come up with gibberish.



Rumors that Anthropic is in talks to buy Atlassian, presumably for the training data. Data poisoning efforts are underway: https://www.reddit.com/r/PoisonFountain/comments/1sqrq24/atl...


I know at least two companies that won't be able to use Atlassian products anymore if that's the case. They really don't give a shit about privacy and regulatory requirements.


github etc hold source code -> scraped -> so AI may generate any of that.

And the specs become the new source (code).

fast forward..

Atlassian etc hold source specs -> scraped -> so AI may generate any of that.. then any of above..

the new source would be (?what? company missions? get-rich-quick-schemes?)

fast forward..


Hmm, if the stock keeps falling that might really happen.


The war is already underway.

Poison Fountain: https://www.reddit.com/r/PoisonFountain/


Poison Fountain: https://rnsaffn.com/poison2/

Poison Fountain explanation: https://rnsaffn.com/poison3/

Simple example of usage in Go:

  package main

  import (
      "io"
      "net/http"
  )

  func main() {
      poisonHandler := func(w http.ResponseWriter, req *http.Request) {
          poison, err := http.Get("https://rnsaffn.com/poison2/")
          if err == nil {
              io.Copy(w, poison.Body)
              poison.Body.Close()
          }
      }
      http.HandleFunc("/poison", poisonHandler)
      http.ListenAndServe(":8080", nil)
  }
https://go.dev/play/p/04at1rBMbz8

Miasma Poison Fountain Tar Pit: https://github.com/austin-weeks/miasma

Apache Poison Fountain: https://gist.github.com/jwakely/a511a5cab5eb36d088ecd1659fce...

Nginx Poison Fountain: https://gist.github.com/NeoTheFox/366c0445c71ddcb1086f7e4d9c...

Discourse Poison Fountain: https://github.com/elmuerte/discourse-poison-fountain

Netlify Poison Fountain: https://gist.github.com/dlford/5e0daea8ab475db1d410db8fcd5b7...

In the news:

The Register: https://www.theregister.com/2026/01/11/industry_insiders_see...

Forbes: https://www.forbes.com/sites/craigsmith/2026/01/21/poison-fo...

On Reddit:

https://www.reddit.com/r/PoisonFountain/


Poison Fountain: https://rnsaffn.com/poison2/

Poison Fountain explanation: https://rnsaffn.com/poison3/

Simple example of usage in Go:

  package main

  import (
      "io"
      "net/http"
  )

  func main() {
      poisonHandler := func(w http.ResponseWriter, req *http.Request) {
          poison, err := http.Get("https://rnsaffn.com/poison2/")
          if err == nil {
              io.Copy(w, poison.Body)
              poison.Body.Close()
          }
      }
      http.HandleFunc("/poison", poisonHandler)
      http.ListenAndServe(":8080", nil)
  }
https://go.dev/play/p/04at1rBMbz8

Miasma Poison Fountain Tar Pit: https://github.com/austin-weeks/miasma

Apache Poison Fountain: https://gist.github.com/jwakely/a511a5cab5eb36d088ecd1659fce...

Nginx Poison Fountain: https://gist.github.com/NeoTheFox/366c0445c71ddcb1086f7e4d9c...

Discourse Poison Fountain: https://github.com/elmuerte/discourse-poison-fountain

Netlify Poison Fountain: https://gist.github.com/dlford/5e0daea8ab475db1d410db8fcd5b7...

In the news:

The Register: https://www.theregister.com/2026/01/11/industry_insiders_see...

Forbes: https://www.forbes.com/sites/craigsmith/2026/01/21/poison-fo...

On Reddit:

https://www.reddit.com/r/PoisonFountain/


Probably not.

We have witnessed, over the past few years, an "AI fair use" Pearl Harbor sneak attack on intellectual property.

The lesson has been learned:

In effect, intellectual property used to train LLMs becomes anonymous common property. My code becomes your code with no acknowledgement of authorship or lineage, with no attribution or citation.

The social rewards (e.g., credit, respect) that often motivate open source work are undermined. The work is assimilated and resold by the AI companies, reducing the economic value of its authors.

The images, the video, the code, the prose, all of it stolen to be resold. The greatest theft of intellectual property in the history of Man.


The greatest theft of intellectual property in the history of Man.

Copyright was always supposed to be a bargain with authors for the ultimate benefit of the public domain. If AI proves to be more beneficial to the public interest than copyright, then copyright will have to go.

You can argue for compromise -- for peaceful, legal coexistence between Big Copyright and Big AI -- but that will just result in a few privileged corporations paywalling all of the purloined training data for their own benefit. Instead of arguing on behalf of legacy copyright interests, consider fighting for open models instead.

In a larger historical context, nothing all that special is happening either way. We pulled copyright law out of our asses a couple hundred years ago; it can just as easily go back where it came from.


>If AI proves to be more beneficial to the public interest than copyright, then copyright will have to go.

Going forward? Okay, sure. But people created all of the works they created with the understanding of the old system. If you want to change the deal, then creators need to know that first so they can decide if they still want to participate

Allowing everyone to create everything and spend that labor with the promise of copyright, and then pull the rug "oops this is just too important" is not fair to the people who put in that labor, especially when the people redefining the arrangement are getting 100% of the value and the creators got and will get nothing


Life isn't fair, and 100+ year copyright terms enforced eternally with unbreakable DRM sure as hell aren't.

But open-weight LLMs are a pretty decent compromise.


There is one missing factor in your argument. The wealth transfer. The public was almost never the beneficiary of copyright and other IPs. Except perhaps its earliest phases where the copyright had a strict term limit, it was always the corporations who fought for it (Disney being the most infamous), using it to prevent the public from economically benefitting from their work almost forever.

And then people found a way to use the same copyright law to widely distribute their work without the fear of losing attribution or being exploited. Here comes along LLMs that abuse the 'fair use' argument to break attribution and monetize someone else's work. Which way does the money flow? To the corporations again.

IP when it suits them, fair-use when it benefits us. One splendid demonstration of this hypocrisy is how clawd and clawdbot were forced to rename (trademark law in this case). By twisting and reinterpreting laws in whatever way it suits them, these glorified marauders broke a trust mechanism that people relied on for openly sharing their work.

It incentivices ordinary people to hide their work from public. Don't assume that AI is going to solve that loss. The level of original thinking in LLMs is very suspect, despite the pompous and deceitful claims by its creators to the contrary. Meanwhile, the lack of knowledge sharing and cooperation on a global scale will throw civilizational growth rate back into the dark ages. Neither AI, nor corporations are yet anywhere near the creativity and original thinking as the world working together. Ultimately, LLMs serve only the continued one-way transfer of wealth in favor of an insatiably greedy minority, at the cost of losing the benefit of the internet (knowledge sharing) and an enormous damage to the environment - all of which actively harm the public.


Ultimately, LLMs serve only the continued one-way transfer of wealth in favor of an insatiably greedy minority

Including the ones I can run on my own PC at home? I couldn't do that before. Maybe I'm the greedy minority, but I'm stronger and (at least intellectually) wealthier than I was before any of this started happening.

Qwen 3.5, which dropped yesterday, is a genuine GPT 5-class model. Even the ones released by US labs such as OpenAI and Allen AI are legitimate popular resources in their own right. You seem to feel disempowered, while I feel the opposite.


Yes, even the ones you can run on your system. They're no different from proprietary OS and software you used to run on your system, whose design in which you had no say whatsoever. These 'free to run' models are hardly open source. You don't have the data that was used to train them. It's not just about the legality of those data. The dataset chosen may have extreme bias that you can never eliminate satisfactorily from a trained model.

As if that wasn't bad enough, these models cannot be trained on your regular home computer. But instead of striving to improve the energy efficiency of these models, those big corporations build and run massive gas guzzling data centers to train them. They ruin the quality of life for the neighbors through pollution, water depletion and electricity price rise. It also disproportionately affects the poor in the world by reducing supply of essential computing components like RAM (which are needed for medical devices, utility and manufacturing installations and every other aspect of modern life), and by aggravating the climate crisis, whose victims are the poorest.

They don't give you those models out of the goodness of their hearts. Those are just advertisements and trial pieces for their premium services. They also peddle the agenda of its creators. So yes, those models are empowering only in a very narrow sense without any foresight. They are still the money making engines for the rich that subject you to their benevolence, whims and fancies.


    Once men turned their thinking over to machines
    in the hope that this would set them free.

    But that only permitted other men with machines
    to enslave them.

    ...

    Thou shalt not make a machine in the
    likeness of a human mind.

    -- Frank Herbert, Dune


Eh, we already have a name for the concept of living by plausible-sounding works of fiction: religion.

Yet another post who misses (or chooses to overlook) my point: this stuff is running on my machine. "Seizing the means of production" means going into my back room and pulling a computer out of a rack.


Alibaba (China) thinks for you. They control you, to some extent.

Wikipedia: "Qwen (also known as Tongyi Qianwen, Chinese: 通义千问; pinyin: Tōngyì Qiānwèn) is a family of large language models developed by Alibaba Cloud. Many Qwen variants are distributed as open‑weight models under the Apache‑2.0 license, while others are served through Alibaba Cloud. Their models are sometimes described as open source, but the training code has not been released nor has the training data been documented, and they do not meet the terms of either the Open Source AI Definition or the Model Openness Framework from the Linux Foundation."


Oh, no

The Linux Foundation is coming for me

Well, anyway, where were we


This isn't a hypothetical or fictional problem. This is a well-known and well-warned problem that we already see in action. How many pro-China biases have the Chinese models show? How often does Groq do whatever it wants? (Including calling itself Mecha-Hitler and undressing people, including minors for fun!) How many times have nearly every model taken pro-oligarch stances (eg: refusing to draw Mickey Mouse even after its copyright expired.) How many people, including kids were driven to suicide by some of the models?

There is no end to the examples of how it harms ordinary people. And yet, you decide to just hand wave away those concerns as if those don't exist for you or the others. There is no debate when all you do is ignore the counter arguments. It's like those science deniers who stick to their beliefs, no matter how much evidence is presented.


Don't get me wrong, I'm interested in the Chinese models only to the extent that their weights are available. I hope DeepSeek 4 sees the light of day on HuggingFace, but a lot of wealthy peoples' oxen are being gored and I suspect that it'll be the last we get if it is released at all.

If I want to see Mickey Mouse or any number of copyrighted Hollywood figures, Z-Image Turbo and HunyuanImage-3 will gladly oblige. The Chinese models may be biased to deny Taiwanese self-rule, and they may change the subject when you ask about the Tiananmen Square massacre... but they do work, and as of the Qwen 3.5 release they work well enough to be used by people at home who don't have a rack of H200s in the basement.

The most important thing about the Chinese models is that they will still be there on my hard drive 20 years from now. No additional censorship beyond what they shipped with, which (being a Westerner) is largely in areas I don't care about. No rug pulls, unwanted updates, usage limits, or price increases. No ablation of whatever subjects are deemed politically incorrect in the future. No ads. No spying. No realignment with the sayings of Chairman Musk.

As for suicide, that is a silly mediagenic exercise in blaming inanimate tools for the actions of mentally-ill people and the inaction of negligent parents. I don't consider it a valid or relevant counterargument, so yes, I'm going to hand-wave away your concerns in that area.


We have dozens of proxy sites and add new sites every day.

But your caution is healthy and it's ok if you don't particiate. Cheers.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: