Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>>The researchers note that there’s a problem with this argument, too, as it violates the likelihood principle. This principle tells us the interpretation should only rely on the actual data observed, not the context in which it was collected.

and then in the publication itself:

>>The likelihood principle [Edwards et al., 1963] is a fundamental concept in Bayesian statistics that states that the evidence from an experiment is contained in the likelihood function. It implies that the rules governing when data collection stops are irrelevant to data interpretation. It is entirely appropriate to collect data until a point has been proven or disproven, or until the data collector runs out of time, money, or patience

Surely there is a difference when you look at someone who played 46 games online in his life and scored 45.5 and when you look at someone who played 46000 games and scored 45.5/46 once.

The difference is that Kramnik wasn't "collecting the data" but looked at the whole Nakamura's playing history and found a streak.

Another example would be looking at coinflips and discarding everything before and after you encounter 10 heads in a row to claim you have solid evidence that the coin is biased.

They are misapplying the principle here. If what they wrote was correct then someone claiming: "Look, Nakamure won 100 out of 100 if you just look at games 3, 17, 21, 117...." would be proving Nakamura cheated if they applied methodology from the paper even assuming one in 10000 guilty players. Just because you can choose sampling strategy and stopping rules (what the likelyhood principle states) doesn't mean you can discard data you collected or cherry pick parts that support your hypothesis.

How the data is collected is absolutely relevant and Nakamura is right to point it out.



General statistical question. If we say extend the coin flip example distribution to say 10B times. Should/would we expect to see a streak of 100 or even 1000 in the distribution somewhere? Intuition alone tells me probably not for 1000 but a smallish chance for 100 (even if 10B in a row i would think a streak of 100 would be unlikely)


Your intuition's not bad. The expected value for the longest run of heads in N total flips of a fair coin is around log2(N) - 1 with a standard deviation that's approximately 1.873 plus a term that vanishes as N grows large. log2(10B) - 1 is approximately 32 and with that standard deviation, even a run of 100 in 10B flips is incredibly unlikely. For more info see Mark F. Schilling's paper, "The Longest Run of Heads" available here https://www.csun.edu/~hcmth031/tlroh.pdf.


Neat! I guess this is a common thing to wonder about :)


That’s a cool result, thanks for the link!




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: