Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Someone better than me at math should say how many votes we need to have a 95% accuracy.


The key issue is being good at statistics, not math. The data can't be relied on to make such an inference because they are not from a random sample of the relevant population.

VOLUNTARY RESPONSE POLLS

One professor of statistics, who is a co-author of a highly regarded AP statistics textbook, has tried to popularize the phrase that "voluntary response data are worthless" to go along with the phrase "correlation does not imply causation." Other statistics teachers are gradually picking up this phrase.

[quote=Paul Velleman]

-----Original Message----- From: Paul Velleman [SMTPfv2@cornell.edu] Sent: Wednesday, January 14, 1998 5:10 PM To: apstat-l@etc.bc.ca; Kim Robinson Cc: mmbalach@mtu.edu Subject: Re: qualtiative study

Sorry Kim, but it just aint so. Voluntary response data are worthless. One excellent example is the books by Shere Hite. She collected many responses from biased lists with voluntary response and drew conclusions that are roundly contradicted by all responsible studies. She claimed to be doing only qualitative work, but what she got was just plain garbage. Another famous example is the Literary Digest "poll". All you learn from voluntary response is what is said by those who choose to respond. Unless the respondents are a substantially large fraction of the population, they are very likely to be a biased -- possibly a very biased -- subset. Anecdotes tell you nothing at all about the state of the world. They can't be "used only as a description" because they describe nothing but themselves.[/quote]

http://mathforum.org/kb/thread.jspa?threadID=194473&tsta...

For more on the distinction between statistics and mathematics, see

http://statland.org/MAAFIXED.PDF

and

http://repositories.cdlib.org/cgi/viewcontent.cgi?article=10...


Surely it's incorrect to make a blanket statement that voluntary response data is worthless. Worthless means there is zero information in the data, which would only be true in extreme cases.


I think Professor Velleman promotes "Voluntary response data are worthless" as a slogan for the same reason an earlier generation of statisticians taught their students the slogan "correlation does not imply causation." That's because common human cognitive errors run strongly in one direction on each issue, so the slogan has take the cognitive error head-on. Of course, a distinct pattern in voluntary responses tells us SOMETHING (maybe about what kind of people come forward to respond), just as a correlation tells us SOMETHING (maybe about a lurking variable correlated with both things we observe), but it doesn't tell us enough to warrant a firm conclusion about facts of the world. The Literary Digest poll

http://aurora.wells.edu/~srs/Math151-Fall02/Litdigest.htm

is a spectacular historical example of a voluntary response poll that didn't give a correct picture of reality at all.


It's effectively worthless because it needs to be qualified so heavily. Dropping the qualification (as will happen sometimes) makes it more like misinformation, which is worse than worthless, so it averages out to zero...

"The median poll selection of a HN visitor who clicked on a poll which was on the front page on the following weekdays (and national holidays in the following countries:...) was..." etc.

If the information content is sufficiently hard to extract usefully that it would be easier to redo the sample in a sensible way, then you could call the sample worthless. (A bit like uneconomic oil reserves).


So, doesn't that just mean that a statistically significant portion of the regular HN users would have to respond?

Of course, you'd have to figure out how to define "regular HN users" and what would be a significant portion of them.

If this poll were to get that much response, it would also tell you something interesting: that HN users respond to polls in statistically significant numbers :)

So, voluntary response polls are not worthless if enough people voluntarily respond to them, but this is a tricky problem.


> So, doesn't that just mean that a statistically significant portion of the regular HN users would have to respond?

No. If 50% of HN visitors voted but 99% of those over 40 didn't (due to whatever confounding factor you like - say embarrasment), you'd still wouldn't be able to extrapolate from the sample to the category of "HN visitors".


Correct. It doesn't matter how large a sample size is if the sample is biased. This should be something that every high school student knows after any high school statistics course, but it is a crucial consideration that is widely ignored.


although we can all think of an edge case.. (especially if we know how the sample is biased)

eg: a gender survey in the general population that transgender respondees do not answer -- ends up with a 99% sample size (all but the transgenders) -- and thus gives a pretty significant bit of data on male:female ratios, etc..

Agreed in principle though.


No information is better than misinformation.


I'd argue that entirely depends on the context :) it's a bit of a blanket statement.


True some of the time, but not all of the time.


For the record I agree that the chances of the data exactly replicating the reality is minimal.

If nothign else because the youngsters here are more likely to respond (and also respond correctly) even though it is anonymous :)

But even allowing for that I would reckon we can use the data to extrpolate a few guesses. For example the high number of under 20's votes (I imagine these beign the most accurate numbers because of the deomgraphic too).

Plus of course this place is probably more likely to generate valid responses because of the general demographic. I would expect the data to be more accurate than a similar poll on, say, welovebritneyspears.com :)

So, yes (and to abuse a much overused maxim :D), voluntary poll data is untrustworthy. But some are more untrustworthy than others ;)


This is very interesting - does anyone know just how many of the polls we see on TV or in magazines are voluntary response polls? Should we be ignoring their results?

Intesting Anecdote: NDTV 24x7 (something like CNN in India) had extensive polls in the last (2004) election where they predicted a huge victory for the BJP party. They had poll results state-wise and interviews with eminent psephologists and so on, all predicting a resounding victory.

In the end, BJP lost - and quite badly too! Lies, damned lies, and statistics...


This is very interesting - does anyone know just how many of the polls we see on TV or in magazines are voluntary response polls?

Polls in general interest magazines, especially women's magazines, are usually wholly unreliable for modeling reality. (One thing that happens is the college-age men stuff the supply of replies full of phony responses, especially if the survey is about some salacious topic such as sexual behavior.)

I responded to a telephone call that came to my home phone number from Gallup Poll a few weeks after the United States presidential election. Gallup attempts to proactively call all households in the United States, and to correct for households that refuse to answer its calls. The interviews are quite long, and they ask a lot of demographic questions for stratifying the data gathered. If my caller I.D. box had not said "Gallup Poll," I never would have picked up my phone for a cold call. I got another poll call just this weeked, and many years ago got a call from the Harris Poll. The major polling companies attempt to actively gather random data samples, so what they do is distinguishable from voluntary response. But a magazine that writes a little sidebar "email your response to [email address]" is simply gathering worthless data. The same is true of TV stations that tell viewers to cell-phone text a response to some number they designate, or websites that have a poll form for anyone who surfs by.

If you are gathering data to improve a website, you are much better off conducting a live usability study in which you observe the user directly than simply polling visitors who voluntarily respond.


I couldn't find a figure on the number of registered HN users, so I used the recent number of daily unique IPs from http://ycombinator.com/newsnews.html. That number is 22,000.

For a confidence level of 95% with a confidence interval of 4%, we'd need 584 votes.

I'd imagine there's a lot of science to polling which this calculation probably ignores.


For a confidence level of 95% with a confidence interval of 4%, we'd need 584 votes.

I'd imagine there's a lot of science to polling which this calculation probably ignores.

Yep, everything about being sure we have an unbiased sample, which is not likely for a question like this.

What does your calculation say about the grouping of the data into categories?


I don't think the calculation says anything about the grouping of data. Or does it?

If you have an unbiased sample of 584 votes and 60% of the sample responded "25 - 30", could you say with 95% confidence that 60% (+/-4%) of the sampled population would select "25 - 30"?


We need the confidence thingy for this online community giving information about their age.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: