Showing posts with label How to lie with statistics. Show all posts
Showing posts with label How to lie with statistics. Show all posts

Monday, April 12, 2021

How to Lie with Statistics – Part 5: The Gee-Whiz Graph

Continuing with my previous post, here I summarize Chapter 4 of the book How to Lie with Statistics. This chapter is titled, The Gee-Whiz Graph.

If this chapter could be summarized in six words it would be the following: It is not what it seems. Here we are illustrating how by merely adjusting the scaling of a graph you can get it to say whatever it is that you are trying to prove. Ten percent growth in one year is a “drop in a bucket” when illustrated in a 50-year graph; but when compared versus, say a 2 year graph, the 10-percent rate can be made to look fantastic. Adjust to quarters, or months, or days, and you can get even more stunning results – just like the doctor ordered. In other words, what would visually be a handsome yearly return is lost when viewed in the context of a longer time frame; and vice versa. So, whenever anyone shows you some graph (yours truly included) always check the scaling. Do not be deceived.

Wednesday, April 7, 2021

How to Lie with Statistics – Part 4: Much Ado About Practically Nothing

Continuing with my previous post, here I summarize Chapter 4 of the book How to Lie with Statistics. This chapter is titled, Much Ado About Practically Nothing.

This brief chapter explains the critical importance of understanding how representative is a sample metric of the population as a whole. We can actually calculate that answer. First, we must understand that there are two figures that can be used to give us a sense of that representation: the probable error and the standard error. My summary will focus on the Standard Error because it is the most commonly used measure these days.

The Standard Error is used to construct confidence intervals, whereby if the calculation is sound then the true value would be contained within those intervals. This means that there is some probability that the true value will in fact be outside those intervals (range of values). For example, you are 95% confidence that the true value is within some confidence interval. Said another way, there is a 5% chance that the true value is outside the range depicted by the confidence interval.

Also of importance is to recognize that the wider the intervals, the less confidence we have as relates to the estimated value. For instance, saying some value is within a range of 5 to 10 will carry greater weight if you are told some value is within 1 to 100. The latter may be more accurate, but precision is way off. This would be a polite way of saying, “take our results with a ‘grain of salt’”; or putting it another way, the results are close to meaningless.   

As you can imagine, rarely you see reported estimates this way. That is why it is important to go to the source documents, not just uncritically receive the headline media reports. Do not be fooled.

Saturday, April 3, 2021

How to Lie with Statistics – Part 3: The Little Figures That Are Not There

Continuing with my previous post, here I summarize Chapter 3 of the book How to Lie with Statistics. This chapter is titled, The Little Figures That Are Not There.

The key thought of this chapter is about what is left unsaid when a particular statistic is illustrated – an average without a range, for example. Today’s post about COVID metric reporting was also illustrative of this point. Take an “average” temperature of any given city and that tells you nothing if a range is excluded. Or better yet, when no information is given as relates to sample size or method of deriving the statistic.

Many times a sample statistic may be passed as representative of the whole, but in fact the sample being singled out is merely one of many taken and conveniently left out (if it doesn’t fit the agenda being pushed, of course).  In addition, “how likely it is that a test figure represents a real result rather than something produce by chance”? In these instances what you should be told is some sort of degree of probability telling you something about the statistical significance of the results. In plain English, how likely are the results truly indicative of the population as a whole? In no uncertain terms you should be told what that likelihood is.

Furthermore, be skeptical of charts that do not have proper scales or at worst deliberately exclude information.

Some Things Are Not What They Seem: The COVID Metric Reporting is Worse Than the Virus

There is a lot of confusion these days whenever you encounter the use of COVID metrics in mainstream media reports. I am reminded of my days in college when a Statistics professor said that in order to get published in a peer-reviewed journal, the idea was to come up with a conclusion and find data that supports that conclusion. I remember when I heard that, even at my developmental stage of learning, something did not sit right with me. It seemed to me some sort of deception simply for personal gain. It was later in life that I fully understood the implication of what my professor said at that time: if the majority of professors and professionals engage in that sort of activity, then the amount of garbage being produced would increase over time. The lines between true and false would be blurred; and if you have a thought that borders absurdity, then you simply make up the data (legally of course) to prove the absurdity. 

This is what why it is helpful to read and understand the simple concepts in the book, How to Lie with Statistics. I've written a few posts summarizing that book; one is here. Simply click on the label with the name of the book.  

Here in this post I will simply highlight some of this absurdity. Back on March 18, 2021, the Financial Times published an article titled, “The US vaccine effect: rapid rollout starts to bear fruit,” where they showed various metrics that illustrated the success of vaccines in reducing the number of death and cases allegedly due to COVID. In the article they produced the following graph:

The pre- (right graph) and post- (left graph) trend is obvious. Down. Actually, the left side is more down than the right. The point the FT is trying to prove is that vaccines are working as intended. But poking a bit deeper into the metrics they report, you will notice some things don’t quite add up. For example, “cases and deaths as a percentage of peak during each wave”, what exactly are those peaks? Nowhere in the article have they told you that. That alone invalidates the analysis; or at a minimum raises the question of how exactly they arrived at the numbers to come up with the graph. If you are left to guess how a metric was derived, then it is likely you are being conned (irrespective of motive by the writer, here the FT). The assumption attributed by the FT to these graphs is that the vaccines are the reason why cases and deaths have declined.

 Now, as a counter example, take a look at the broader COVID case data reported by the CDC. The graph as of April 1, 2021 is here:

Unfortunately the graph is not as clear when copying here, so you can go to the CDC and look at the data there. When you focus on the time period since the FT article was written (mid-March 2021 to the present). What do you see? Yes, the trend is up, albeit slightly. So, if the FT logic is sound, we must then conclude that vaccines are not bearing fruit. Indeed, the FT metric and the ones I am showing you are not identical. But the point I am making is that if you have a particular conclusion, you can find data to fit that conclusion.

Tuesday, March 23, 2021

How to Lie with Statistics – Part 2: The Well-Chosen Average

Continuing with my previous post, here I summarize Chapter 2 of the book How to Lie with Statistics. This chapter is titled, The Well-Chosen Average.

The key thought in this chapter is to remember to ask more questions when you hear or read about some “average”, such as the “average” income, the “average” pay, or “average” price – just to name a few. The first thing to remember is to ask, what kind of average and how that metric is calculated. An arithmetic average, which you simply divide a total amount by the number of observations, will yield a different outcome compared with the median average, which gives you the middle amount after sorting the data in ascending or descending order.

When averages are particularly highlighted, that should alert us to be on the lookout that a deceptive figure may be coming our way. If you are not told exactly how he average is calculated, do not readily accept the conclusions being submitted.

Any dataset that is skewed, that is, when graphed you see that the data particularly bunched in one side, the relevant “average” is the median. The arithmetic average will always be deceptive in such skewed data. So, be aware, be alert when you come across the “average” something.

Saturday, March 20, 2021

How to Lie with Statistics – Part 1: The Sample with the Built-in Bias

In times of confusion I have a general principle that I often apply: go back to basics! In other words, when things may appear complex, gibberish, or outright confusing, you should still be able to filter down key pieces of information in such a way that will allow almost any individual with some curiosity to discern the quality of the results. All you need are some simple tools. A hammer goes a long way to hammering nails, as opposed to using a piece of rock. So, it is with the intent of passing along some useful tools that I write the following book summary from the classic How to Lie with Statistics by Darrell Huff. This is a fantastic little book that gives insight as to the tricks used by many people (particularly those who have an agenda) in order to fit the data to their conclusions. Beware of such people.

I will write a short summary of each chapter and then posted it here periodically. Here is Chapter 1.

The Sample with the Built-in Bias

When we want to get a sense of a general trend in some population, we normally select a sample from that population. That sample must be representative of the whole. When it is not, then any metric we calculate will be an aberration, an illusion, or an incorrect figure. A key metric, or more commonly referred to as an statistic, is the “average”: the “average” person, the “average” temperature, the “average” employee, etc. Be on the alert when you come across such phrase because the proverbial “average” may be ripe with all sorts of bias – some of which is seen and some of which is unseen.

As I mentioned, a sample is supposed to be representative of the population. When it is not and we subsequently extrapolate conclusions based on the sample metric, we will draw all sorts of wrong implications. And by extrapolation I mean that you take some metric and apply conclusions beyond the boundaries that the metric allows you. For example, let us say a prestigious business school publishes the “average” salaries obtained from respondents who submitted surveys. The likely outcome is that a very handsome salary will be the average and thus published to the public. But consider the following as relates to this survey: self-respondents may tend to exaggerate, or those not choosing to report might do so because of their low salaries and not wanting to be perceived as failures. In our example, if these two factors are not disclosed or accounted for, you might as well disregard the “average” salary metric as bogus. 

Precision of a metrics is also a red flag: for example, when you read/hear that an “average” family has “2.05” children, or that someone on average takes “7.85” showers a week, etc. Such number is so precise that it frankly renders is meaningless. There are many things in life that are nearly impossible to measure with such precision.