Brad DeLong explains some mechanics of the ESPN/FiveThirtyEight polls-only forecast:
- Takes recent past polls to estimate a current state of the race…
- Estimates an uncertainty of our knowledge of the current state…
- Projects that that current state of the race will drift away from its current state in some Brownian-motion like process…
- Calculates the chance that that process–starting from the estimated present–would produce 270 electoral votes for one candidate or the other.
This is also a near-perfect description of the “random-drift” calculation in the second line of the banner above. My main quibble here is that he says it’s not a forecast. Actually, it is. It assumes random drift from now to November 8th, Election Day. That is a forecast!
The issue is that it is a forecast based on current polls only. In my view, such a forecast has the conceptual flaw of overweighting the freshest data. It amounts to reporting a snapshot of today’s conditions, blurred out a bit. To make an analogy, this is like forecasting the temperature a week from now using a single temperature reading. You should not predict weather using a single reading taken on a hot day or in the middle of the night. Jon Stewart had a routine like this once to make fun of deniers who say that because it is snowing today, climate change must be a hoax. (at night: “The sun will never rise again!”)
People who get into a tizzy over post-convention bounces are a little like that. We are now in the middle of a post-DNC convention bounce for Hillary Clinton. What if her numbers come back down? In this sense, DeLong is correct – extrapolating from current conditions is not really a forecast. Though in that case, the “Brownian motion” (note: polling changes do not actually act like Brownian motion, so it would be better to just say random drift) plays no meaningful role, and should be dropped.
To generate a long-term prediction (their polls-plus forecast), FiveThirtyEight establishes a prior based on other factors such as the economy. (No, dear reader, I do not want you to tell me what those factors are!) This method can work – Drew Linzer at Votamatic and Benjamin Lauderdale have used it to make a long-range forecast, and have analyzed their model rigorously.
In my view, FiveThirtyEight’s polls-only approach does not take full advantage of this year’s polling data. It is possible to generate a long-term prediction using polls only. Here at PEC, that’s what we do to create a prior expectation for where polls will drift to. Think of it as “polls-plus-more-polls.” Here is how we do it.
As I explained in May, we have lots of data on how far polls can move during an election year. Here are time series from 16 Presidential campaigns. The red traces shows the ±1 standard deviation interval for the Democratic-minus-Republican margin, relative to the final outcome. Campaigns fall between the red traces about two-thirds of the time (68% of the time, to be exact).
Given this behavior, the optimal estimator for the final outcome is not a snapshot of conditions today. Instead, it is a weighted average of snapshots, where we would weight each margin using the square of the standard deviation to get “m bar”:
In this formula, m is the two-candidate margin (for example, the national polling margin; another example is the PEC Meta-Margin), Greek sigma is the value of the red trace at the corresponding date, and the summation is done over all dates d for which we have data. The result, m-bar, is an estimate of where voting on Election Day will end up.
One remaining challenge is to calculate the uncertainty on m-bar, which would allow converting it to a probability. That is a harder problem because on any given day, m is dependent on what it does on previous and later dates. It is a pain.
To escape this problem, I did something else: I used m-bar to establish a prior. This achieves the same effect as the “-plus” in “polls-plus” – but it uses older polls instead of econometric factors. I use m-bar for all of 2016 to estimate where the race is centered. I then estimate the range +/-S around m-bar, over which the final result may fall.
In the final step, PEC uses that range as a Bayesian prior. We take the random-drift calculation (see our banner) and combine it with the prior to get an estimate of where the election is likely to end up. That generates the “Bayesian win probability” in the banner above.
Because it is based on state polls and because the prior holds things in place, the resulting win probability does not move very much:
Now, this leaves the problem of how to estimate S. S is closely related to the red trace for the standard deviation in the first graph. Before 2004, S was large – up to 6 percentage points. Since 2004, S has been smaller, less than 3 percentage points. I believe this to be a symptom of polarization in politics, in which people choose sides.
So…which S do we choose? If 2016 is like 2004-2012, then S is 2-3 percentage points. If all bets are off this year and one party has broken the partisan logjam, then S could be larger, 6-7 percentage points, a value that gives Trump a fighting chance – but also opens the possibility of a massive Clinton victory. What will 2016 be like?
In 2012, I estimated S as the standard deviation of the Meta-Margin. This year, I am currently using national polls (current average, Clinton +4.7%) and a value for S of 6 percent (a “de-polarized scenario”). Sometime in Septemnber, I will transition over to using the average and standard deviation of this year’s Meta-Margin. As of August, that Meta-Margin is looking pretty stable (see the graph below). If it holds up, it will make the forecast considerably more certain.
In addition to the prior, the PEC November forecast is stabilized by the fact that it only uses state polls. That means that any change is integrated into the aggregate slowly, over a period of several weeks.
Broadly, I think it is a mistake to want a November prediction and a sharp snapshot of current conditions in the same calculation. To my taste, the FiveThirtyEight calculation would be improved by removing the random-drift step to give a more sensitive snapshot of today’s conditions, the same way that the HuffPollster national Clinton-v.-Trump average does. Then use a polls-plus-fundamentals (or in our case, polls-plus-more-polls) approach to give a true prediction.
@williamjordann @sleavitt1 the now cast is like if huffpo’s average got drunk and started swinging wildly at people. exists for clicks
— (((Will Cubbison))) (@wccubbison) August 1, 2016
One final thought, about econometrically-based priors (Ray Fair, Drew Linzer; but also Norpoth and others). In such a weird year, we may be past the point where an econometrically-based prior is useful. Linzer and Lauderdale have analyzed their own econometric approach, and found that the uncertainty on such a model was such that it could be useful for making predictions under normal, non-polarized conditions (high S). But we live in Polarized America (low S) and Trump America (highly unusual candidate). An econometric prior based on “bread and peace” does not take into account Hillary Clinton and Donald Trump’s personal negatives and positives.
In summary, I have doubts about whether anyone should be using an econometric model for prediction. However, an econometric prior can be useful as a research tool to calculate how the incumbent and challenging parties ought to be doing under normally extrapolated conditions – and give us a way of objectively evaluating whether Clinton or Trump was a particularly weak candidate this year.




Leave a Reply