Claims about fulfilled biblical prophecy occupy an unusual position in Christian apologetics. Unlike arguments from personal religious experience, prophecy appears capable of offering something approaching an objective test. A text was written. It predicted an event. The event subsequently occurred. If the prediction was sufficiently specific, sufficiently improbable and demonstrably written before the event, then perhaps coincidence eventually becomes an inadequate explanation.
At least, that is the argument.
The attraction is obvious. It apparently transforms revelation from a matter of faith into a statistical question. Peter W. Stoner famously attempted precisely this in Science Speaks, calculating that the probability of one person accidentally fulfilling eight selected Messianic prophecies was approximately one in 1017. He illustrated this with the memorable image of covering Texas two feet deep in silver dollars, marking one, and asking a blindfolded person to select it.
There is, however, a rather fundamental difficulty. The impressive number is the output of the calculation. The interesting questions concern the assumptions that went into it.
The proper question therefore is not simply, “What are the odds that Jesus fulfilled all these prophecies?” It is:
That is a different problem. Conveniently, it is one that Bayesian reasoning is designed to address.
Two competing hypotheses
Let us begin with two deliberately broad hypotheses.
HP - Predictive Revelation. At least some passages in the Hebrew Bible contain genuine information about future events ultimately fulfilled by Jesus which could not reasonably have been obtained by ordinary means.
HN - Natural Historical and Literary Development. The apparent fulfilments arose through some combination of coincidence, broad prediction, retrospective interpretation, deliberate action, legendary development, literary dependence, theological construction and the normal human tendency to notice successful matches while ignoring unsuccessful ones.
Notice what HN does not assert. It does not require Christianity to be false. It does not require Jesus not to have existed. It does not require the Gospel writers to have been dishonest. It does not even require every purported fulfilment to be mistaken.
It merely represents the collection of ordinary mechanisms already known to operate in historical and religious literature.
For each alleged prophecy we can then ask how much more likely the evidence would be under HP than under HN. In simplified form:
A large Bayes factor would favour predictive revelation. A value around 1 would tell us essentially nothing. A value below 1 would mean that the evidence was actually more expected under ordinary historical and literary processes.
The important point is that we do not begin by assuming that every claimed prophecy is an independent miraculous bullseye. We make each one earn its evidential weight.
What makes a prophecy impressive?
Consider two imaginary predictions. The first says, “A great ruler shall arise and there shall be wars.” The second names a particular person, place, date, time and highly unusual event centuries in advance.
If both were unquestionably written beforehand, the second would plainly be more remarkable if fulfilled. Specificity matters. But specificity is only one variable.
Predictive status
First we must establish that the alleged prophecy was actually a prediction. This sounds embarrassingly obvious, yet it immediately causes trouble for some familiar Christian examples.
Matthew 2:15 applies Hosea 11:1 to Jesus: “Out of Egypt I called my son.” But Hosea is not predicting a future Messiah. In its original context the passage refers retrospectively to Israel and the Exodus. Matthew is performing a theological or typological reading of scripture.
That may be perfectly legitimate theology. But it is not straightforward prediction.
Call this variable P - predictive status. A clear, unambiguous prediction receives a high value. A retrospective passage subsequently reinterpreted prophetically receives a low one.
Specificity
Next comes S - specificity. “Someone will suffer” deserves virtually no evidential weight. A prediction identifying an unusual person, place, action, sequence and time deserves considerably more.
The critical question is how many plausible outcomes would count as fulfilment. A prediction capable of accommodating almost anything predicts almost nothing.
Dating confidence
Then there is D - confidence that the prediction predates the alleged fulfilment. A spectacularly accurate prediction written afterwards is not prophecy.
For many Hebrew Bible texts we can establish comfortably that some form existed before the first century CE. The Dead Sea Scrolls are particularly valuable here. Individual passages nevertheless require attention to textual history, redaction and dating.
The relevant question is not merely, “Is this in the Old Testament?” It is, “How confidently can we establish that this particular predictive wording existed before the event it supposedly predicts?”
Historical confidence in the fulfilment
Call this F. Suppose an ancient text predicts that the future Messiah will sneeze seven times beneath an olive tree. We then discover a biography written decades after the supposed Messiah's death which says that he did exactly that “so that the prophet might be fulfilled”.
We cannot simply treat the biography's statement as an independently verified historical event.
The evidential question is whether we have historical grounds for believing the event occurred independently of the document attempting to demonstrate the prophecy's fulfilment.
This matters enormously when dealing with the Gospels. Their authors knew Jewish scripture, and Matthew repeatedly and explicitly presents events in Jesus's life as fulfilments of scripture. That does not prove that the events did not happen. It means that literary dependence must be included in the probability model.
Literary dependence
We can represent that as L. If the author reporting a purported fulfilment knows the alleged prophecy, that knowledge provides an ordinary mechanism by which correspondence can arise.
Sometimes the influence may simply affect wording. Sometimes it may affect interpretation. Sometimes an author may organise genuine historical memories around scriptural motifs. Sometimes scripture may contribute directly to the construction of a narrative.
Prediction and fulfilment are not independent observations merely because they occur on different pages.
Controllability
Now consider C - whether the alleged fulfilment could deliberately have been produced.
If tomorrow's horoscope says I shall wear black and, suitably impressed, I put on a black shirt, astrology has not scored a remarkable prediction. I have.
This becomes relevant to events such as entering Jerusalem on a donkey. If Jesus knew Zechariah and deliberately enacted the symbolism, the correspondence may be historically genuine while providing little evidence of supernatural prediction.
A birthplace, assuming it is independently established, therefore carries more potential evidential weight than choosing a mode of transport. Death circumstances controlled by an executioner may carry more than words deliberately spoken by a person familiar with the relevant scripture.
The independence problem
Suppose I calculate probabilities of 1/100 for A, B and C, multiply them, and announce odds of one in a million. That multiplication is justified only under appropriate independence assumptions.
Alleged Messianic prophecies frequently are not independent. Several may derive from the same passage, describe aspects of the same event, be reported by the same Gospel author, depend upon the same theological interpretation or derive from one narrative tradition.
Multiplying them as though they were independent artificially manufactures astronomical numbers.
A serious model therefore requires a dependency graph. If two fulfilments share a source, event or interpretative mechanism, their evidential contributions must be correlated rather than simply multiplied.
The selection effect
Why are we examining these prophecies? Because Christians have spent approximately two millennia identifying passages that appear to correspond with Jesus.
This creates an enormous selection effect.
Give me the collected works of Shakespeare and the biography of Winston Churchill. Allow me to search every line, use metaphor, typology, translation variants and partial quotations, and decide retrospectively which events in Churchill's life count as matches. I confidently predict that I shall discover some impressive “prophecies”.
The probability of any preselected correspondence might be small. The probability of finding some correspondence after searching thousands of passages and hundreds of biographical events can be very large.
We remember the hits. We quietly forget the misses.
Preregistering God
This suggests an experiment. Before calculating anything, construct a fixed corpus of alleged Messianic prophecies.
For each one, preregister the original passage, probable date, original historical context, predictive status, precise fulfilment criteria, failure criteria, permitted date range, historical security of the alleged fulfilment, controllability, literary dependence and the number of alternative people or events capable of satisfying the same criteria.
Only after those rules are fixed do we examine the candidate.
Failed prophecies must count
If someone makes 1,000 predictions and gets ten strikingly right, presenting only those ten produces an impressive prophet. Presenting the other 990 produces a rather different statistic.
A prophecy corpus therefore cannot consist solely of passages Christianity eventually decided Jesus fulfilled. We must also include candidate Messianic passages that apparently were not fulfilled, expectations applied differently by contemporary Jewish interpreters, predictions whose supposed fulfilment requires substantial reinterpretation, prophecies whose historical fulfilment is unknown, and apparent failures.
Otherwise we are effectively calculating P(success | selected successes), which, with commendable obedience, produces rather a lot of success.
The control group
Now comes the most interesting part of the experiment. Apply exactly the same procedure to people other than Jesus.
Take a large ancient literary corpus written before Augustus. Allow exactly the same degree of metaphor, typology, translation variation, textual reinterpretation and historical uncertainty permitted in the Jesus experiment. How many apparent predictions of Augustus can we discover?
Repeat the exercise for Alexander the Great, Muhammad, Sabbatai Zevi and Napoleon.
Then include deliberately awkward controls: fictional characters such as Harry Potter, Paul Atreides or Luke Skywalker. If sufficiently flexible prophetic interpretation produces “astonishing” correspondence there too, we have learned something important about the method rather than about the divine status of the control subject.
Monte Carlo prophets
Many variables will inevitably be uncertain. Instead of pretending otherwise, represent them as probability distributions.
Perhaps expert assessments place the probability that a passage constitutes genuine prediction between 0.3 and 0.6. Perhaps historical confidence in the fulfilment ranges between 0.4 and 0.7. Perhaps literary dependence is almost certain.
Rather than choosing whichever number produces the result we prefer, repeatedly sample from the plausible distributions. Run the model 100,000 times, or a million.
We then obtain a distribution of outcomes rather than one theatrically precise number.
More importantly, perform sensitivity analysis. Which assumptions drive the conclusion? If predictive revelation is strongly favoured across almost every reasonable combination, that would be interesting. If the result favours revelation only when apologetically convenient values are assigned to several disputed variables, that would be equally interesting.
What would strong evidence actually look like?
This framework is not constructed so that prophecy must fail.
Imagine archaeologists discover an unquestionably pre-Christian manuscript dated to 300 BCE naming a future teacher, his birthplace, an exact date, Pontius Pilate and highly specific circumstances of death. Suppose independent Roman records then confirm every detail.
I would be extremely interested.
Literary dependence could not readily explain it. Deliberate fulfilment could not explain most of it. Vague interpretation would have little room to operate. Dating would be secure. The prediction would be specific. The fulfilment would be independently corroborated. The Bayes factor could consequently become enormous.
The methodology does not exclude supernatural evidence. It merely requires supernatural claims to meet evidential standards capable of distinguishing them from ordinary explanations.
Bethlehem as a worked example
Consider the familiar claim that Micah 5 predicts the Messiah's birth in Bethlehem. At first glance this looks promising. Micah predates Christianity. Bethlehem is specific. Matthew and Luke both associate Jesus's birth with Bethlehem.
But the model immediately asks further questions. Is Micah 5 straightforwardly predicting the birthplace of a distant individual Messiah? How confident are we historically that Jesus was born in Bethlehem? To what extent are Matthew and Luke independent on this point? How should we treat their knowledge of Jewish scripture and the theological importance of David's city?
None of those questions proves Jesus was not born in Bethlehem. They determine how much evidential weight the alleged fulfilment deserves.
“Micah says Bethlehem and Matthew says Bethlehem” is not the end of the calculation. It is where the calculation begins.
Hosea as a different kind of case
Now consider “Out of Egypt I called my son.” Hosea 11:1 reads in context as a reference to Israel's past: “When Israel was a child, I loved him, and out of Egypt I called my son.” Matthew applies this to Jesus returning from Egypt.
Whatever the theological merits of Matthew's typology, the passage was not originally an explicit prediction that the Messiah would travel to Egypt. Its predictive-status score should therefore be low.
That does not mean Matthew made a mistake. It means modern apologetics makes a category mistake when it takes an ancient Jewish typological interpretation and feeds it into a probability calculation as though Hosea had written a dated forecast concerning Jesus.
The donkey problem
Zechariah's king entering Jerusalem on a donkey presents another category. Here we have a recognisably predictive royal image. But Jesus could know the passage.
If Jesus deliberately rode into Jerusalem in conscious enactment of Zechariah, the historical correspondence might be excellent. Its supernatural evidential value would nevertheless be modest.
A person deliberately doing something predicted about him is not statistically surprising.
What Stoner's number actually tells us
This brings us back to the famous one in 1017.
Such a number can only be as reliable as the prophecy selection, assigned individual probabilities, independence assumptions, historical reliability of each fulfilment, absence of literary dependence, absence of deliberate fulfilment, treatment of failed predictions and size of the search space from which successful matches were selected.
Those are not minor technical details.
The experiment I would like to see
So here is the City of Dis prophecy experiment.
Take perhaps fifteen of the most commonly cited prophecies concerning Jesus. Freeze the list. Publish the inclusion rules and scoring criteria. Invite Christian, Jewish, atheist and religiously unaffiliated scholars to challenge the values assigned to each variable. Construct the dependency graph. Run the Bayesian model.
Then repeat the experiment against control figures using exactly the same rules.
Finally, run Monte Carlo simulations across the plausible range of disputed assumptions. Publish everything: code, data, failed cases, alternative assumptions and sensitivity plots.
No hiding inconvenient prophecies. No assigning a probability of one in a million because something “sounds unlikely”. No multiplying dependent events. No treating Gospel claims as independent archaeological confirmation of themselves. No changing the definition of fulfilment once we have seen the results.
If the Jesus dataset produces results dramatically different from Augustus, Alexander, Muhammad, Napoleon and the fictional controls, we have discovered something genuinely interesting.
If it does not, we have discovered something equally interesting.
What the experiment is really testing
Human beings are extraordinarily talented pattern-recognition machines. Usually this is useful. Sometimes we see Orion in unrelated stars, faces in clouds, messages in random noise and predictions in old texts.
Religious traditions add another mechanism: generations of intelligent interpreters repeatedly examining the same corpus, preserving successful interpretations and developing increasingly sophisticated explanations for apparent failures.
Given enough text, enough history and enough interpretative freedom, impressive correspondences are not merely possible. They are inevitable.
This does not prove that biblical prophecy is false. It establishes the null hypothesis against which prophecy must compete.
The interesting probability is therefore not how unlikely selected fulfilments appear after Christianity has spent two thousand years selecting them. It is how probable it is that a sufficiently large ancient corpus, subjected to centuries of motivated interpretation, would produce apparently remarkable correspondences with the life of its religion's central figure even if no supernatural prediction occurred.
Until that probability is considered, numbers such as 1017 are not evidence.
They are decoration.
A universe under no obligation to prophesy
Cosmicism begins with an uncomfortable methodological principle: the universe is not obliged to arrange itself around human expectations.
That applies just as readily to religious texts.
Perhaps genuine revelation occurred. Perhaps somewhere among the prophetic literature of antiquity there really is information that ordinary historical processes cannot adequately explain.
But we do not discover that by beginning with the desired conclusion and marvelling at the improbability of arriving there.
We construct competing explanations. We specify beforehand what evidence would distinguish them. We count failures as well as successes. We penalise flexible interpretation. We account for dependence. We test our method against controls. And then we allow the evidence to move us wherever it happens to lead.
There is something almost pleasingly perverse about applying this to prophecy. For centuries apologists have insisted that fulfilled prophecy constitutes objective, even mathematical evidence for revelation.
Very well.
References and further reading
Peter W. Stoner and Robert C. Newman, Science Speaks, Moody Press, revised edition, 1963. The famous “one in 1017” calculation is discussed in the section on Messianic prophecy.
Matthew 2; Matthew 21; Hosea 11; Micah 5; Zechariah 9. Biblical passages should be read in their surrounding literary and historical contexts rather than solely through later fulfilment formulae.
For Bayesian reasoning and Bayes factors, see Richard E. Kass and Adrian E. Raftery, “Bayes Factors”, Journal of the American Statistical Association 90, no. 430 (1995), pp. 773-795.
For the methodological importance of multiple comparisons, dependence and preregistration, see the broader statistical literature on researcher degrees of freedom, selective reporting and confirmatory analysis.
Loading this essay's discussion...