A Controlled Test, a Clear Result, and a Culture That Carried On Regardless
Prayer is among the oldest and most universal of human practices. Across every culture, every century, and every theological tradition that has ever developed, human beings have addressed petitions to powers they cannot see, asking those powers to intervene in the physical world on behalf of themselves or others. The particular form varies enormously: it may be silent or spoken aloud, individual or communal, spontaneous or liturgically prescribed, embedded in a formal rite or offered in the darkness of a hospital corridor. The theological rationale differs from tradition to tradition. But the core claim, at least in its intercessory form, is consistent across all of them. There is an agent beyond the natural world who can be influenced by human supplication, and that influence can produce outcomes that would not otherwise have occurred. This is not a metaphorical or a poetic claim. It is a claim about causation in the physical world, which means it is, in principle, testable.
In 2006, a team of researchers led by Herbert Benson of Harvard Medical School published the results of the Study of the Therapeutic Effects of Intercessory Prayer, known by its acronym STEP, in the American Heart Journal. It was the largest, most rigorously designed, most adequately powered randomised controlled trial of intercessory prayer ever conducted, and it found nothing that the prayer hypothesis would predict. More precisely, it found no statistically significant difference in medical outcomes between patients who were prayed for and patients who were not, while simultaneously finding that patients who were told with certainty that they were being prayed for fared slightly worse than those left uncertain about their prayer status. The study cost approximately 2.4 million dollars, took nearly a decade to complete, enrolled 1,802 patients across six American hospitals, and produced a result that was, by any scientific standard, about as clear as results get. The religious world absorbed this finding, explained it away through a variety of inventive theological manoeuvres, and proceeded with no discernible change in the practice or promotion of intercessory prayer.
This essay examines what the STEP trial actually did, what it actually found, and why a result this methodologically robust has had essentially no effect on the culture that generated the question in the first place. The point is not to mock people who pray. Prayer as a private practice, as a form of meditation, as a means of ordering one’s thoughts and intentions, may well carry psychological value that has nothing to do with supernatural causation, and this essay will address that distinction directly. The point is the evidence: the specific, empirical, publicly made claim that petitionary prayer directed at another person’s recovery can alter their medical outcomes, and what happens to that claim when it is subjected to the kind of scrutiny that any other causal hypothesis would face as a matter of course.
1. The Question Behind the Study
Before examining the STEP trial itself, it is worth establishing why a controlled study of this kind was considered necessary and how it came to be funded and executed at the scale it was. The casual observer might assume that the scientific community simply decided to investigate a folk belief out of idle curiosity. The actual history is more instructive than that, and understanding it clarifies what the study was and was not attempting to show.
By the time the STEP trial was conceived in the mid-1990s, there was already a body of smaller studies claiming to have found positive effects for intercessory prayer on various medical outcomes. The most cited was a 1988 study by Randolph Byrd, published in the Southern Medical Journal, which claimed that cardiac care patients prayed for by Christians showed significantly better outcomes on several measures than patients in the control group. A subsequent 1999 study by William Harris and colleagues, published in the Archives of Internal Medicine, claimed similar findings with cardiac patients in a Kansas City hospital. Both studies received considerable attention in religious and popular media. Both were also subject to serious methodological criticism: small samples, multiple outcome comparisons without appropriate statistical correction, inconsistent measurement approaches, and questions about the integrity of blinding procedures. The scientific consensus was that neither study was remotely sufficient to establish the claim it appeared to support, but neither provided the definitive refutation that would have settled the question conclusively for a general audience.
STEP was designed explicitly to provide that definitive test. Herbert Benson, its principal investigator, was himself a researcher with a genuine interest in the effects of belief and contemplative practice on health outcomes; he had pioneered research into what he called the relaxation response and was not a reflexive sceptic about the potential of mind-body interactions. The study was funded in substantial part by the John Templeton Foundation, an organisation whose stated mission is the investigation of the relationship between science and faith and which has historically been sympathetic to the possibility that spiritual practice might have measurable physical effects. This point deserves emphasis, because the study has sometimes been characterised in religious media as an atheist attempt to discredit prayer: it was not. It was designed by a sympathetic investigator, funded by a sympathetic foundation, with a methodology intended to give the prayer hypothesis every reasonable opportunity to demonstrate an effect. If there was any bias in the study’s design, it leaned, if anything, toward finding a positive result.
2. What the STEP Trial Actually Did
The published paper, Benson et al., “Study of the Therapeutic Effects of Intercessory Prayer (STEP) in cardiac bypass patients,” American Heart Journal, 151(4), 934-942, 2006, describes a study that enrolled 1,802 patients at six hospitals who were scheduled to undergo coronary artery bypass graft surgery. The patients were randomised into three groups. The first two groups contained patients who were told that they might or might not be prayed for: one of these groups was actually prayed for, and the other was not. The third group consisted of patients who were told with certainty that they would be prayed for, and who were indeed prayed for. The primary outcome measure was the occurrence of any complication or death within 30 days of surgery.
The prayer itself was conducted by members of three Christian organisations: Silent Unity, a Missouri-based organisation associated with Unity Church; the Community of Teresian Carmelites in Worcester, Massachusetts; and the monks and nuns of St. Paul’s Monastery in Minnesota. These were not casual or informal prayers. The intercessors were experienced in contemplative prayer, they were provided with each patient’s first name and the first initial of their last name, and they were asked to pray for “a successful surgery with a quick, healthy recovery and no complications,” beginning the night before the patient’s surgery and continuing for fourteen days. The intercessors did not know which hospital the patients were in or anything about their specific medical condition beyond the fact of the upcoming surgery. The patients in the prayed-for groups did not know which specific organisation was praying for them.
The study was powered to detect a clinically meaningful difference in complication rates with reasonable statistical confidence. The investigators pre-specified their primary outcome and their analysis plan before examining the data, which is the standard practice for a rigorous trial and the procedure that most effectively guards against the manipulation of results after the fact. The blinding was as robust as the nature of the intervention permitted: the medical staff treating the patients were unaware of the patients’ group assignments, and the patients in the uncertain groups did not know whether they were being prayed for or not. The fundamental methodological requirements of a controlled trial were met to a degree that the smaller, earlier prayer studies had not approached.
The complications monitored covered a wide range of outcomes: major events such as death, cardiac arrest, pulmonary embolism, and stroke, as well as minor events including atrial fibrillation, wound infection, prolonged ventilation, and numerous other postsurgical complications. The 30-day window was chosen because it captures the acute postsurgical period during which any divine intervention in recovery would be most plausibly expressed. There was nothing arbitrary or designed-to-fail about the measurement framework. It was exactly the kind of rigorous outcome tracking that a genuinely positive result would have required in order to be taken seriously as evidence of an effect.
3. What the STEP Trial Found
Among patients who were uncertain whether they were being prayed for, complications occurred in 52 percent of those who were actually prayed for, and in 51 percent of those who were not prayed for. The difference is a single percentage point and carries no statistical significance whatsoever. In the language of the paper itself, there was no significant effect of intercessory prayer on complication-free recovery from CABG surgery. The prayer hypothesis, tested under the most rigorous conditions yet applied to it, produced no detectable effect on the primary outcome.
The third group’s result was, if anything, more striking and more uncomfortable for the prayer hypothesis. Among patients who were told with certainty that they were being prayed for and who were indeed prayed for, the complication rate was 59 percent, compared with 52 percent in the prayed-for uncertain group and 51 percent in the not-prayed-for uncertain group. The difference between the certain-prayer group and the uncertain-not-prayed group was statistically significant at conventional thresholds. Patients who knew they were being prayed for fared measurably worse than patients who did not know this. The investigators offered a cautious interpretation: the knowledge of being prayed for may have introduced a form of performance anxiety, an awareness of being observed and evaluated in some way that generated psychological stress and thus somewhat elevated physiological risk in the postsurgical period. This interpretation is speculative, as the investigators acknowledged, but it is the only hypothesis consistent with the data that does not require the abandonment of basic causal reasoning.
The primary finding, then, is not merely the null result that prayer provided no benefit. It is the directional finding that certainty of being prayed for was associated with worse outcomes for the prayed-for patients. Whatever mechanism prayer’s proponents have in mind when they claim it works, the STEP data do not support it. The study’s own abstract states: “Intercessory prayer itself had no effect on complication-free recovery from CABG, but certainty of receiving intercessory prayer was associated with a higher incidence of complications.” The paper is unambiguous, the sample size is large, the methodology is sound, and the result is the opposite of what the prayer hypothesis predicts.
It is worth pausing on the scale of effort involved in arriving at this negative result. 1,802 patients is not a small convenience sample. Six hospitals across the United States is not a localised pilot study. A decade of planning and execution is not a rushed investigation. Nearly two and a half million dollars of research funding is not a trivial commitment. All of this was brought to bear on a question that hundreds of millions of people answer confidently every day, in every country on earth, without evidence. The STEP trial was, by any reasonable standard, a serious and sustained attempt to find out whether the answer those hundreds of millions assume is correct has any empirical basis. The evidence establishes that it does not, and the size and rigour of the study mean this conclusion is not easily dismissible as a fluke of inadequate design.
4. Why This Result Should Have Mattered More Than It Did
The STEP trial was published in April 2006 in a peer-reviewed cardiology journal with a respectable impact factor. It received moderate coverage in the general press, largely because the story of a prayer study has obvious popular appeal regardless of which direction the result falls. Then, with a speed and thoroughness that would be remarkable if it were not entirely predictable, the result was absorbed into the existing framework of religious belief without disturbing it in any meaningful way. The apologists went to work, and within weeks the study was being characterised in religious media not as a refutation of the prayer hypothesis but as irrelevant to it. Understanding the specific moves made in that apologetic response is essential to the wider argument of this essay, because they reveal something important about the relationship between religious belief and empirical evidence.
The first and most common response was to argue that God cannot and should not be tested, that the very design of a controlled trial is incompatible with the nature of divine action. God, this argument runs, is not a mechanical force that operates according to fixed rules detectable by human instruments. He acts according to his sovereign will, which is not subject to human experimental design. Prayer is not a lever that produces outcomes on demand; it is a relationship, and relationships cannot be quantified in the way a pharmaceutical trial quantifies drug efficacy. This is a rhetorically sophisticated position, and it has the undeniable virtue of being unfalsifiable, which is precisely what makes it scientifically worthless. The difficulty with claiming that God acts in ways immune to detection is that this claim is indistinguishable, in every observable respect, from the claim that God does not act at all. An agent whose presence and actions produce no measurable difference from his absence is not, in any scientifically meaningful sense, an agent operating causally in the physical world.
The second common response was to argue that the study measured the wrong kind of prayer. The intercessors were given a specific form of words and a specific intention; real prayer, critics contended, is more spontaneous, more personal, more relationally embedded. Members of a community who know and love a patient pray differently from contemplatives who have been given a name on a piece of paper. This objection is worth taking seriously for approximately as long as it takes to notice that it proves too much. If prayer only works when it is sufficiently personal and relationally embedded, then the prayer offered by strangers in distant cities on behalf of patients they will never meet is, by this account, not really efficacious prayer. But this is exactly the kind of intercessory prayer that billions of people perform every week in churches, mosques, temples, and living rooms around the world: prayer for strangers, for distant suffering, for those they have heard of but will never know. If this kind of prayer does not work, then a substantial proportion of the world’s devotional practice is operating on a false premise, and the institutions that promote it bear some responsibility for saying so.
The third response was to argue that God was not obliged to answer prayers during a scientific study, that he may choose to withhold intervention precisely when that intervention would constitute observable evidence, because such evidence would remove the possibility of faith. On this account, the absence of a measurable prayer effect is not evidence against divine action but may even be evidence for it, since it reveals a God who values faith over proof. The circularity here is spectacular. This argument renders prayer effects unfalsifiable by definition while simultaneously explaining every negative result as consistent with, or even confirmatory of, the theological framework. Jerry Coyne, in Faith vs. Fact, identifies the underlying logic with admirable precision: “Theists’ typical response to these failures (i.e. of prayer to affect rates of healing) is to say either ‘God won’t let himself be tested’ or ‘That’s not what prayer is about: it’s simply a way to converse with God.’ But you can bet that had these studies shown a large positive effect, the religious would be noisily flaunting this as evidence for God. The confirmation bias shown by accepting positive results but explaining away negative ones is an important difference between science and religion.” This is precisely what happened. The apologetic response to STEP was not a serious engagement with evidence; it was the structure of confirmation bias made explicit and elevated to the status of theology.
Mark Twain had already identified the essential evasion in the faith-cure movements of his own era, more than a century before the STEP trial was designed. Writing in Christian Science (1907), he observed: “Within the last quarter of a century, in America, several sects of curers have appeared under various names and have done notable things in the way of healing ailments without the use of medicines. There are the Mind Cure, the Faith Cure, the Prayer Cure, the Mental Science Cure, and the Christian-Science Cure; and apparently they all do their miracles with the same old, powerful instrument, the patient’s imagination. Differing names, but no difference in the process. But they do not give that instrument the credit; each sect claims that its way differs from the ways of the others.” The observation maps precisely onto the landscape of post-STEP apologetics: each theological tradition claims that its prayer is different from the kind tested, and therefore the test does not apply. The mechanism is always elsewhere, always untestable, always just beyond the reach of whatever instrument has been brought to measure it.
5. Confirmation Bias as Theological Method
There is a general principle at work in the religious response to the STEP trial that deserves to be named directly, because it is not unique to prayer research; it is the operating logic of faith-based engagement with empirical questions across a wide range of domains. When a study appears to show a positive effect of prayer, that study is cited enthusiastically as evidence that prayer works. When a study shows no effect, the study is dismissed as methodologically inadequate, spiritually inappropriate, or theologically naive. The standard of evidence shifts depending on whether the result supports the prior belief. This is not an incidental feature of how religious communities engage with prayer research. It is the structural logic of faith applied to empirical questions, and it is visible in the history of prayer research with unusual clarity because the prayer hypothesis is specific enough to have generated actual data.
The Byrd study of 1988, with its small sample, its methodological weaknesses, and its positive result, was celebrated in religious media for nearly two decades as compelling evidence of divine intervention. The STEP trial, with its far larger sample, its more rigorous methodology, its sympathetic investigator, and its null result, was dismissed or ignored within months of publication. The asymmetry is not a coincidence. It is the operating system of confirmation bias applied to a question that the believer has already answered before the data were collected.
This asymmetry carries practical consequences that extend well beyond the academic debate about prayer research. If a clinical intervention, whether a drug, a surgical technique, or a rehabilitation programme, were subject to the same asymmetric evidential standards, the results would be dangerous. Medicines would be approved on the basis of small, poorly controlled trials showing positive results, and the large, rigorous trials failing to replicate those results would be dismissed as methodologically naive or somehow inappropriate to the intervention being tested. This does not happen in evidence-based medicine, at least not without vigorous challenge and institutional correction, because the scientific community has developed procedures, among them pre-registration, peer review, replication requirements, and systematic review, specifically designed to counteract the confirmation bias that is the default setting of human cognition. The theological community has developed no analogous corrective. The apologetic response to negative evidence is as reliable as the response to positive evidence, and the two together produce a system that is, by design, impervious to disconfirmation.
6. The Specific Claim Being Tested and Why It Matters
Some defenders of prayer have argued, with a degree of philosophical sophistication that deserves acknowledgement, that the STEP trial was not really testing prayer as theologians understand it. Prayer, on this account, is not primarily about changing external circumstances but about transforming the one who prays: about aligning the will of the petitioner with divine will, about expressing relationship, about submission and gratitude rather than demand and expectation. This is a genuinely interesting theological position, and if it were the only claim being made about prayer, the STEP trial would indeed be irrelevant to it. A study measuring postsurgical complication rates cannot determine whether someone’s prayer life has deepened their spiritual understanding or brought them closer to their conception of God.
But this is not the only claim being made about prayer, and it is demonstrably not the claim that most people engaged in petitionary prayer believe they are acting on. When a mother kneels beside a hospital bed and prays for her child’s recovery, she is not primarily engaged in an exercise of self-alignment. She believes, or at the very least hopes with some conviction, that her prayer might make a difference to the outcome for her child. When a congregation holds a prayer vigil for a member undergoing cancer treatment, the expressed intention is that the prayer will contribute to that member’s healing. When a church website maintains a prayer request board where members can submit names of sick relatives for communal intercession, the implicit premise is that prayer for those relatives might help them. The claim tested by STEP, that intercessory prayer can improve medical outcomes for the person prayed for, is not a straw man constructed by hostile atheists. It is the actual stated and practised belief of hundreds of millions of people in every major Christian denomination and its analogues in Islam, Judaism, and numerous other traditions.
The retreat, after a negative result, to the position that prayer was never really about changing outcomes is a post-hoc redefinition of the hypothesis. It is the theological equivalent of moving the goalposts after the ball has missed, and then insisting the match was never really about goals. If prayer is not about influencing external events, if it is purely a practice of internal transformation with no expectation of external effect, then the vigils, the prayer chains, the intercessory ministries, and the hospital chaplains who pray beside beds for patients’ recovery are all operating under a misapprehension about what they are doing. The honest application of the sophisticated theological position would be to inform all of these practitioners that their prayers cannot and are not intended to change anything beyond their own inner state. That message is not, in practice, being delivered. The retreat to the purely internal definition of prayer happens when a study produces a null result; it does not happen on Sunday mornings when the congregation is asked to pray for those in need of healing, or in the pastoral letter that urges believers to hold a sick child’s family in prayer. The inconsistency is telling.
7. The Prior History of Smaller Studies and What It Shows
The STEP trial did not emerge from a vacuum. It was the culmination of a decades-long research programme that had produced a chaotic and largely uninterpretable body of literature. A systematic review published in the Cochrane Database of Systematic Reviews, covering randomised controlled trials of intercessory prayer across various medical conditions, has consistently found no convincing evidence of benefit while noting significant methodological limitations across the existing literature. The pattern across these smaller studies is consistent with what statisticians call the file drawer problem, combined with the exploitation of multiple outcomes: researchers who test prayer across many outcome variables will, by chance alone, find statistically significant effects on some of them, and if only the positive findings are published while the negative ones remain unpublished, the literature will appear to support the hypothesis even when the underlying phenomenon is null.
The Byrd study, which appeared to show benefit, suffered from exactly this problem. It measured multiple outcomes without correcting for multiple comparisons, a procedure that systematically inflates the probability of false positive results. Its positive findings appeared on a handful of outcomes drawn from a much larger list, which is precisely what one would expect from chance variation alone in an underpowered study with loose outcome pre-specification. The Harris study of 1999 used a different outcome measure from Byrd and found positive effects on its chosen measure, but the two studies’ positive results fell on different outcomes, which is inconsistent with a genuine underlying effect and entirely consistent with spurious findings driven by selective reporting and multiple comparisons exploited without correction.
STEP was designed to eliminate these weaknesses systematically. It pre-specified a single primary outcome, it was adequately powered for that outcome, and it was conducted at a scale that made a chance false negative result extremely unlikely. If there were a genuine effect of the magnitude implied by the earlier positive studies, the STEP trial had more than adequate statistical power to detect it. The study found nothing consistent with such an effect. The responsible scientific conclusion is that the earlier positive findings were artefacts of inadequate methodology, and that the most rigorously designed test of the hypothesis produced the result that the null hypothesis would predict.
This is how science is supposed to work. Small, exploratory studies generate hypotheses. Larger, more rigorous studies test those hypotheses under controlled conditions. When the rigorous study finds no effect, the responsible conclusion is that the hypothesis is not supported. In medical research, this process, applied to interventions such as homeopathy, certain nutritional supplements, and various alternative therapies, has led to the withdrawal of those interventions from recommended clinical practice. In the case of intercessory prayer, the same process has produced no analogous correction, because the practice is embedded in a framework of belief that is not, at its foundation, responsive to empirical evidence in the way that medical practice is designed to be.
8. The Harm in the Claim
The question of whether intercessory prayer works is not merely an abstract philosophical curiosity. It has practical consequences that fall most heavily on people who are already in the most vulnerable positions: the ill, the bereaved, the frightened, and the desperate. Understanding these consequences is not a detour from the main argument; it is essential to grasping why the argument matters at all.
The most direct harm occurs when the belief in prayer’s efficacy substitutes for or delays medical treatment. This is not a hypothetical concern. There is a documented, recurring pattern across Christian Science communities, certain Pentecostal and charismatic groups, and a number of faith-healing movements, in which parents have chosen prayer over medical treatment for sick children, with fatal results. These are cases where the belief that prayer can change medical outcomes has been taken to its logical conclusion, and children have died from treatable conditions: meningitis, diabetes, appendicitis, pneumonia. The legal systems of several American states have spent decades wrestling with how to address this pattern without criminalising sincere religious belief, which is itself a consequence of the cultural weight that prayer claims carry when they are never seriously challenged by the institutions that endorse them. When the claim that prayer heals is treated as axiomatic, the minority who take that claim with maximum seriousness have no internal check available to them, and the consequences can be lethal.
A subtler but more pervasive harm is the psychological burden placed on those whose prayers are not answered in the way they hoped. If prayer genuinely works, then when a loved one dies despite extensive intercessory prayer by an entire community, the bereaved must find an explanation for the failure. The available explanations are limited and uniformly painful: perhaps there was insufficient faith; perhaps some sin or spiritual shortcoming blocked the prayer’s efficacy; perhaps God chose not to answer for inscrutable reasons that must simply be accepted. Each of these explanations adds a layer of guilt, confusion, or theological dissonance to grief that is already difficult to bear without additional burdens. The family that prays extensively for a dying child and loses that child anyway has been handed not only the loss but also the implicit suggestion that something they did or failed to do contributed to the outcome. The evidence cannot support this burden, because the evidence establishes that prayer made no systematic difference to the outcome. Placing this burden on bereaved people in the name of a practice whose efficacy has been tested and found wanting is a specific, identifiable, and unnecessary harm.
There is also the broader cultural harm of normalising a standard of evidence that, if applied outside the specific context of religious practice, would be regarded as intellectually reckless. When public figures, including politicians and healthcare professionals, attribute recoveries to prayer, or when hospitals include prayer in their standard-of-care documentation without appropriate qualification, the implicit message is that prayer is a legitimate medical intervention. The STEP trial, and the wider body of research it represents, establishes that it is not. The continued institutional endorsement of intercessory prayer as potentially effective in medical contexts, in the absence of evidence that it is, contributes to a cultural climate in which the distinction between evidence-based intervention and wishful thinking is systematically blurred. That blurring has consequences that extend well beyond the prayer ward.
Robert G. Ingersoll, writing in 1897 about the religious response to epidemic disease, identified the same institutional preference for prayer over practical engagement that the STEP result exposes in a different register. Reflecting on how the church treated the Black Death, he observed: “The church regarded epidemics as the messengers of the good God. The ‘Black Death’ was sent by the eternal Father, whose mercy spared some and whose justice murdered the rest. To stop the scourge, they tried to soften the heart of God by kneelings and prostrations, by processions and prayers, by burning incense and by making vows. They did not try to remove the cause. The cause was God. They did not ask for pure water, but for holy water. Faith and filth lived or rather died together. Religion and rags, piety and pollution kept company. Sanctity kept its odor.” The STEP trial is a modern chapter in that same long institutional history. The cause, in this case the false empirical claim that intercessory prayer changes medical outcomes, remains unaddressed, while the practice continues because the institutions that promote it have decided that their authority does not extend to acknowledging when their promoted practices have been tested and found ineffective.
9. The Philosophy of Testing the Untestable
There is a genuinely interesting philosophical debate, separate from the empirical question, about whether scientific methodology is even appropriate for investigating claims about divine action. This debate deserves a serious answer rather than a dismissive one, because the people making the methodological objection are not always simply trying to escape a negative result. Some of them have a substantive point about the philosophy of science and its limits, and engaging with that point honestly is more useful than ignoring it.
The argument runs roughly as follows. Science investigates regularities in nature, patterns that repeat reliably under specified conditions. Divine action, if it occurs, is not a natural regularity but a free exercise of divine will. A controlled trial assumes that the experimental treatment operates according to fixed causal principles such that it produces a consistent effect across a sufficiently large sample. But God, precisely because he is understood to be sovereign and free, is not bound by fixed causal principles. He may choose to answer prayer in some circumstances and not in others, according to reasons that are not transparent to human researchers. A controlled trial cannot capture this kind of irregular, will-dependent causation, and so it is not the right instrument for the job.
This argument has genuine philosophical content. There is a serious discussion in the philosophy of science about what kinds of phenomena are amenable to controlled experimental investigation and what kinds are not. But notice what happens to the intercessory prayer hypothesis if we accept this argument fully. We are left with a God who answers prayers in ways that are entirely indistinguishable from chance variation, because any pattern large enough to be statistically detectable would have been detected by a study like STEP, and STEP found no such pattern. If divine action is genuinely free and not subject to regular patterns, then from the human perspective, prayers are followed by outcomes at exactly the rate that chance and biology would predict, with no systematic deviation from the background rate. The claim that prayer works but works in ways we cannot detect or predict is not a claim about an observable phenomenon. It is a claim about an agent whose observable effects are identical to his non-existence. The question that follows, and it is an honest rather than a rhetorical one, is what kind of evidence could in principle distinguish this God from no God at all. The honest answer is that none has been identified.
For those interested in the wider question of what kinds of evidence would be required to establish or refute supernatural claims, this connects directly to the problem of unfalsifiable hypotheses explored in more detail in the essay on why the proof for God never comes. The prayer question is, in many ways, a specific instance of that more general problem: the supernatural claim perpetually retreats to a position just beyond the reach of whatever instrument has been brought to evaluate it.
10. What Would Convince a Scientist, and What Would Convince a Believer
The asymmetry between these two questions is instructive, and working through it carefully reveals something important about the nature of the disagreement between scientific and faith-based ways of evaluating causal claims.
What would convince a scientist that intercessory prayer works? The answer is straightforward, if demanding: a well-powered, pre-registered, randomised controlled trial with a clearly specified primary outcome that showed a statistically significant improvement in medical outcomes for prayed-for patients, replicated across independent research groups in different settings, with no plausible alternative explanation for the effect. This is the same standard applied to any other claimed therapeutic intervention. It is not designed to be impossible to meet. If prayer produced a genuine, large, consistent effect on medical outcomes, a properly designed study would detect it. STEP was designed to detect it. STEP found nothing of the kind.
What would convince a believer that intercessory prayer does not work? This is a harder question, and the responses to STEP suggest that for many believers, the answer is: nothing that a scientific study could show, because every methodological move available to negative evidence is already accounted for in advance by the theological framework. God cannot be tested. The study measured the wrong kind of prayer. The intercessors were not sufficiently faithful or spiritually prepared. The outcome measure was too narrow. The time window was too short. God chose, in his sovereignty, not to act within the study period. Each of these escape routes is available before any data are collected, and each of them will be reached for after a null result regardless of how carefully the study was designed. This is not dishonesty on the part of individual believers; it is the predictable consequence of holding a belief within a framework that has made itself structurally immune to disconfirmation.
The essay on pseudoscience and the evidence of God develops this point in the context of other unfalsifiable religious claims, but the prayer case makes it unusually vivid because the hypothesis was actually specific enough to generate a testable prediction. Most theological claims are too vague to generate testable predictions at all. The intercessory prayer hypothesis is not vague: it says that people prayed for will do better, medically, than people not prayed for. This is exactly the kind of claim that a controlled trial is designed to evaluate. The STEP trial evaluated it. The claim failed the evaluation. The community that held the claim then made it too vague to test. This sequence, specific claim, rigorous test, null result, post-hoc retreat to vagueness, is a diagnostic marker of pseudoscientific reasoning wherever it appears.
11. The Research That Followed, and the Silence
One of the most revealing aspects of the STEP trial’s legacy is not what happened immediately after publication but what failed to happen in the years that followed. In a functioning scientific field, a study of this size and methodological quality would generate substantial follow-up activity: attempts to identify moderating variables, pre-registered replications at other sites, meta-analyses incorporating the new data alongside prior studies, and a gradually converging consensus about what the totality of evidence shows. Some of this academic activity did occur. The Cochrane reviews of intercessory prayer have incorporated STEP’s findings, and the consensus among researchers not committed to demonstrating prayer’s efficacy has moved clearly toward the view that intercessory prayer has no measurable effect on medical outcomes.
What did not happen is equally significant. There was no large-scale attempt by the religious organisations that endorse intercessory prayer, the denominations, the prayer ministries, the hospital chaplaincy programmes, to engage seriously with the STEP findings and revise their understanding of what prayer does and does not accomplish. The John Templeton Foundation, which funded STEP and whose mission is precisely the rigorous investigation of spiritual claims, did not, to the author’s knowledge, use the STEP result as the basis for a fundamental reappraisal of how prayer is promoted or described in religious contexts. The theological publishing industry did not produce a wave of honest reckoning with the evidence. Parish prayer lists did not acquire footnotes noting that controlled trials have found no effect of intercessory prayer on medical outcomes. The practice continued, was promoted unchanged, and was described to the faithful in exactly the same terms that had been used before the study was published.
This is what institutional immunity to disconfirmation looks like in practice. It is not the result of any conspiracy or deliberate suppression of evidence. It is simply the consequence of operating within a framework in which the foundational claims are not regarded as subject to revision on the basis of empirical evidence, regardless of what that evidence shows. The prayer habit, the prayer culture, the pastoral comfort provided to the dying by the presence of a chaplain, none of these things required the intercessory prayer hypothesis to be empirically true in order to exist and to serve real human needs. The practice was never really resting on the hypothesis in the way that a medical intervention rests on evidence of efficacy. It was resting on something much more robust and much less rational: tradition, community, the need for comfort in extremity, and the deep human reluctance to say that a practice which feels meaningful is doing nothing beyond what it feels like.
12. What Prayer Might Actually Do, and Why That Does Not Save the Claim
It would be intellectually dishonest to discuss the evidence on intercessory prayer without acknowledging that other forms of prayer, self-directed contemplative prayer, mindfulness practices with religious framing, communal worship as a form of social bonding, have been associated in some studies with measurable psychological benefits. There is a legitimate and interesting body of research on the effects of contemplative practice on stress, anxiety, immune function, and overall well-being, and some of that research draws on samples from practising religious communities in which prayer is a central activity. This evidence is genuine and it would be wrong to dismiss it.
But notice what this evidence does not show. It does not show that a third party’s prayer for a specific person improves that person’s medical outcomes. The psychological benefits of a personal contemplative practice, to the extent they are real and replicable, operate through entirely natural mechanisms: reduced cortisol, improved sleep, enhanced social cohesion, the placebo effects of expectation and belief. These are mechanisms that require no supernatural agent and no actual communication with a divine being. They would operate in exactly the same way whether or not God exists, which means they provide no evidence for God’s existence or for his responsiveness to prayer in the intercessory sense. The claim tested by STEP was not “does praying make the person who prays feel calmer and better supported?” but “does one person’s prayer change the medical outcomes of another person who may not even know they are being prayed for?” These are entirely different claims, and the first is fully compatible with a purely naturalistic account of prayer’s function while the second is not.
The move that religious apologists make when confronted with the STEP result is to collapse the distinction between these two types of claim. They argue that prayer’s real value lies in its effects on the person praying, its effects on community solidarity, its effects on the mood and attitude of those gathered around a sick person, and therefore that STEP’s null result does not undermine prayer’s value. This is a legitimate observation about prayer’s non-intercessory functions. But it concedes the argument about intercessory prayer precisely by abandoning the intercessory claim. One cannot simultaneously argue that prayer’s real value lies in its effects on the person praying and also maintain that communities should pray for the sick in the expectation that it will help them medically. The two positions are incompatible, and the willingness to retreat to the first when the second is falsified, only to reassert the second in pastoral contexts where no one is currently testing it, is another instance of the asymmetric evidential standards that characterise faith-based reasoning about empirical questions.
For a broader examination of how the power of prayer is represented in culture and religious practice relative to what the evidence actually supports, the analysis of the power-of-prayer myth provides useful context alongside the specific trial findings discussed here.
13. The Institutions That Did Not Change Their Minds
Religious institutions occupy a position of cultural authority that places particular responsibilities on them when their publicly promoted practices are subjected to rigorous empirical investigation and found wanting. The major Christian denominations that endorse intercessory prayer, and whose members pray weekly and sometimes daily for the sick, the dying, and those in need of healing, were not simply bystanders to the STEP trial. They were, in an important sense, the primary interested parties. The trial was testing a claim that their theology endorses and that their pastoral practice promotes. The result should have mattered to them in a way that it evidently did not.
The argument is sometimes made that religious institutions should not be expected to revise their theological positions on the basis of empirical studies, because theology operates in a different domain from empirical science and answers different kinds of questions. There is a version of this argument that is coherent, though it requires the theologian to be consistently honest about which claims are empirical and which are not. The problem is that the claim “intercessory prayer improves medical outcomes for the person prayed for” is an empirical claim, not a theological abstraction. It is about causation in the physical world. It predicts a measurable difference between prayed-for and not-prayed-for patients. When an institution promotes this claim in its pastoral practice, it is making an empirical commitment, and it bears the same responsibility to revise that commitment in the face of disconfirming evidence that any institution making empirical claims bears. The argument that theology operates in a different domain only works if the institution is willing to specify clearly which of its claims are empirical and which are not, and then to actually revise the empirical ones when the evidence demands it. No major institution has done this with respect to intercessory prayer.
The pattern of institutional non-response to the STEP trial fits a broader historical pattern in which religious bodies have treated empirical challenges to their promoted practices as occasions for theological creativity rather than honest revision. When Galileo’s observations challenged geocentric cosmology, the institutional response was not correction but condemnation; it took the Church 350 years to formally acknowledge the error. When germ theory challenged the providential account of epidemic disease, the institutions that had promoted prayer as the primary response did not pivot immediately to endorsing public health measures; they incorporated germ theory slowly and under considerable pressure. The STEP trial is a smaller episode in the same pattern. The empirical challenge is clear. The institutional response has been, largely, silence and continuation. Whether this pattern eventually shifts, as it has, slowly and incompletely, on other empirical questions, remains to be seen.
14. The Evidence for Psychological Benefit Is Not Evidence for Supernatural Causation
A persistent strand of pro-prayer argumentation attempts to import evidence for prayer’s psychological benefits into the debate about intercessory prayer’s physical efficacy, as though establishing that one person benefits from praying somehow establishes that the person prayed for benefits as well. This conflation is so common and so consequential that it requires sustained attention before the argument can be resolved clearly.
There is reasonably good evidence that regular contemplative practice, including practices labelled as prayer in religious contexts, is associated with reduced anxiety, improved subjective well-being, and in some studies modest effects on physiological markers of stress. This evidence base is not as strong or as consistent as its popular representation suggests, and many of the relevant studies suffer from methodological limitations including self-selected samples, poorly matched control conditions, and inadequate blinding. But the core association, that people who engage regularly in reflective or contemplative practices tend to report higher subjective well-being than those who do not, is robust enough to take seriously as an empirical finding.
The question is what this finding establishes. The most parsimonious explanation is that reflective practice has beneficial effects through mechanisms entirely continuous with the rest of cognitive and physiological psychology: it reduces rumination, it promotes attentional focus, it may enhance a sense of social embeddedness when conducted in community, and it can contribute to the subjective sense of meaning and purpose that is associated with positive well-being outcomes across a wide range of research programmes. None of these mechanisms requires a divine listener. None of them distinguishes between prayer to a God who exists and prayer to a God who does not. The psychological benefits of the practice are fully explicable by natural mechanisms that would operate with equal efficiency whether or not the theological premises of the practice are true. This means that evidence of psychological benefit from prayer is evidence of something genuinely interesting about contemplative practice and human psychology, but it is precisely not evidence for the causal claim that the prayer hypothesis requires: that a supernatural agent receives the petition and acts on it.
The relevance of this distinction to the STEP debate is direct. When apologists argue that prayer produces real benefits and therefore the STEP trial missed something important, they are typically drawing on the psychological benefit literature. But the STEP trial was not testing whether the intercessors felt better for having prayed. It was testing whether the patients they prayed for experienced improved medical outcomes. These are different populations, different outcomes, and different causal pathways. The psychological benefits to the person praying, even if they are entirely real and produced by genuinely interesting mechanisms, are simply irrelevant to the question of whether a third party’s prayer improves the medical outcomes of the person prayed for. Treating them as relevant is either a misunderstanding of the question or an attempt to obscure what the question actually was.
15. The Honest Conclusion and What Follows from It
The honest conclusion from the STEP trial and the wider body of prayer research is not complicated, though it is uncomfortable for a significant proportion of the world’s population. Intercessory prayer, defined as one person’s petitionary prayer directed at improving another person’s medical outcomes, has been subjected to the most rigorous controlled investigation yet applied to it, and it has been found to have no measurable effect. The largest and best-designed study of the hypothesis produced a null result on the primary outcome and a slightly negative result on a secondary finding. The smaller studies that appeared to show positive effects were methodologically inadequate in ways that are now well understood, and the more rigorous the design of any given study, the less evidence for an effect appears. This is the pattern one expects when a genuine effect is absent and earlier apparent effects were artefacts of poor methodology and selective reporting. The totality of the evidence points in one direction.
What follows from this conclusion is not that people should be discouraged from prayer as a private practice, or that the sincere intentions of the intercessors who participated in the STEP trial were without value to themselves. What follows is that the specific causal claim, that prayer changes the medical outcomes of the person prayed for, is not supported by evidence, and that promoting it as though it were, especially in medical contexts, is not a benign or neutral practice. It raises false hope in people who are frightened and vulnerable. It places an implicit burden of responsibility on bereaved people whose prayers were not answered as they had hoped. It contributes to an evidential culture in which the standards applied to supernatural claims are systematically lower than the standards applied to everything else, and that lowering of standards has consequences that extend well beyond any individual prayer request.
The STEP trial did its job. It tested a specific hypothesis under controlled conditions and produced a clear result. The question of why that result has had so little practical effect on the culture it addressed is, in a sense, more interesting than the empirical question itself, because the answer reveals something more fundamental than the specific science of prayer research. It reveals the structure of a belief system that has organised itself so that no empirical finding can threaten it, while simultaneously making empirical claims about the physical world that it refuses to subject to empirical scrutiny in any sustained or honest way. The evidence for intercessory prayer is clearly negative. The conversation that has not yet taken place, and that the evidence has more than earned, is the one in which the institutions that endorse this practice acknowledge that plainly, and ask themselves what it means for how they describe prayer to the people in their care.
That conversation would require a standard of intellectual honesty that the evidence deserves and that the people in those pews deserve in equal measure. Whether it arrives is a different question, and the history of religious institutions responding to disconfirming evidence does not generate optimism. But the evidence is there, the result is clear, and the argument that prayer changes medical outcomes for the person prayed for has had its most rigorous test and failed it. The honest position is to say so plainly, without rancour toward the people who pray, and without softening a sound conclusion into false balance.
References
Benson, H., Dusek, J. A., Sherwood, J. B., Lam, P., Bethea, C. F., Carpenter, W., and Hibberd, P. L., “Study of the Therapeutic Effects of Intercessory Prayer (STEP) in cardiac bypass patients: a multicenter randomized trial of uncertainty and certainty of receiving intercessory prayer,” American Heart Journal, 151(4), 934-942, 2006.
Byrd, R. C., “Positive therapeutic effects of intercessory prayer in a coronary care unit population,” Southern Medical Journal, 81(7), 826-829, 1988.
Harris, W. S., Gowda, M., Kolb, J. W., Strychacz, C. P., Vacek, J. L., Jones, P. G., and McCallister, B. D., “A randomized, controlled trial of the effects of remote, intercessory prayer on outcomes in patients admitted to the coronary care unit,” Archives of Internal Medicine, 159(19), 2273-2278, 1999.
Masters, K. S., Spielmans, G. I., and Goodson, J. T., “Are there demonstrable effects of distant intercessory prayer? A meta-analytic review,” Annals of Behavioral Medicine, 32(1), 21-26, 2006.
Roberts, L., Ahmed, I., Hall, S., and Davison, A., “Intercessory prayer for the alleviation of ill health,” Cochrane Database of Systematic Reviews, Issue 2, 2009.
Coyne, J. A., Faith vs. Fact: Why Science and Religion Are Incompatible, Viking, 2015.
Ingersoll, R. G., “A Thanksgiving Sermon,” 1897.
Twain, M., Christian Science, Harper and Brothers, 1907.