Borderline p-Values in Peer Review
A borderline p-value is one that falls just inside the 0.05 cut-off, close enough that a single data point or a defensible change of test could push it back out. Reviewers raise it when a result at p = 0.047 is written up with the same confidence as one at p = 0.001. The answer is not to argue the threshold but to check how stable the result is, report the effect size and interval beside the p-value, and match your wording to the evidence.
A p-value of 0.049 clears the conventional line and a p-value of 0.051 does not, yet the two are almost identical evidence. When a reviewer flags a borderline result, that near-identity is what they are pointing at, and the work is to find out whether a number sitting just inside the threshold would survive a small disturbance before you call it a finding.
What the reviewer wrote
The primary comparison reaches significance only marginally (p = 0.049), and I would encourage the authors to temper the strength of the conclusions in the abstract accordingly.
p = 0.048 is not a robust result. The entire headline claim rests on a number that barely clears the threshold, and I am not persuaded.
The manuscript is a useful contribution and my comments are minor. On the analysis, the key effect in Table 2 carries a p-value of 0.047; since this is the sole statistical support for the main finding, some acknowledgement of how close it sits to the conventional cut-off would be appropriate.
What they actually mean
The reviewer is not disputing your arithmetic. A p of 0.049 is below 0.05, and they know it. What they are objecting to is the distance between the confidence your language carries and the confidence the number earns, because a p at 0.001 and a p at 0.049 both clear the same line while the first is strong evidence and the second is weak. They are asking you to calibrate: report the borderline result honestly, with its effect size and interval, and soften any wording that presents it as settled. They are not asking for a different threshold, a deleted result, or a defence that 0.049 is technically under 0.05, since they have already granted that. The commonest misread is to answer the arithmetic when the objection was about overstatement.
This is the opposite of the complaint that a p just above 0.05 is being called a trend, which is covered in Trending Toward Significance. There the p is outside the line and the worry is a claim of momentum; here it is inside the line and the worry is overconfidence.
Why they are asking
A p sitting right at the edge of 0.05 is unstable in a way a small p is not. Because it lies a hair inside the line, one influential observation, or a switch from one legitimate test to another equally defensible one, can move it back across, so a reader who sees "significant" is trusting a verdict that a minor perturbation would have reversed. The effect estimate behind a borderline p usually comes with a confidence interval whose lower bound sits close to the null value, which means the data are also compatible with an effect too small to matter. Borderline results are also the ones most inflated by selective reporting, because when only the analyses that cleared 0.05 get written up, a p of 0.048 is exactly the value you would expect to find among the survivors. None of this makes a borderline result wrong; it makes it weak evidence, and the reviewer wants the write-up to say weak rather than settled. What a p-value can and cannot tell you is set out in What p-values mean, and the case against treating 0.05 as a bright line is in the p-value controversy. The American Statistical Association's statement on statistical significance makes the same point formally, cautioning that scientific conclusions should not be based on whether a p-value passes a specific threshold (Wasserstein and Lazar, 2016).
How to check it
Two things are worth reading off a borderline result: the interval around the effect, and how far the p-value moves when you disturb the data slightly. Take the PlantGrowth data, where a control group of plants is compared with a second treatment group on dried weight. A Welch t-test returns the p-value and, in the same output, the effect and its confidence interval.
The p-value is 0.0479, just under the conventional 0.05, so by the usual rule the difference is significant. The estimated difference is 0.494 grams, 5.526 against 5.032, and its 95% confidence interval runs from 0.005 to 0.983. That lower bound sits almost exactly at zero, so the same data that produced a significant p are also consistent with an effect of essentially nothing.
The second check is whether the verdict depends on the exact data and the exact test. Refitting the comparison while dropping each treatment plant in turn, and rerunning it as a rank-based Wilcoxon test that makes no normality assumption, shows how far the p-value travels.
The test the authors reported gives 0.048. The Wilcoxon test, a defensible choice for a sample this size, gives 0.063, on the other side of the line. Dropping a single treatment plant is enough to push the reported test itself as high as 0.085. The finding is real in the sense that some analyses put it under 0.05, but it is balanced on the threshold, and which side it lands on depends on choices no reader can see behind the word "significant".
What to do about it
You are fine
A borderline p-value is not a defect by itself. If the comparison was your single, pre-specified primary outcome, if the effect and its interval are reported, and if the result holds when you drop each observation in turn, then a p of 0.047 rather than 0.007 means the evidence is less decisive, not that it is flawed. Demonstrate the stability directly: rerun the test leaving out each data point and show the p-value stays under 0.05 throughout, so the reader can see the result does not hinge on any one observation. With a single planned comparison there is no multiplicity correction to apply, and you can say so. The one repair even here is the wording, so that a modest result is described as modest rather than as decisive.
It is fixable
Most often the problem is in the prose, not the analysis. If the abstract or discussion says a borderline result "confirms", "demonstrates", or "clearly shows" an effect, those verbs claim more than a p of 0.049 can support, so the repair is to report the estimate with its interval and pick language calibrated to it, which changes the wording without touching a single number. The difference of 0.49 grams (95% CI 0.01 to 0.98, p = 0.048) reads more honestly as modest, provisional evidence of an effect than as a demonstration of one. If you have a pre-specified secondary outcome or a sensitivity analysis pointing the same way, report it alongside, because a borderline result gains far more from a second independent line of support than from firmer adjectives.
It is a real problem
The result is in trouble when it is both the sole support for the headline claim and fragile, and the running example is exactly that. The PlantGrowth study has three groups, not two, so the control-versus-treatment comparison is one of three pairwise contrasts, and the correction the design calls for changes the picture.
Adjusted for the three comparisons the experiment actually makes, the difference that was significant at 0.048 now carries an adjusted p of 0.198, and its interval runs from -0.20 to 1.19 grams, so it includes zero. Because the result fails the analysis its own design requires, no change of wording can recover it. The honest response is to report the corrected comparison, describe it as non-significant, and revise the abstract and discussion so the study reads as inconclusive on this point rather than as evidence of an effect. Whether a correction is required, and which one, is covered in Multiple Comparisons; the point here is that when a borderline result depends on ignoring a correction the design demands, conceding is the only defensible move.
How to word your response
One reply for each outcome, each following the four-part pattern: restate the concern, say what you checked, say what it means for the conclusion, and point to where the change now appears.
If you are fine
The reviewer is right that the primary comparison is significant only marginally (p = 0.047). We have confirmed the result is stable rather than an artefact of a single observation: refitting the model while omitting each data point in turn leaves the p-value below 0.05 throughout. The comparison was the single pre-specified primary outcome, so no correction for multiple testing applies. We have added the effect size and its confidence interval to the Results (Methods, page X) and adjusted the abstract to describe the evidence as modest rather than definitive, so the strength of the language now matches the strength of the result.
If it is fixable
We thank the reviewer for this point. On rereading the discussion we agree that "demonstrates a clear effect" overstated a borderline result. We have reworded the passage to report the estimate with its interval, a difference of 0.49 grams (95% CI 0.01 to 0.98, p = 0.048), and now describe it as modest and provisional evidence rather than a demonstration (Results, page X). The analysis is unchanged; only the description has been calibrated to what the test returned. We have also added the pre-specified secondary outcome, which points the same way, to that paragraph.
If it is a real problem
The reviewer is correct that the finding rests on a single borderline comparison. On revisiting it we found the comparison is one of three the design makes, and after the correction appropriate to that design it is no longer significant (adjusted p = 0.198, 95% CI -0.20 to 1.19). We have accordingly revised the manuscript: the comparison is now reported as non-significant after correction (Results, page X), and the abstract and discussion describe the study as inconclusive on this outcome rather than as evidence of an effect (page X). We would rather report this accurately than rest the conclusion on a result the fuller analysis does not support.
Practice
A reviewer writes: "Several of the group differences you call significant have p-values close to 0.05, and I am not convinced the effects are real." One of the comparisons the objection sweeps up is the effect of the vitamin C delivery method at the lowest dose, which stands in here as orange juice against ascorbic acid at the 0.5 mg level in the ToothGrowth data. Run the comparison and decide which of the three outcomes it falls into.
Click to reveal solution
The reflex, under a comment that lumps several results together as borderline, is to soften all of them. For this comparison that would be the wrong call. The p-value is 0.0064, not a borderline 0.049, and the effect is large: mean tooth length is 13.23 under orange juice against 7.98 under ascorbic acid, a difference of 5.25 with a 95% confidence interval from 1.72 to 8.78 that stays well clear of zero. The result also holds when you drop any single animal, with the p-value ranging only from 0.004 to 0.016.
So you do not soften this one. The comparison is not borderline, its interval excludes a trivial effect, and it is stable across leave-one-out, which puts it in the first outcome above. Reply that this specific p-value is 0.006 rather than near 0.05, report the effect and its interval, and reserve the softening for whichever comparisons genuinely sit close to the line. That is why the actual p-value is worth reading before agreeing to temper anything, since a reviewer's blanket description of "borderline" will not always fit the particular result in front of you.