
Why "sex-specific" findings often aren't
The search for sex differences is one of the fastest-growing areas of biomedical study. Reports are common that an experimental manipulation had a sex-dependent effect—that it “worked” better in one sex than another. Such findings influence our thinking about disease, our future research, and our clinical decisions.

Reports of sex-dependent effects
Reports of sex-dependent effects are published in the Neurosciences at three times the rate any other field.
As the number of reports explodes, it becomes increasingly important to ask whether these claims are supported by solid evidence.
Our recent article in the Proceedings of the National Academy of Sciences shows that in the behavioral and brain sciences, major claims of sex-dependent effects are supported by appropriate evidence less than 25% of the time.
-
We sampled 200 recent research articles with claims of sex-dependent effects in the title (articles using the term “gender” were also included).
-
The articles came from fields such as neuroscience, psychology, and psychiatry, and included human and non-human animal research.
-
For each article, we examined the statistical evidence supporting the claim in the title.
In a sample of 200 recent studies in the behavioral and brain sciences,
claims of sex-dependent effects were supported only infrequently.

In 48 articles (24%), the sexes were compared statistically and the results supported the claim of a sex-dependent effect.
In 18 articles (9%), the sexes were compared statistically but the results of that comparison were missing.
In 19 articles (9.5%), the sex difference was reported as not statistically significant, which was incompatible with the claim in the title.
In 115 articles (57.5%), the sexes were not statistically compared—no test was performed that could have supported the claim in the title.
If scientists aren’t comparing the sexes statistically, what are they doing?
In most of the articles we sampled, the researchers use a statistical approach widely known to be invalid: they split their sample into two groups (female and male), tested the hypothesis independently within each group, and declared a sex difference because the effect was statistically significant in only one sex. This error has been called the Difference in Sex-Specific Significance (DISS) error. It's an easy error to make -- the senior author of our study has made it a lot!

In the above example, the effect of treatment is statistically significant in the males but not the females; nonetheless, these effects do not differ not meaningfully between the sexes.
Why is the DISS approach invalid?
If we divide our subjects into sub-groups and test them separately, we can fail to detect a real effect in one of those group simply because we have fewer subjects.
A classic example comes from a clinical trial in the 1980s showing that aspirin reduced mortality from heart attacks. The evidence was clear: compared with placebo, aspirin provided an obvious benefit.
To demonstrate the statistical effects of dividing a sample, cardiologist Peter Sleight reanalyzed the data by participants' astrological signs. Once the trial was split into 12 zodiac groups, the benefit was no longer statistically detectable in the Libras and Geminis. When Sleight presented these results, his colleagues immediately understood the joke—that dividing a large sample into smaller subgroups can make a highly significant effect seem to disappear, even when the treatment is clearly effective.

But what is happening with sex categories isn’t as funny. Funding agencies and medical journals often mandate separate analyses by sex. Thus, it’s easy to end up with just enough subjects of one sex to detect an effect, but not quite enough in another. We can wrongly conclude that a treatment works in only one sex, when in fact the effect applied to the entire sample.
What is the solution?
To show appropriate evidence for a sex-dependent effect, the effect should be compared statistically between the sexes. For more information, see our Best Practices article in eLife and our interactive website, SexDifference.org, where you can make your own graphs like the ones above and test for a sex-dependent effect.
But don’t good journals screen out papers with flawed methods?
No. We found that articles with flawed statistical approaches to sex-based data were just as common in high-impact journals and cited just as often as papers with valid approaches.

The appropriateness of statistical evidence was unrelated to the impact factor of the journal. Quartile 1 = highest impact, Quartile 4 = lowest.
The appropriateness of statistical evidence was unrelated to the article's citation rate.
CNCI = Category-Normalized Citation Index.
Why does it matter if false sex differences are reported?
Many of the articles in our sample called for separate treatments for women and men. For example, we saw calls for sex-specific approaches to suicide prevention, stress-related psychiatric disorders, substance use disorder, and psychopathy, all without statistical comparisons across sex. In biomedicine, effects truly “specific” to one sex are likely rare. Recommending different treatments for men and women, without convincing evidence of large sex differences, could create unnecessary or even discriminatory barriers to care, preventing patients from receiving the most beneficial interventions.
Sex-inclusive biomedical research will fulfill its promise only when inclusion is matched by analytical rigor.
Reading list for researchers
