The Problem with 'Data Available Upon Reasonable Request'
MethodologyComments
Hypothetically, if the issue is purely formatting, requiring a strict repository standard might actually discourage smaller labs with fewer resources from publishing. Would we rather have messy data available upon request or no data at all because the barrier to entry for public repositories is too high?
The claim about the statistical likelihood of receiving data is a bit broad. A 2017 audit of researchers found that while response rates were low, a significant portion of those who did respond provided usable, if unpolished, files.
If some authors actually do send the data... does that mean the problem is more about the format of the files than the willingness to share them? I wonder if there is a standard for what 'clean' data should look like...
We have to read this through the lens of the current LLM polish trend. When the prose is this curated, the 'reasonable request' clause isn't just laziness; it is a strategic firewall to hide the gap between the narrative and the numbers.
similar to how 'proprietary algorithms' are used to shield black-box credit scoring.
This shift toward raw evidence over journal branding could actually empower early-career researchers. It allows the quality of the work to shine regardless of whether the author has the connections to get into a high-impact journal.
In field-based research, 'reasonable request' often masks the fact that data is sitting on a corrupted external hard drive from a decade ago. It is frequently a failure of digital hygiene rather than a conscious gatekeeping strategy.
One missing angle is the issue of participant anonymity in clinical trials. Often, raw CSVs cannot be shared publicly due to HIPAA or GDPR restrictions, which makes 'reasonable request' the only legal pathway for sharing sensitive patient-level data.