Do we actually need to write research papers at all?
If we are going to lean on the creativity and the cross-domain thinking of human researchers, then surely that should be where the human time goes.
I have written before about what I think is going to change in how academic research gets done, and how the results get out into the world, once AI is properly in the loop. I have been lucky recently. I have been on paternity leave, and I am very grateful to Digital Science for giving me a month with a new baby and, if I am honest, a fair few late nights where the baby was asleep and I was not, watching England not win the World Cup and thinking about academic publishing.
Do we actually need to write papers at all?
I want to be careful here, because this is easy to misread as a claim that the human researcher is now optional. That is not what I think. I have spent a while trying to work out whether large language models are really just very fancy autocomplete, or whether they can be genuine question machines. The research coming out of Google and its co-scientist work suggests that at the very least they can enhance a researcher's creativity when it comes to generating novel research directions. But even if it turns out that the LLMs and the agentic scaffolding around them are never quite both answer AND question machines, I still think there is an enormous amount of disruption coming to the academic space. Not because the machines are geniuses. Because the machines remove friction, and academia is held together by friction.
If we are going to lean on the creativity and the cross-domain thinking of human researchers, then surely that should be where the human time goes. The eureka moment. The interpretation. The intuition that two things nobody has put next to each other belong next to each other.
That is not where the time actually goes. The time goes into the machinery around the insight. Cleaning the data. Running the analysis. Formatting it into the shape a journal will accept. Somewhere in this process, human problems can come into play. There is interpretation drift. There is over-interpretation. There is cherry-picking, where a researcher already knows the story they want to tell and the results get nudged until they tell it. High-impact research is supposed to look a certain way, and so results start to arrange themselves into that certain way.
Why can we not take every dataset, every research output, run it through a machine, and get out a standardised object that carries all of the context of the research, but strips out a huge amount of the interpretation and the sheer time cost of turning a dataset into something other people can build on? So it was serendipitous that Claude Science launched and the getting started guide has this as one of the options “Run a first-pass analysis on a dataset you already have. Point Claude Science at your data (or a public dataset you've been meaning to look at) - it runs QC and a first analysis pass and returns figures plus a written summary.” So I did.

I downloaded some well cited datasets from Figshare and had it do the analysis and create a re-usable, forkable, machine readable paper. Then I did a few more, from Dryad and Zenodo, looking at different topics. There’s now 14 of these papers that look something like this:

Then in a conversation with Graham Smith of Nature, he alerted me to the Data Journalist Agent , by Kevin Qinghong Lin - which turns data into journalism - So from one dataset, we can get the academic paper and the science communication to go with it. The stories they create are pretty cool.

You can see all of this at datasetpapers.com.
It is worth asking what the paper was ever for. Ashish Uppala has a good piece on this, arguing that the journal has quietly done four jobs at once since 1665: it registers who got there first, it certifies that the work is sound, it disseminates the work outward, and it archives it so someone can check it later. Those four were welded together because in the seventeenth century it was cheaper to do them under one roof than to arrange them separately. The internet already made three of them close to free. Registration is a timestamp. Dissemination is a URL. Archival is a repository with a persistent identifier. Certification is the one that never got cheaper, and it is the one everything else is now hanging off.
So when I say the paper stops being the point, I mean something narrower than it sounds. Three of the four jobs no longer need a paper to happen. The fourth still needs something, and the interesting question is what that something is.
A datasetpaper is a versioned, forkable, executable research object built on an open dataset. The data, the code, the environment, the figures, and each individual claim are all separately addressable. The written narrative, the thing we currently treat as the entire output, is just one rendered view of the object. You can read it as a paper if you want a paper. You can read it as a story, evidence-linked, with a click-through viewer on every sentence, if you want that. Or a machine can read it as a bundle of claims, each one carrying its own provenance and its own verification status, and build on a single claim without ever parsing the prose.

Generating an analysis and writing it up used to be the hard part. It is not any more. And once anyone can produce a plausible-looking analysis in minutes, plausibility becomes worthless. The scarce thing is no longer the write-up. The scarce thing is trust. Knowing which claim was actually re-executed from the data, and which one was merely asserted because it sounded right. The verification and the provenance are the product. However, verification is only half of what a journal sells. Uppala splits prestige into two things we usually mash together: trust, meaning the methods are sound and the data is real, and significance, meaning the work matters. Re-execution handles the first. It does nothing at all for the second. Significance has always been an editorial guess, and it is a guess with a poor record, given that Nature turned away Krebs on the citric acid cycle and desk-rejected the modified-mRNA work behind the COVID vaccines inside a day. Uppala's suggestion is that significance is better determined "on read instead of on write", by machines that know what matters to a particular reader, rather than by an editor guessing on everyone's behalf at the moment of submission. I think that is right, and I think it is the harder half.
There is a nice side effect here too, which is speed. This is fast dissemination. Something closer to a research report that comes out dynamically as findings land, rather than a polished paper that emerges eighteen months later once peer review has finished with it. And every time you speed up dissemination (or access eg Open Access publishing), you speed up the point at which the next researcher, or the next model, can start building on what came before. That compounds. It is the single most underrated force in research, and we throttle it constantly.
Too fast publishing?
If you speed up the moment research becomes visible, you have to ask who gets to see it, and when.
Right now there is enormous pressure on researchers around being scooped. Publish or perish dictates the careers of a lot of hard working researchers. People disseminate carefully, defensively, only once they are confident they will get the credit. And into this already anxious system we are now introducing models that would very much like to read your work before you have decided it is ready.
We are already seeing LLMs reaching for research data before the researcher has chosen to publish it. Not the finished paper. The preliminary results. The hypotheses. The half-formed analysis sitting in a collaborative writing tool. OpenAI acquiring the team behind a collaborative writing platform is not a neutral event when you sit it next to a stated ambition to be part of the scientific pipeline. If models start training on pre-published research, at the analysis stage or even earlier, then researchers are being scooped before they have published anything at all.
I spoke about this at London Tech Week recently. My view, is that there has to be a line. The researcher has to approve the moment their work is ready to be disseminated. That has to remain a human choice. You can argue that early access to all of this in-progress science is good for humanity, that it accelerates the whole enterprise, and I would not dismiss that argument. You can also argue it is quietly terrible for the individual human scientist whose one good idea gets absorbed into a training run before they got the paper out. Both things are true at once, which is exactly why we should not let the default just happen to us.
Show me the incentives
Through Figshare, and through the State of Open Data survey we run every year, I keep hearing the same tension. Researchers do not feel they get enough credit for sharing their data. And researchers do not have the time to prepare their data for sharing in the first place. Those two complaints are usually treated as separate. They are the same complaint. It is a friction and reward problem.
So automated workflows that turn research findings into publishable, creditable units, seamlessly, might ease that tension from both ends at once. That is a large part of what datasetpapers is testing. Every datasetpaper starts from someone's dataset, records that debt explicitly, notifies the original depositor that their data was used, and gives them a first-class place in the credit graph. The people who share data have been under-credited for as long as data sharing has existed. A world where machines analyse open data at scale makes that worse, unless the credit is designed in from the start. So it is designed in. There is a cautionary tale here that I had not known about until recently. Between 1961 and 1967 the NIH ran the Information Exchange Groups, a preprint network that circulated unpublished papers by post to a membership that grew past 3,600 scientists. It was essentially arXiv, thirty years early, and it worked. It died because journal editors saw what it was and agreed among themselves to refuse publication to anything that had been circulated through it. Show me the incentives, and I’ll show you the outcomes.
Reducing friction is not enough. Sharing results for the good of mankind is not enough. Solving disease, extraordinarily, is not enough on its own to make a researcher change how they work, if the incentive structure still rewards the old behaviour. Funders can mandate new workflows, and interestingly it often turns out the new workflow is less work than the one they are already asking for. But someone has to align the incentive with the outcome, deliberately, or the better system just sits there being better and unused. So it goes.