# OpenResearch.wtf > FAIR data, metadata and research infrastructure for an AI-driven science. Essays and working prototypes on what machines need before they can discover. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### About this site URL: https://www.openresearch.wtf/about/ Last updated: 2026-07-31T12:52:03.000Z Open Access, Open Data, Open Science and specifically the metadata, identifiers and infrastructure that decide what machines can actually find, read and reuse. The argument running through these essays is that most research output is still shaped for human readers, and that this has quietly become the bottleneck. If AI is going to do science, the substrate underneath it has to be better than it is. So the writing here tends toward the unglamorous end: data citation corpora, DOI prevalence by country, whether "technically FAIR" and "actually reusable" are the same thing (they are not), and what peer review looks like when a large share of submissions are machine-written. ## Working prototypes Alongside the writing there are experiments: small, public, and built to produce data rather than to become products. You can find more about them at [https://InfiniteResearchers.com/](https://infiniteresearchers.com/?ref=openresearch.wtf) - [**OpenScience.ai**](https://openscience.ai/?ref=openresearch.wtf) — what claims are already latent in open databases like ClinVar, GTEx and Open Targets that nobody has written up? - [**Preprints.ai**](https://preprints.ai/?ref=openresearch.wtf) — how much of peer review is mechanical enough to score automatically? - [**OpenAccess.ai**](https://openaccess.ai/?ref=openresearch.wtf) — if you automate the editorial stack, what happens to the unit economics of publishing? - [**FAIRdata.ai**](https://fairdata.ai/?ref=openresearch.wtf) — can badly described public datasets be made genuinely reusable without a human curator? Each one exists to answer a question I could not answer by arguing about it. ## Who writes this I'm Mark Hahnel. I founded [Figshare](https://figshare.com/?ref=openresearch.wtf) while finishing a PhD in stem cell biology at Imperial College London, on the fairly naive thesis that if you make research data openly available, good things happen that you cannot predict. Figshare now provides research data infrastructure for institutions, publishers and funders globally, and I'm VP Open Research at Digital Science. I sat on the board of [DataCite](https://datacite.org/?ref=openresearch.wtf) and the advisory board for the [Directory of Open Access Journals](https://doaj.org/?ref=openresearch.wtf). I've served on the judging panel for the NIH / Wellcome Trust Open Science Prize and advised on the Springer Nature master classes. More about me, and everything else I write, at [**hahnel.org**](https://hahnel.org/?ref=openresearch.wtf). ## Posts ### Whoever writes the paper, the tools have to be trustworthy URL: https://www.openresearch.wtf/whoever-writes-the-paper-the-tools-have-to-be-trustworthy/ Last updated: 2026-08-06T14:12:07.000Z # In the last post I argued [that the paper might stop being the point](https://www.openresearch.wtf/do-we-actually-need-to-write-research-papers-at-all/), and that machine-generated, forkable research objects could change what a research output even is. That is the speculative end of my thinking and I stand by it. But there is a more immediate version of the same problem sitting on the desk of every researcher right now. Whether the next paper is written by a human, by a machine, or by the two of them arguing in the margins, the tools they use have to be trustworthy, reproducible, and they have to leave the work where it belongs, which is with the person who did it. This week my colleagues at Digital Science launched [Papers AI](https://papers.ai/?ref=openresearch.wtf). One research workspace for your writing, your data and your code, with an assistant that has the whole project in context rather than the paragraph you happened to paste into a chat window. I will come back to what it does. First I want to make the argument for why the shape of the tool matters more than the feature list, because I do not think we are having that argument properly. [![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/08/Screenshot-2026-08-05-at-13.59.10.png)](papers.ai) ## The board is visible in chess. It is not visible in a lab. There is an essay from Thinking Machines, [The Future Worth Building Is Human](https://thinkingmachines.ai/blog/the-future-worth-building-is-human/?ref=openresearch.wtf), that makes the following distinction. There are domains where raw intelligence really is sufficient on its own, and where AI can race ahead without anyone in the loop. Chess is one. Mathematics is increasingly another. What those two have in common is that the goal is static and can be written down, and there is no hidden knowledge. The board is visible to everyone. Nothing that matters is sitting in someone's head, unrecorded. Research is not like that, and I think this is the thing that people building AI-for-science tools keep getting wrong. Most of what a lab knows is not in its papers. It is in the three approaches you already tried that did not work, which will never appear in any published record anywhere. It is in the reason you chose this cohort and not that one, and the argument you had with your collaborator about it. Michael Polanyi called this tacit knowledge, and Hayek's argument about why central planning fails is the same argument: the knowledge that matters is local, provisional, held privately by the people who acquired it through doing the work, and it does not survive being aggregated into one place. Life science researchers have an enormous context window. The warning Thinking Machines make is about AI that "extracts a snapshot of it and replaces it with a standard offering". Apply that to a manuscript in progress and it stops being abstract. A model reading your half-finished draft is not reading a paper. It is reading the tacit layer, the bit that never gets published, at exactly the moment it is most exposed and least protected. ## Show me the incentives Regular readers will know this is where I always end up, so let me apply it to the AI labs rather than to researchers for once. If your product is a single frontier model rented to millions, then every piece of specialised knowledge you can pull into the weights makes the product better for everyone else, and the researcher who supplied it has handed over the one thing that made them worth hiring. [Prism](https://openai.com/index/introducing-prism/?ref=openresearch.wtf) launched on 27 January, built on Crixet, a cloud LaTeX platform OpenAI acquired. Jonathan Schaeffer, emeritus AI professor at Alberta, made a useful split when he [spoke to Decrypt](https://decrypt.co/356259/openai-science-platform-prism-experts-warn-privacy-concerns?ref=openresearch.wtf): there are two things happening when you write a paper, composing the text and doing the actual research. Tools like this help with the former, the drafting and the proofreading and the literature search, and that is real and useful. The problems are everything around the writing. As Schaeffer put it. Standard protocol is that when you write a paper you are documenting your own research and you own it. Route that through a large multinational's model and, in his words, "you're actually exposing your intellectual property to a multinational company". Whether the provider would ever have a legal claim on what you produce inside their tool is, as he says, a question where the devil is in the details. By default, ChatGPT content, and therefore Prism content, may be used to train future models. [Prism is now being folded into codex](https://x.com/thsottiaux/status/2046265345699405880?ref=openresearch.wtf). ## Who holds the line I use frontier models a lot. I see it as a trade off between giving them my thoughts to train on and the incredible stuff I get back. I'm not an academic researcher and this may need more thought for them, because it is a decision about exactly where a line gets drawn across the research lifecycle, and the person drawing it is not the researcher. Drag the line below to see what I mean ## The model reads the work before the researcher has published it. STAGE 01 Drafting & analysis Working data, unfinished drafts, collaborative writing tools STAGE 02 Preprint, dataset, poster The moment the researcher chooses to go public STAGE 03 Peer-reviewed paper & code After review, revision and formal publication Private to the researcher Read by the model Drag the line, tap a stage, or use the arrow keys The point of it is not that one of those three stages is objectively correct. I have my view, which is that the second one is where the trade-off is least bad, but I know serious people who would put it at the third and have good reasons. Wherever the line goes, it should be the researcher who puts it there. Not a vendor, not a default, not a setting three menus deep that got flipped during an update you did not read. If a researcher decides their unpublished analysis is not ready for a model to learn from, that decision has to be enforceable, not merely requested. Almost every tool in this space puts the line at stage one. ## Papers AI moves the line into the architecture Which brings me back to [Papers AI](https://papers.ai/?ref=openresearch.wtf), and to why I think its shape matters more than its feature list. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/08/Screenshot-2026-08-05-at-13.55.54.png) You can write in Word, LaTeX, Markdown or Typst. Real .docx with tracked changes and comments that survive conversion, which anyone who has collaborated across a Word-and-LaTeX divide will tell you is not a small thing. Your datasets and Jupyter notebooks live in the same project as the draft, so the analysis and the prose stop drifting apart. A reviewer asks you to re-run something, you update the cell, and the figure regenerates where the manuscript already is. The assistant reads the whole project, and pulls citations from your own reference library, verifying them before insertion. Every change it makes appears as a reviewable diff. You accept or reject. The assistant's edits are attributed and reversible, the same as a human collaborator's. And you can bring your own model. Use the built-in assistant, or point it at something local running on Ollama or LM Studio or any OpenAI-compatible endpoint. Files stay on your device. Compilers run in your browser. Collaboration happens when you invite it, project by project. [![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/08/Screenshot-2026-08-05-at-13.52.35.png)](https://papers.ai/?ref=openresearch.wtf) That last part is usually written up as a privacy feature. I think it is better understood as the answer to the question the slider asks. Point the assistant at a model running on your own hardware and the embargo line stops being a policy and becomes a fact about where the electricity is going. Nothing has to be promised, because nothing left. Keep the files on the device and the compilers in the browser and the line holds even if every term of service on earth changes tomorrow. If the risk is that a centralised model absorbs what makes your lab distinctive, the mitigation is not a stronger assurance from the provider. It is not having sent it. You cannot extract a snapshot of a lab's tacit knowledge from a model that never left the building. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/07/data-src-image-f9d8d4f4-ae5c-4d5f-9327-809ecf6635d4.png) Local inference is not exotic in 2026, and anyone could build this way. Building the version where the researcher keeps the line is a decision about who the tool is for, and it is the decision I would want made by whoever I trusted with an unpublished manuscript. We ran a webinar on exactly this last week, [AI Uncorked: guardrails, compliance and IP protection for the enterprise](https://www.overleaf.com/events/webinars/ai-uncorked-guardrails-compliance-and-ip-protection-for-the-enterprise?ref=openresearch.wtf), with John Lees-Miller, who co-created Overleaf, alongside Mark Bloomfield and the ECB's Maximilian Freier. The whole conversation was about the question underneath all of this. Not whether to adopt AI in research, because that ship has sailed, but how to do it without putting sensitive data, proprietary research and IP at risk. How to build guardrails so teams can use these tools with confidence. How to make sure that the research which is yours, stays yours. In the [last piece](https://www.openresearch.wtf/do-we-actually-need-to-write-research-papers-at-all/) I got excited about a future where research outputs are forkable, verifiable, machine-readable objects that people and agents can build on at speed. I still want that future. But a forkable research object is only as trustworthy as the tools that generated it and the tools that check it. Reproducibility is not a nice-to-have bolted on at the end, it is the [substrate](https://www.openresearch.wtf/ground-the-model-or-it-invents-the-evidence/). And if we are going to hand more of the writing, the citing and eventually the analysis to machines, then the tools doing that work have to verify before they assert, keep provenance attached, and leave ownership with the human whose name is on the paper. In chess, intelligence alone gets you all the way, because there is nothing hidden. Life science research is not chess (yet). Maths and Physics will be solved long before wet lab research. Most of what a lab knows was never written down, and the tools we adopt now will decide whether that knowledge gets cultivated where it is, or quietly extracted and sold back to us as a standard offering. Whoever ends up writing the papers, humans, machines, or both, they need better tools and [Papers AI](https://www.openresearch.wtf/the-perpetual-research-cycle-ais-journey-through-data-papers-and-knowledge/) is one of them. ### Do we actually need to write research papers at all? URL: https://www.openresearch.wtf/do-we-actually-need-to-write-research-papers-at-all/ Last updated: 2026-07-31T12:35:38.000Z I have written before about what I think is going to change in how academic research gets done, and how the results get out into the world, once AI is properly in the loop. I have been lucky recently. I have been on paternity leave, and I am very grateful to Digital Science for giving me a month with a new baby and, if I am honest, a fair few late nights where the baby was asleep and I was not, watching England not win the World Cup and thinking about academic publishing. Do we actually need to write papers at all? I want to be careful here, because this is easy to misread as a claim that the human researcher is now optional. That is not what I think. I have spent a while trying to work out whether large language models are really just very fancy autocomplete, or whether they can be genuine question machines. The research coming out of Google and its [co-scientist](https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/?ref=openresearch.wtf) work suggests that at the very least they can enhance a researcher's creativity when it comes to generating novel research directions. But even if it turns out that the LLMs and the agentic scaffolding around them are never quite both answer AND question machines, I still think there is an enormous amount of disruption coming to the academic space. Not because the machines are geniuses. Because the machines remove friction, and academia is held together by friction. If we are going to lean on the creativity and the cross-domain thinking of human researchers, then surely that should be where the human time goes. The eureka moment. The interpretation. The intuition that two things nobody has put next to each other belong next to each other. That is not where the time actually goes. The time goes into the machinery around the insight. Cleaning the data. Running the analysis. Formatting it into the shape a journal will accept. Somewhere in this process, human problems can come into play. There is interpretation drift. There is over-interpretation. There is cherry-picking, where a researcher already knows the story they want to tell and the results get nudged until they tell it. High-impact research is supposed to look a certain way, and so results start to arrange themselves into that certain way. Why can we not take every dataset, every research output, run it through a machine, and get out a standardised object that carries all of the context of the research, but strips out a huge amount of the interpretation and the sheer time cost of turning a dataset into something other people can build on? So it was serendipitous that [Claude Science](https://claude.com/product/claude-science?ref=openresearch.wtf) launched and the getting started guide has this as one of the options “Run a first-pass analysis on a dataset you already have. Point Claude Science at your data (or a public dataset you've been meaning to look at) - it runs QC and a first analysis pass and returns figures plus a written summary.” So I did. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/07/data-src-image-abfe512d-d9b7-40a7-bc4b-a735081bc43d.png) I downloaded some well cited datasets from Figshare and had it do the analysis and create a re-usable, forkable, machine readable paper. Then I did a few more, from Dryad and Zenodo, looking at different topics. There’s now 14 of these papers that look something like this: ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/07/datapapers3.gif) Then in a conversation with Graham Smith of Nature, he alerted me to the [Data Journalist Agent](https://data2story.github.io/?ref=openresearch.wtf) , by [Kevin Qinghong Lin](https://arxiv.org/search/cs?searchtype=author&query=Lin%2C+K+Q&ref=openresearch.wtf) \- which turns data into journalism - So from one dataset, we can get the academic paper and the science communication to go with it. The stories they create are pretty cool. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/07/datasetstory-1.gif) You can see all of this at [datasetpapers.com](https://datasetpapers.com/?ref=openresearch.wtf). It is worth asking what the paper was ever for. [Ashish Uppala has a good piece on this](https://www.shishyko.com/supercritical/coordination-tech-in-science?ref=openresearch.wtf), arguing that the journal has quietly done four jobs at once since 1665: it registers who got there first, it certifies that the work is sound, it disseminates the work outward, and it archives it so someone can check it later. Those four were welded together because in the seventeenth century it was cheaper to do them under one roof than to arrange them separately. The internet already made three of them close to free. Registration is a timestamp. Dissemination is a URL. Archival is a repository with a persistent identifier. Certification is the one that never got cheaper, and it is the one everything else is now hanging off. So when I say the paper stops being the point, I mean something narrower than it sounds. Three of the four jobs no longer need a paper to happen. The fourth still needs something, and the interesting question is what that something is. A [datasetpaper](https://datasetpapers.com/?ref=openresearch.wtf) is a versioned, forkable, executable research object built on an open dataset. The data, the code, the environment, the figures, and each individual claim are all separately addressable. The written narrative, the thing we currently treat as the entire output, is just one rendered view of the object. You can read it as a paper if you want a paper. You can read it as a story, evidence-linked, with a click-through viewer on every sentence, if you want that. Or a machine can read it as a bundle of claims, each one carrying its own provenance and its own verification status, and build on a single claim without ever parsing the prose. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/07/Screenshot-2026-07-24-at-14.59.29.png) Generating an analysis and writing it up used to be the hard part. It is not any more. And once anyone can produce a plausible-looking analysis in minutes, plausibility becomes worthless. The scarce thing is no longer the write-up. The scarce thing is trust. Knowing which claim was actually re-executed from the data, and which one was merely asserted because it sounded right. The verification and the provenance are the product. However, verification is only half of what a journal sells. Uppala splits prestige into two things we usually mash together: trust, meaning the methods are sound and the data is real, and significance, meaning the work matters. Re-execution handles the first. It does nothing at all for the second. Significance has always been an editorial guess, and it is a guess with a poor record, given that *Nature* turned away Krebs on the citric acid cycle and desk-rejected the modified-mRNA work behind the COVID vaccines inside a day. Uppala's suggestion is that significance is better determined "on read instead of on write", by machines that know what matters to a particular reader, rather than by an editor guessing on everyone's behalf at the moment of submission. I think that is right, and I think it is the harder half. There is a nice side effect here too, which is speed. This is fast dissemination. Something closer to a research report that comes out dynamically as findings land, rather than a polished paper that emerges eighteen months later once peer review has finished with it. And every time you speed up dissemination (or access eg Open Access publishing), you speed up the point at which the next researcher, or the next model, can start building on what came before. That compounds. It is the single most underrated force in research, and we throttle it constantly. ## Too fast publishing? If you speed up the moment research becomes visible, you have to ask who gets to see it, and when. Right now there is enormous pressure on researchers around being scooped. Publish or perish dictates the careers of a lot of hard working researchers. People disseminate carefully, defensively, only once they are confident they will get the credit. And into this already anxious system we are now introducing models that would very much like to read your work before you have decided it is ready. We are already seeing LLMs reaching for research data before the researcher has chosen to publish it. Not the finished paper. The preliminary results. The hypotheses. The half-formed analysis sitting in a collaborative writing tool. OpenAI acquiring the team behind a collaborative writing platform is not a neutral event when you sit it next to a stated ambition to be part of the scientific pipeline. If models start training on pre-published research, at the analysis stage or even earlier, then researchers are being scooped before they have published anything at all. I spoke about this at London Tech Week recently. My view, is that there has to be a line. The researcher has to approve the moment their work is ready to be disseminated. That has to remain a human choice. You can argue that early access to all of this in-progress science is good for humanity, that it accelerates the whole enterprise, and I would not dismiss that argument. You can also argue it is quietly terrible for the individual human scientist whose one good idea gets absorbed into a training run before they got the paper out. Both things are true at once, which is exactly why we should not let the default just happen to us. ## Show me the incentives Through Figshare, and through the [State of Open Data](https://doi.org/10.6084/m9.figshare.30823079?ref=openresearch.wtf) survey we run every year, I keep hearing the same tension. Researchers do not feel they get enough credit for sharing their data. And researchers do not have the time to prepare their data for sharing in the first place. Those two complaints are usually treated as separate. They are the same complaint. It is a friction and reward problem. So automated workflows that turn research findings into publishable, creditable units, seamlessly, might ease that tension from both ends at once. That is a large part of what datasetpapers is testing. Every datasetpaper starts from someone's dataset, records that debt explicitly, notifies the original depositor that their data was used, and gives them a first-class place in the credit graph. The people who share data have been under-credited for as long as data sharing has existed. A world where machines analyse open data at scale makes that worse, unless the credit is designed in from the start. So it is designed in. There is a cautionary tale here that I had not known about until recently. [Between 1961 and 1967 the NIH ran the Information Exchange Groups](https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.2003995&ref=openresearch.wtf), a preprint network that circulated unpublished papers by post to a membership that grew past 3,600 scientists. It was essentially arXiv, thirty years early, and it worked. It died because journal editors saw what it was and agreed among themselves to refuse publication to anything that had been circulated through it. Show me the incentives, and I’ll show you the outcomes. Reducing friction is not enough. Sharing results for the good of mankind is not enough. Solving disease, extraordinarily, is not enough on its own to make a researcher change how they work, if the incentive structure still rewards the old behaviour. Funders can mandate new workflows, and interestingly it often turns out the new workflow is less work than the one they are already asking for. But someone has to align the incentive with the outcome, deliberately, or the better system just sits there being better and unused. So it goes. ### BioSingularity. The substrate determines the ceiling URL: https://www.openresearch.wtf/biosingularity-the-substrate-determines-the-ceiling/ Last updated: 2026-07-31T12:35:14.000Z ## Why the infrastructure underneath the models decides how far they can go. In the past two years a phrase has been circulating in the AI-in-biology community that I find genuinely useful for thinking about where the field is heading. People talk about a “biosingularity”, meaning the point at which biology becomes computationally tractable in the way that, weather or fluid dynamics already are. The cell becomes a system you can simulate. The drug becomes a thing you can design before you make. Jensen Huang has been making the same argument from the NVIDIA side, framing it as biology shifting from a life science into life engineering. I agree with this hypothetical end goal. The path to it is still very bloody complex. An AI system is only ever as good as the verifiability and the provenance of the thing it reasons over. We call this the substrate. The substrate determines the ceiling. ## **Biology is trying to build a substrate it can trust** ![scales_of_biology.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-8e7984f9-7549-48df-84d8-ce5febd622a1.png "scales_of_biology.png") *Figure 1\. Foundation models now span almost every scale of biology. The substrate they rely on is the cross-cutting requirement.* There are now foundation models at almost every scale of biology. AlphaFold and its successors at the molecular level. Models like [scGPT](https://www.nature.com/articles/s41592-024-02201-0?ref=openresearch.wtf) and [Geneformer](https://www.nature.com/articles/s41586-023-06139-9?ref=openresearch.wtf) at the single-cell level. [UNI and Virchow](https://arxiv.org/abs/2509.15482?ref=openresearch.wtf) at the pathology level. Organ digital twins above that. It is fantastic that we are building these things. The Cell Perspective by the Chan Zuckerberg-led group on the [AI Virtual Cell](https://virtualcellmodels.cziscience.com/?ref=openresearch.wtf), along with the Arc Institute [Virtual Cell Challenge](https://virtualcellchallenge.org/?ref=openresearch.wtf) launched in 2025, gives a sense of how much organised effort (and money) is now going into building the next layer of the stack. There is however, a caveat. A model fit to observational data only ever learns the correlation structure of that data. It learns the patterns, and it learns the biases, and it cannot tell the difference between them without help. The single-cell company Noetik has described its own model’s outputs as “[not at all guaranteed to be causal, merely strongly correlated](https://www.noetik.blog/p/how-do-you-use-a-virtual-cell-to-6c6?ref=openresearch.wtf).” Recent benchmarking by [Ahlmann-Eltze and colleagues](https://www.nature.com/articles/s41592-025-02772-6?ref=openresearch.wtf) found that five state-of-the-art perturbation models could not beat a simple linear baseline when tested on gene knockouts they had genuinely never seen. The models are excellent at interpolating within what they have observed, and unreliable outside of that. The fix the field has converged on is to engineer causal structure into the substrate itself. Rather than gathering more observational data, the move is to feed the models perturbation experiments, the results of actually intervening on a system rather than just watching it. You build a substrate that encodes cause as well as correlation, and the model’s ceiling rises accordingly. The Virtual Cell Challenge is, in a sense, a community-scale attempt to do exactly this in public. ## **Software can do this, because its substrate was verifiable all along** The reason AI agents arrived in software engineering before almost anywhere else is that code is executable and feedback-rich. [Kenny Workman of LatchBio](https://blog.latch.bio/p/agentic-biology-is-shaped-like-software?ref=openresearch.wtf) makes this argument well. You can run code, watch it fail, read the error, and try again. The substrate hands you ground truth on every loop. Or as he puts it “Agentic biology is shaped like software“ ![verifiability_gradient.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-0af91096-b580-4c5a-8a99-4bb2c52cafc5.png "verifiability_gradient.png") *Figure 2\. The gradient from locally verifiable to globally verifiable. After Kenny Workman (2026).* Take a hard scientific question, like whether a particular set of gene mutations drives a disease. At the top level it has no global verifiability, with no single correct answer to grade against. But it decomposes into smaller steps, some of which are locally verifiable. Did this cell line pass quality control? Do these genes show differential expression? Each of those has a crisp answer. The way to answer this is to build up from the locally verifiable steps toward the globally uncertain claim. That gradient, from locally verifiable to globally verifiable, is the single most useful lens I have found for the biology problem. It tells you where progress is fast (the left end) and where it is slow (the right end). It also tells you what kind of infrastructure each end needs. ## **What this looks like in drug discovery** For pharma, the locally verifiable layer is where most of the immediate value sits. Was this target mentioned in a paper? Was that paper funded by which grants? Did the trials supporting it use the right comparator? Has the molecule been characterised in patents? Each of these is a crisp yes-or-no, and each is the kind of question that AI ought to be able to answer with confidence, but only if the underlying substrate is structured and provenance-bearing rather than a soup of free text. This is where the kind of resource we have spent a long time building at Digital Science comes in. The [Dimensions Knowledge Graph](https://www.dimensions.ai/products/all-products/dimensions-knowledge-graph/?ref=openresearch.wtf) links publications, grants, patents, clinical trials, datasets and policy documents into a single typed structure. Every entity carries a persistent identifier. Every relationship can be traced back to its source. If a model says “this target appears in seventeen recent papers funded by this consortium and three clinical trials at this phase,” that string is not coming from a statistical pattern. It is coming from edges in a graph. The model becomes a query interface to a substrate, and the substrate carries the truth. The same principle holds one layer down. Figshare exists in part because the perturbation data that biology is now scrambling to generate needs persistent identifiers, version history and provenance metadata, or it cannot be reused. A virtual cell model trained on undocumented data is a virtual cell model whose ceiling is set by the lowest-quality dataset it ingested. The persistent identifier (DOI for publications and datasets, ORCID for researchers, ROR for institutions) is the plumbing that decides whether any of this works downstream. ## **FAIR has been pointing here for a decade** The FAIR data principles, findable, accessible, interoperable, reusable, have been making this argument since 2016, long before anyone needed a large language model to care. For most of that decade they sounded like infrastructure housekeeping, the sort of thing you nod along to and deprioritise. What has changed is that AI has suddenly made the cost of an unverifiable substrate visible, and urgent, and measurable in failed drug trials. A trial fails for many reasons, but a meaningful fraction of late-stage failures trace back to evidence chains that no one ever audited at the start. If the original target paper relied on a misinterpreted figure, or the dose-finding work cited a study with a flawed control arm, that information was always in principle recoverable. It just was not connected to the system that picked the molecule. A substrate with provenance built in does that connecting by default. ## **We’re on a slope as opposed to waiting for a Eureka moment** The biosingularity is probably better understood as a slope than a moment. The labs and companies that climb it fastest will be the ones whose models reason over the most verifiable, best-provenanced substrate, rather than the ones with the cleverest models. The ceiling is set by the ground, not by what stands on it. We need to continue to improve provenance and reproducibility in research. We need to continue to push for Open Science. The infrastructure work, persistent identifiers and knowledge graphs and perturbation datasets and FAIR-compliant repositories, is what determines how high the ceiling rises. We are getting better at this as a research community. Preprints help. So does open data. The substrate beneath biology is next, and the alternative is the same one playing out elsewhere in the research ecosystem. Ungrounded generation rushes in to fill the gap, with things that were never true. As mentioned in my last post “[We have a deadline, because something is already busy filling the substrate with things that were never true.](https://www.openresearch.wtf/ground-the-model-or-it-invents-the-evidence/)” ### Ground the model, or it invents the evidence URL: https://www.openresearch.wtf/ground-the-model-or-it-invents-the-evidence/ Last updated: 2026-07-31T12:35:41.000Z ## Trusted, curated, subject-specific models are needed for academia Earlier this year [a thread from Tom Dietterich](https://x.com/tdietterich/status/2055000956144935055?ref=openresearch.wtf), who chairs arXiv’s computer science section, made the rounds. He was clarifying the consequences for unchecked LLM output on the platform. If a submission contains “incontrovertible evidence” that the authors did not check what the model produced (hallucinated references, or worse, model meta-comments still sitting in the manuscript like “here is a 200 word summary, would you like me to make any changes?”) the authors are banned from arXiv for a year. In my opinion, the policy is needed and appropriate. The [underlying problem](https://www.openresearch.wtf/academia-has-a-new-preprints-problem/) is bigger than arXiv. ![citation_rates_by_repository.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-ecbba5e5-bd0c-44a9-afb5-a54c535eaf79.png "citation_rates_by_repository.png") *Figure 1\. Citation-hallucination rates and monthly counts by repository, as of August 2025\. Source: Zhao et al. (2026), arXiv:2605.07723.* A letter in [The Lancet by Maxim Topaz](https://www.thelancet.com/journals/lancet/article/PIIS0140-6736%2826%2900603-3/fulltext?ref=openresearch.wtf) and colleagues found that one in every 277 PubMed-indexed papers in early 2026 cited something that does not exist. In 2023 the figure was one in 2,828\. A much larger audit by Zhao, Yin and colleagues at Cornell, Berkeley, Tsinghua and UCLA looked at 111 million references across arXiv, bioRxiv, SSRN and PubMed Central. Their conservative estimate is that at least [146,932 hallucinated citations entered the scientific record in 2025](https://arxiv.org/pdf/2605.07723?ref=openresearch.wtf) alone. All four repositories climbed sharply from mid-2024\. By August 2025 the hallucination rate ranged from 0.21% on bioRxiv to 1.91% on SSRN, and because PubMed Central is so much larger, it carried the highest monthly count even at a rate of 0.27%. These numbers are largely about general-purpose language models being asked to do something they were not designed for. Generating prose is what they do well. Verifying that a citation refers to a real paper is not in their job description, and the architecture does not support it. They are statistical prediction engines that produce the most plausible-looking string given everything they have absorbed. This is a problem when it claims to build on the permanent record of research. ## **The hallucinations have a direction** When fabricated citations name real authors, they do not name them at random. This is likely true of existing research practices, but if we’re entering a new era of research, it would be cool if we got rid of some of the old problems. The Cornell team checked whether this bias was confined to the obviously fake citations, or whether it showed up in the perfectly real ones sitting beside them in the same bibliographies. It does. The real citations in papers that also contain hallucinations skew toward the same prominent authors. ![credit_skew.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-9808cc09-d80a-421b-88eb-4482a3da32ec.png "credit_skew.png") *Figure 2\. How hallucinated citations differ from matched genuine ones. Source: Zhao et al. (2026), arXiv:2605.07723.* Compared with matched genuine citations, hallucinated citations credit scholars with 68.8% more prior publications and 58.3% more citations. The credited names also skew male, by 6.4 percentage points. Prominence and plausibility are entangled in the training data, since the most-cited authors appear in the most reference lists. The model has learned that their names belong in reference lists. When it fabricates, it fabricates those who would most likely turn up in a reference list. ![matthew_loop.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-cef65fbd-740b-4f7a-8c3c-de800667b911.png "matthew_loop.png") *Figure 3\. The Matthew effect, automated.* The [Matthew effect](https://en.wikipedia.org/wiki/Matthew%5Feffect?ref=openresearch.wtf), Robert Merton’s name for the rich-get-richer dynamic in citation, has always existed. What is new is that we may be about to automate it at a scale and speed that no funding panel or promotion committee can keep up with. ## **Why a general model is the wrong tool for research** The pattern in the data points to a specific kind of mistake. A general-purpose language model has no representation of the world’s actual citation graph. Asking it whether a paper is real, or whether it supports a claim, is asking a question its architecture does not know how to answer correctly. The result is a confident-sounding answer that is roughly correlated with the truth and fails in predictable ways. This is something I saw when building [https://datacitations.com/](https://datacitations.com/?ref=openresearch.wtf), A knowledge graph linking 9.3 million data citations across nearly 2,000 data repositories, that demonstrates how FAIR principles power trustworthy, traceable AI answers grounded in structured evidence. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-13b0940c-931e-4320-b4dd-274e3e9cd708.png) The interesting observation, and the reason I think this problem is more tractable than the discourse suggests, is that science already has the infrastructure to do this properly. The infrastructure simply has not been wired up to the place where the generation is happening. ![architectures.png](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-76932412-7e8d-4281-9144-f5230a350c4e.png "architectures.png") *Figure 4\. Two architectures, two failure modes.* Every published paper has a DOI. Every dataset can have one too. Every author can have an ORCID. Every institution can have a ROR identifier. Crossref will resolve any DOI in milliseconds. OpenAlex, DataCite and Semantic Scholar expose this as queryable open data. At Digital Science the [Dimensions Knowledge Graph](https://www.dimensions.ai/products/all-products/dimensions-knowledge-graph/?ref=openresearch.wtf) holds these relationships in typed, persistent form, connecting publications, grants, patents, clinical trials, datasets and people. Figshare hosts versioned datasets with their own DOIs. The Cornell team’s own validation pipeline matched 95% of references on first pass with off-the-shelf tools. The reason 78% of fabricated citations still pass arXiv moderation comes down to one thing. The validation is not being run. The substrate exists. The plumbing connecting it to the systems that produce citations does not. A subject-specific model that consults this substrate before it speaks has a different failure mode than a general model that does not. It will sometimes say “this paper is not in my graph” rather than inventing one. It will sometimes return a smaller set of citations than the user expected. It will sometimes flag conflicts in the underlying evidence. None of these are exciting demo moments. All of them are what scientific reasoning is supposed to look like. ## **Will we move to a GitHub-like way of building?** The longer I sit with this, the more the analogy to software seems instructive. Software engineering went from a craft where you copy-pasted from forums to one where every dependency has a version number, a hash, an authorship trail, a license, and a public history of who changed what when. It is the reason agents can now reason about software at all, because the substrate is verifiable on every loop. The equivalent in scholarly publishing would be a citation graph where every reference resolves to a persistent identifier, every claim has a recorded provenance trail back to source data, every researcher’s contribution is attached to their ORCID, and every dataset has a version on a repository like Figshare with its own DOI. None of these are technologically hard. Most of them already exist, it feels like we’re close, but research is not (yet) happening this way. arXiv’s new policy is one small step in that direction. So are journal requirements to validate references against Crossref at submission. So is the slow expansion of registered reports and data-availability requirements. None of these is a complete solution. They are pieces of a substrate. ## **The deadline** The thing that makes this urgent is the second-order effect, more than the absolute number of hallucinated citations today. Hallucinated citations are getting indexed in Google Scholar as standalone bibliographic entries. They are being ingested into the training corpora of the next generation of models. The contamination, as Maxim Topaz put it in the coverage of his paper, does not go away when the AI gets better. The longer the substrate is left ungoverned, the more of the next-generation models will be trained on something that includes things that were never true. We have a deadline, because something is already busy filling the substrate with things that were never true. I’m an optimist. This is a problem science is well-equipped to solve, because every piece of the infrastructure already exists. The problem is that research has a lot of problems that could be solved if cultural changes amongst the research community happened faster. In this new age of acceleration, maybe we’ll finally start optimising the research process. ### The new era of going “Fast and Far” in research URL: https://www.openresearch.wtf/the-new-era-of-going-fast-and-far-in-research/ Last updated: 2026-07-31T12:35:57.000Z ### Small teams have always innovated. But AI means a) We can split research into big and small teams b) The small can do absolutely massive things now ### In October 2024, a Nobel Prize in Chemistry went, in effect, to a piece of software. More precisely it went to the people who built it. Demis Hassabis and John Jumper at Google DeepMind took half the prize, sharing the rest with David Baker in Seattle. The thing they had made, AlphaFold, did in a few years what structural biology had been chipping away at for half a century. It predicted the three-dimensional shape of a protein from its amino acid sequence, and then it did it for almost every protein science had ever named. Around 200 million of them. What really amazed me was the size of the project team. Just 19 names share the lead credit on the 2021 AlphaFold2 paper, the block marked as having contributed equally, and the full author list runs to thirty-three. This is a problem that had defeated the global biochemistry community for 50 years, only to be solved by an AI-enabled (and compute) small team. Does this mean that the old adage, “If you want to go fast, go alone. If you want to go far, go together,” is no longer true in an AI-powered, agentic world? Usually attributed to West Africa, usually wheeled out to justify whatever the speaker had already decided to do. And for as long as I have been in research, it made sense. It described a genuine trade-off. You picked one. Speed or distance. The lone researcher sprints and burns out. The consortium endures and crawls. This is measurable. Lingfei Wu, Dashun Wang and James Evans went through more than 65 million papers, patents and software products spanning 60 years of output, and [published the result in Nature in 2019](https://doi.org/10.1038/s41586-019-0941-9?ref=openresearch.wtf). Large teams develop. Small teams disrupt. Small teams reach further back into older ideas, produce the more destabilising work, and pay off further into the future, if they pay off at all. Big teams take an existing line and push it forward. Go far, or go fast. ## **The trade-off breaks** The AlphaFold team was disruptive in exactly the Wu, Wang and Evans sense, reaching back to an old problem most people assumed was decades from solution. But it did not stay narrow, and it did not take a generation to land. Its predictions have since been used by more than two million researchers in 190 countries, and the core work was done inside about 18 months. Fast and far, from the same 19 people. AI is what collapsed it. A small team with enough model leverage now reaches as far as a consortium and keeps the speed of a startup. Go fast and far, alone. A lot of the data and models are open. Are we about to see a new class of small-team research disruptors?I pulled the author lists on the breakthroughs of the past three years where AI did the discovering, across several fields. The pattern holds. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-69e05632-b347-45bc-8397-363b0b13539b.png) *AI-enabled scientific breakthroughs of the past three years, by team size. Sources are Nature and Science. AlphaFold3, at 48, is the largest team in the set.* Six people built [GNoME](https://www.nature.com/articles/s41586-023-06735-9?ref=openresearch.wtf) and predicted 2.2 million new crystals, around 380,000 of them stable, roughly a tenfold expansion in the stable inorganic materials known to science. Six more, at Huawei, built [Pangu-Weather](https://arxiv.org/abs/2211.02556?ref=openresearch.wtf) and beat the leading operational forecast. Sixteen built AlphaMissense and produced a verdict on 89 per cent of all possible human missense mutations, a resource clinical geneticists now draw on. And before anyone objects that this is big tech doing it with pure compute, look at the wet-lab end. [The new structural class of antibiotics that kills MRSA](https://www.nature.com/articles/s41586-023-06887-8?ref=openresearch.wtf), the first new class in 60 years, came from screening more than 12 million compounds with a graph neural network and then synthesising the survivors and testing them in mice. Twenty-one people, across MIT, the Broad and two other institutes. [Berkeley's A-Lab](https://www.nature.com/articles/s41586-023-06734-w?ref=openresearch.wtf), a robotic chemistry lab that planned, ran and interpreted its own experiments to make dozens of new compounds, took 16\. The two largest teams in the set are biology at the cutting edge, the Baker lab's protein design at 28 and AlphaFold3 at 48, and both are still small by the standards of the field they sit in. ## **So are these really small teams?** A short author list can hide a long shadow. Behind those six names on the Pangu-Weather paper sits Huawei Cloud. Behind AlphaFold sits the whole of DeepMind. And every one of these models learned from a dataset that took a large, slow, collaborative effort to build. So it's not like “six people did this alone”. It is “six people did this on top of what an institution spent decades assembling”. But either way, these small teams are doing in months what used to take a generation and a large workforce. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/Screenshot-2026-06-04-at-12.18.45.png) So both things are true at once. The teams really are small, and they really are doing what once took huge teams. The trick is that the huge team already ran. It built the substrate and left it in an open database where a handful of people with a good model could pick it up. Tier one takes an army and a decade. Tier two takes a dozen people and a few months. ## **Standing on the shoulders of giants** This is the bit that matters most if you care about research infrastructure. AlphaFold did not arrive from nowhere. It learned to predict structures by training on the Protein Data Bank, which is to say it learned from 50 years of slow, expensive, collaborative experimental work. Every structure in that archive was solved by someone at a beamline or an NMR machine or a cryo-EM rig, then deposited, curated and shared. EMBL had been generating that ground truth for decades. Jumper has said it plainly: public data were essential. So the small fast team did not replace the large slow collaboration. It stood on its shoulders. Once you see that, the two tiers stop looking like rivals and start looking like what they are, which is stacked. We’re seeing big philanthropic orgs picking up the heavy lifting when it comes to big data and models - [biohub](https://biohub.org/ai-models/?ref=openresearch.wtf), [Allen AI](https://allenai.org/?ref=openresearch.wtf), [Astera](https://astera.org/foundation/?ref=openresearch.wtf) are all making massive amounts of information open to all. With folks predicting the[ third wave of American philanthropy](https://nanransohoff.substack.com/p/the-third-wave-of-american-philanthropy), that AI will likely add \~$37–100B per year in intended philanthropic spend in the near future - the future for open science looks bright. It looks like the “go far, go together” side of things is taken care of. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/06/data-src-image-42ef772a-093d-4d8c-bec5-ba9ce4a20cf6.png) *The same era's defining results, plotted by team size on a log scale. The Higgs boson mass paper carried 5,154 authors. A later ATLAS collaboration paper reached 8,778\. The AI breakthroughs sit three orders of magnitude below, with almost nothing in between.* ## **Two tiers, one stack** That gap in the middle of the chart is the story. Modern science is settling into two tiers. Tier one is the slow business of building the substrate: the atlases, the reference datasets, the standardised measurements no single lab could produce alone. It shows up in those author counts of five and eight thousand. It is how you measure a Higgs boson, and the most ambitious biological example of it now is the effort to model the cell. The Human Cell Atlas set out to catalogue every human cell type and had to invent the methods as it went. The Chan Zuckerberg Initiative has since put real weight behind a virtual cell, standing up one of the largest non profit GPU clusters in the life sciences, launching a Billion Cells Project with 10x Genomics and Ultima Genomics, and committing more than 10 billion dollars to basic research over a decade. The Arc Institute trained its first virtual cell model on more than 170 million cells. These are not papers. They are infrastructure, and they move at the pace of consortia. Tier two is the small team with AI leverage. The PDB worked as a launchpad because it was open, it was standardised, it carried identifiers, and a machine could read it without a human in the loop. That is FAIR data doing the quiet work nobody hands out prizes for. Build the cell atlases of this decade to the same standard and they become launchpads too. Build them as a thousand incompatible spreadsheets behind a thousand logins and they don’t. ## **However small the team, you are still a team** Not one of those breakthroughs has a single author (yet). A team of 19 co-authoring a 60page paper with 32 component algorithms is not 19 people working alone. It is 19 people who have to think, write and ship as if they were one. A small team can go fast not because collaboration stopped mattering, but because its friction dropped close to zero. There is a reason why collaborative tools like [Overleaf](https://www.overleaf.com/?ref=openresearch.wtf) have 25 million users. Everyone writing together, in the same place, in real time, with the version history and the structure and the maths held in one place. Modern problems require modern solutions. So I think the proverb needs a rewrite for the age of large models. If you want to go fast and far, go as a small team, give it real leverage, stand it on an open foundation someone else built, and give it tools that make collaboration close to effortless. ### What parts of the academic knowledge creation & dissemination pipeline can we automate with AI? URL: https://www.openresearch.wtf/what-parts-of-the-academic-knowledge-creation-dissemination-pipeline-can-we-automate-with-ai/ Last updated: 2026-07-31T12:36:00.000Z As a tech founder who can, at best, describe himself as a hacky coder, I've spent most of my career having to either hire engineers or do without when an idea was too technically demanding. That's changed quite a lot recently. Tools like Claude Code mean I can now actually build the things I want to test, which makes experimenting in this space genuinely good fun in a way it wouldn't have been two or three years ago. So with that said - what parts of the academic knowledge creation and dissemination pipeline can we automate thanks to AI? Below, I describe 4 platforms that try to individually answer parts of this question, and when combined, attempt to answer a little bit more of it. As a disclaimer, everything discussed here is experimental. The platforms described are not robust scholarly infrastructure. They should not be cited, relied upon, or treated as production systems. They're a hobby project built on open APIs, open data, and open infrastructure designed to test what's technically possible. **The problems worth solving** We've built good infrastructure for sharing research outputs, yet most of the value locked in those outputs goes unrealised. Meanwhile, the research community is under pressure in almost every dimension: time, funding, reviewer bandwidth, editorial capacity. A disproportionate amount of effort goes on tasks that are fairly mechanical - formatting metadata, screening submissions for integrity, summarising prior literature, running standard statistical checks. New academic knowledge creation and basic research has never been more in demand. AI will never replace this. But there is a subset of academic knowledge gap filling that new tools seem capable of filling. Can they? What if some of those tasks could be automated? I am not attempting here answer the question of “What should be automated?” **The four platforms, briefly** ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-e18c4691-4914-46fe-8ded-d3686858732b.png) [FAIRdata.ai](http://fairdata.ai/?ref=openresearch.wtf) starts at the data layer. A large proportion of datasets deposited on generalist repositories are technically FAIR but practically unusable. The metadata is incomplete, the variable labels are opaque, the context is missing. FAIRdata.ai tries to fix that automatically, enriching metadata to improve FAIR scores. It also tries to make use of surprisal scoring and automated discovery techniques to surface findings within datasets that may never have been noticed by the original depositors. Those signals feed into [OpenScience.ai](http://openscience.ai/?ref=openresearch.wtf).You can read more details about [the technical capabilities and dependencies of FAIRdata.ai here.](https://www.openresearch.wtf/fairdata-ai-making-open-data-fair-er/) [OpenScience.ai](http://openscience.ai/?ref=openresearch.wtf) is the hypothesis engine. It queries open scientific databases. eg. gnomAD, Ensembl, AlphaFold, the literature graph, and any hypotheses that [FAIRdata.ai](http://fairdata.ai/?ref=openresearch.wtf) sends it and tries to identify patterns that haven't yet been written up as research articles. It then generates hypotheses, assembles supporting evidence, and drafts those hypotheses into structured manuscripts. It's not trying to replace the experimental scientist; it's trying to explore the space of already-existing open data for combinations that nobody has looked at yet. You can read more details about [the technical capabilities and dependencies of OpenScience.ai here.](https://www.openresearch.wtf/openscience-ai-using-open-data-to-generate-research-that-doesnt-yet-exist/) [Preprints.ai](http://preprints.ai/?ref=openresearch.wtf) asks a fairly direct question - How much of peer review can we automate? It runs a multi-agent review pipeline assessing both research integrity, methodology, statistical validity, reproducibility, and novelty. The output is a structured assessment graded against a rubric, with deliberative rounds between agents before a consensus is reached. You can read more details about [the technical capabilities and dependencies of Preprints.ai here.](https://www.openresearch.wtf/preprints-ai-how-much-of-peer-review-can-we-automate/) [OpenAccess.ai](http://openaccess.ai/?ref=openresearch.wtf) closes the loop. If we're generating articles via OpenScience.ai and reviewing them via Preprints.ai, we need somewhere to publish them. OpenAccess.ai is that place - an open access journal built on the assumption that the editorial process can run largely on agents, and that the unit cost of academic publishing can be meaningfully reduced as a result. You can read more details about [the technical capabilities and dependencies of OpenAccess.ai here.](https://www.openresearch.wtf/openaccess-ai-how-cheap-can-we-make-academic-publishing/) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-55f13fad-b68e-46ab-b618-6c0c1d74d681.png) All four platforms are built exclusively on open data and open APIs. If you can build a plausible end-to-end research automation pipeline using only what's publicly available, that says something useful about the current state of open science infrastructure. The constraint on scientific productivity isn't data availability; it's the capacity for machines to process and synthesise it. I personally have access to considerably richer data through my work at Digital Science (citation graphs, full-text repositories, institutional metadata). I haven't used any of it here. What you're seeing is what's possible with no proprietary advantage, which I think makes the findings more generalisable. **Humans in the loop** Framing this as "minimal human intervention" is accurate for certain parts of the pipeline and misleading for others. There are stages where the automation is fairly robust. Metadata enrichment, FAIR scoring, integrity screening, novelty checks against the existing literature. These are tasks where the output is verifiable and the failure modes are containable. They're also, not coincidentally, the tasks that consume a lot of researcher time without requiring much researcher judgement. But there are other stages where a human needs to be in the loop, and probably will for a while. Editorial decisions - whether a piece of work is actually worth publishing, whether the framing is intellectually honest, whether the conclusions follow from the evidence. These are not tasks I'd be comfortable delegating to agents at this point. The current OpenAccess.ai setup allows human editors to review agent assessments before making final decisions, and I think that hybrid model is the right one for this experiment. It also leans very heavily on [eLife’s Publish, Review, Curate model](https://elifesciences.org/inside-elife/dc24a9cd/open-science-what-is-publish-review-curate?ref=openresearch.wtf), which is a model I see gaining more traction and uptake in the coming years. Similarly, the hypothesis generation in OpenScience.ai needs a human sanity check before anything gets treated as a serious research direction. The system can identify patterns in open data, but it can't yet reliably distinguish a genuinely interesting finding from a statistical artefact or a confound that any domain expert would spot immediately. I've had to be quite disciplined about this after discovering early on that the pipeline would confidently produce fabricated statistics if you didn't anchor every claim to verifiable data from a canonical source. **What's next** Each platform has a reasonably long list of things it still can't do well. OpenScience.ai's anti-fabrication infrastructure is solid now but the discovery logic is still fairly shallow. Preprints.ai's local inference falls over on structured output for complex multi-agent workflows. FAIRdata.ai's scoring is imprecise for accession-based repositories. OpenAccess.ai's cost economics need testing at meaningful volume before any of the unit cost claims are worth much. I'm writing this up now in order to experiment in the open. The idea of chaining these capabilities together is the more interesting thing to put into the world, even if the implementations are still rough. If you're working on any of these problems and want to compare notes, or if you spot something obviously wrong with the approach, I'd be glad to hear from you. ### OpenAccess.ai: How Cheap Can We Make Academic Publishing? URL: https://www.openresearch.wtf/openaccess-ai-how-cheap-can-we-make-academic-publishing/ Last updated: 2026-07-31T12:35:48.000Z ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-90d1cdf6-9b9d-403e-b470-f543da9eaf9b.png) ## **What Problem Are We Trying to Solve?** Article Processing Charges at major journals run to thousands of dollars per paper. The editorial infrastructure - submission systems, copyediting, typesetting, hosting - is expensive. What I want to know is how much of the process can be automated, particularly for machine generated academic content. If you can automate a substantial fraction of the editorial process, (as [Preprints.ai](https://preprints.ai/?ref=openresearch.wtf) is attempting to do) what happens to the unit economics? [OpenAccess.ai](https://openaccess.ai/?ref=openresearch.wtf) is an experiment designed to provide some data. There's a second motivation. [OpenScience.ai](https://openscience.ai/?ref=openresearch.wtf) (in theory) generates research articles. Those articles need somewhere to go. Conventional open access infrastructure is designed for human-authored papers, with submission systems and APC structures that don't map cleanly onto a machine-generated pipeline. OpenAccess.ai is built to accept submissions via API, which means it can serve as the publishing endpoint for the entire experimental pipeline I’m building, without friction. ## **The Technical Architecture** ### **Core Stack** OpenAccess.ai is a [Next.js 14](https://nextjs.org/?ref=openresearch.wtf) application deployed on [Netlify](https://www.netlify.com/?ref=openresearch.wtf) with [Supabase](https://supabase.com/?ref=openresearch.wtf) for PostgreSQL, auth, and file storage, and [Railway](https://railway.app/?ref=openresearch.wtf) for microservices. Serverless-first, no dedicated servers, minimal ops overhead. If we're testing whether publishing can be radically cheap, the infrastructure costs need to be radically cheap too. | Layer | Technology | Purpose | | -------------------- | --------------------------------------------------------- | ------------------------------------------------- | | Frontend | Next.js 14, React 18, Tailwind CSS | Article reader, submission wizard, dashboard | | Database | Supabase (PostgreSQL + RLS) | Submissions, reviews, user profiles, API accounts | | File Storage | Supabase Storage | Manuscripts, figures, cover images | | Serverless Functions | Netlify Functions (26s limit) | Upload processing, metadata extraction, checkout | | Long-running Tasks | Supabase Edge Functions (150s limit) | PDF text extraction via Claude Vision | | Microservices | Railway (Docker) | GROBID, figure extraction | | Payments | [Stripe](https://stripe.com/?ref=openresearch.wtf) | Low submission fee, API credit purchases | | Email | [Postmark](https://postmarkapp.com/?ref=openresearch.wtf) | Submission confirmations, status notifications | ### **The Document Processing Pipeline** When someone uploads a PDF, the platform needs to extract the full text, identify sections, parse references, pull out figures, and generate structured metadata. Stage 1: Fast extraction The upload route runs [pdf-parse](https://www.npmjs.com/package/pdf-parse?ref=openresearch.wtf) for basic text extraction in parallel with a Supabase Storage upload. For DOCX files, [mammoth.js](https://www.npmjs.com/package/mammoth?ref=openresearch.wtf) handles the conversion. This produces something immediately. However, for PDFs with embedded fonts (which is most academic PDFs), pdf-parse returns garbled text. Stage 2: AI vision extraction When pdf-parse fails (detected by our garbled-text heuristic that checks for XMP metadata patterns, font encoding artifacts, and low information density), the client triggers a [Supabase Edge Function](https://supabase.com/docs/guides/functions?ref=openresearch.wtf) that sends the PDF to [Claude's](https://www.anthropic.com/claude?ref=openresearch.wtf) document vision API. Claude Haiku reads the visual layout of each page and produces clean Markdown with proper sections, author lists, figure captions, and references. Stage 3: Structured metadata extraction The extracted text is sent to Claude Haiku again, this time with a structured prompt that returns JSON with title, abstract, authors with affiliations, keywords, and subject area (mapped to [bioRxiv's taxonomy](https://www.biorxiv.org/alertsrss?ref=openresearch.wtf) of 25 subject areas). This populates the submission form automatically. Stage 4: Figure extraction A [PyMuPDF](https://pymupdf.readthedocs.io/?ref=openresearch.wtf)\-based microservice on Railway extracts embedded images from the PDF, filters by minimum dimensions, and stores them in Supabase Storage. The figure URLs are matched to the document AST by order. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-dfed17c9-c9b8-4f24-9bc8-326ac221a164.png) ### **The AST Document Model** All documents are converted to an internal Abstract Syntax Tree inspired by [MyST Markdown](https://mystmd.org/?ref=openresearch.wtf) (itself derived from [mdast](https://github.com/syntax-tree/mdast?ref=openresearch.wtf) and [Stencila](https://stenci.la/?ref=openresearch.wtf)). The AST is the canonical representation - HTML, JATS XML, PDF, and all citation formats are rendered from it on demand. RootNode ├── meta: { title, abstract, sections\[\], wordCount, figureCount, referenceCount } ├── children: SectionNode\[\] │ ├── SectionNode (id: "introduction") │ │ ├── HeadingNode │ │ ├── ParagraphNode\[\] │ │ └── FigureNode { url, caption, label } │ ├── SectionNode (id: "methods") │ └── SectionNode (id: "results") └── references: ReferenceNode\[\] (CSL-JSON format) Sections are first-class nodes, not just CSS classes on headings. This means you can programmatically extract "the methods section" or "all figure captions" without parsing HTML — which is essential for the AI review pipeline. ## **Peer Review via Preprints.ai** Every submission is automatically sent to [Preprints.ai](https://preprints.ai/?ref=openresearch.wtf) for assessment. The review system uses 8 AI reviewers that evaluate the paper across multiple dimensions: - Integrity scoring (A–E grade scale) - Novelty assessment (1–5 tiers: minimal to groundbreaking) - Methodology checks (statistical validity, sample sizes, controls) - Transparency audit (ethics statement, data/code availability, funding disclosure, COI, pre-registration) - Provenance analysis (for AI-generated papers: model plausibility, prompt injection detection, reproducibility signals) The grade, strengths, concerns (with severity: critical/major/minor), and per-reviewer scores are stored as structured JSON and displayed on the article page. The review is fully transparent — readers see exactly what the AI reviewers found, not just a binary accept/reject. ## **Output Formats and Academic Standards** A publishing platform is only useful if it produces outputs that the academic ecosystem can consume. OpenAccess.ai generates: ### **For Indexing and Discovery** - [JATS XML](https://jats.nlm.nih.gov/?ref=openresearch.wtf) \- Journal Article Tag Suite, the standard for CrossRef DOI deposit, PubMed indexing, and PMC archiving. Generated on-demand from the AST via astToJats(). - [Dublin Core](https://www.dublincore.org/?ref=openresearch.wtf) meta tags - for library catalog discovery (Zotero, Mendeley) - [Highwire Press](https://scholar.google.com/intl/en/scholar/inclusion.html?ref=openresearch.wtf#indexing) tags - citation\_title, citation\_author, etc. for Google Scholar - [Schema.org JSON-LD](https://schema.org/ScholarlyArticle?ref=openresearch.wtf) \- ScholarlyArticle type with full provenance metadata ### **For Citation** - Citation styles: APA 7th, Vancouver, Chicago, Harvard, MLA 9th, BibTeX, RIS - generated client-side from article metadata using [Citation Style Language](https://citationstyles.org/?ref=openresearch.wtf) conventions - [CFF](https://citation-file-format.github.io/?ref=openresearch.wtf) (Citation File Format) - for software-like citations ### **For Machine Learning and Linked Data** - [Croissant ML](https://mlcommons.org/croissant/?ref=openresearch.wtf) \- ML dataset metadata standard, making the corpus discoverable for training pipelines - [RO-Crate](https://www.researchobject.org/ro-crate/?ref=openresearch.wtf) \- Research Object packaging with full provenance - [W3C DCAT](https://www.w3.org/TR/vocab-dcat-3/?ref=openresearch.wtf) \- Data Catalog Vocabulary for dataset registries - [DataCite](https://datacite.org/?ref=openresearch.wtf) JSON - DOI registration metadata ### **Persistent Identifiers** - [Handle System](https://www.handle.net/?ref=openresearch.wtf) \- Each article gets a globally unique persistent identifier (e.g., 20.500.14906/oai.biology.2026.00001) registered via Handle.net's REST API ## **The Article Reader** The article page is modelled on [eLife](https://elifesciences.org/?ref=openresearch.wtf)'s reading experience - Tabbed interface. Full text, Figures & data, Peer review (with full Preprints.ai assessment), Summary - Sticky section navigation. Left sidebar with table of contents generated from AST section nodes, highlights current section on scroll - Inline citations. Hover to see reference details, click to scroll to bibliography - AI-powered Q&A. Claude-powered chat that answers questions about the article content - Branded PDF export: Print-optimized HTML with OpenAccess.ai header, proper title block, A4 pagination - Article metrics. Views, downloads, citations tracked per article ### **Authentication** Users sign in via [ORCID](https://orcid.org/?ref=openresearch.wtf) OAuth or [GitHub](https://github.com/?ref=openresearch.wtf) OAuth. ORCID profiles are enriched with works count and affiliation data. The affiliation field uses the [ROR API](https://ror.org/?ref=openresearch.wtf) (Research Organization Registry) for standardised institution lookup. ## **Agent-First Architecture** OpenAccess.ai exposes an [MCP](https://modelcontextprotocol.io/?ref=openresearch.wtf) (Model Context Protocol) endpoint at /api/mcp with tools for: - submit\_paper - Programmatic manuscript submission (API key + credits) - check\_submission - Query review status - get\_article - Fetch in any format (JSON, JSON-LD, BibTeX, Croissant, etc.) - search\_articles - Full-text search - get\_subject\_areas - Valid taxonomy lookup Plus a REST API documented at [/api/openapi.json](https://openaccess.ai/api/openapi.json?ref=openresearch.wtf) and discoverable via [/llms.txt](https://openaccess.ai/llms.txt?ref=openresearch.wtf), [.well-known/ai-plugin.json](https://openaccess.ai/.well-known/ai-plugin.json?ref=openresearch.wtf), and an [agent card](https://openaccess.ai/.well-known/agent-card?ref=openresearch.wtf). The API uses credit-based billing: generate an API key, purchase credits via Stripe, spend 1 credit per submission. This is what enables [OpenScience.ai](https://openscience.ai/?ref=openresearch.wtf) to generate a paper and submit it to OpenAccess.ai without any human in the loop. ## **Tools That Inspired the Platform** Several open-source projects shaped the architecture: | Project | Influence | | ------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------- | | [eLife Enhanced Preprints](https://github.com/elifesciences?ref=openresearch.wtf) | Article reader UX, tabbed layout, assessment display, JATS to HTML to PDF | | [GROBID](https://github.com/kermitt2/grobid?ref=openresearch.wtf) | PDF to TEI XML extraction pipeline | | [MyST Markdown](https://mystmd.org/?ref=openresearch.wtf) / [mdast](https://github.com/syntax-tree/mdast?ref=openresearch.wtf) | AST document model | | [Octopus.ac](https://www.octopus.ac/?ref=openresearch.wtf) | Modular publication types (Problem to Method to Results chain) | | [bioRxiv](https://www.biorxiv.org/?ref=openresearch.wtf) | Subject area taxonomy | | [Citation Style Language](https://citationstyles.org/?ref=openresearch.wtf) | Multi-format citation generation | | [Stencila](https://stenci.la/?ref=openresearch.wtf) | Executable document concepts | | [AnyStyle](https://anystyle.io/?ref=openresearch.wtf) | Reference string parsing to CSL-JSON | ## **What We Can Look to Do Next** We need to run a meaningful volume of articles through the full pipeline - generate via [OpenScience.ai](http://openscience.ai/?ref=openresearch.wtf), review via [Preprints.ai](http://preprints.ai/?ref=openresearch.wtf), publish via [OpenAccess.ai](http://openaccess.ai/?ref=openresearch.wtf) \- and measure the per-article cost at each stage. Figure extraction needs improvement. The PyMuPDF service extracts embedded images but doesn't yet match them to figure captions reliably. We're evaluating eLife's [enhanced-preprints-image-server](https://github.com/elifesciences/enhanced-preprints-image-server?ref=openresearch.wtf) (IIIF-based) for production-grade image delivery. Reference enrichment is the next major quality improvement. Replacing AnyStyle with [jats-ref-refinery](https://github.com/elifesciences/jats-ref-refinery?ref=openresearch.wtf) would dramatically improve citation linking by querying four external databases rather than relying on a single CRF parser. Indexing and registration remain open questions. A journal that publishes AI-generated research will not easily get into PubMed or [Dimensions.ai](http://dimensions.ai/?ref=openresearch.wtf). But understanding what metadata, what provenance signals, and what quality documentation would be required for AI-generated research to be treated as a first-class research output is itself a useful question for the field to be working on. The most interesting longer-term question is whether a hybrid editorial model makes sense - AI review as the default pathway, with human editorial escalation for papers that score ambiguously or that make claims in domains where agent reviewers have known weaknesses.If you're working on any of these problems and want to compare notes, or if you spot something obviously wrong with the approach, I'd be glad to hear from you. ### Preprints.ai: How Much of Peer Review Can We Automate? URL: https://www.openresearch.wtf/preprints-ai-how-much-of-peer-review-can-we-automate/ Last updated: 2026-07-31T12:35:53.000Z ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-ccdfdf7f-1495-48bd-bdd7-149c2f5645ed.png) ## **What problem are we trying to solve?** [Peer review is structurally failing](https://www.openresearch.wtf/have-we-already-hit-the-peer-review-breaking-point/). The volume of preprint submissions has grown faster than the reviewer pool for decades. Average review turnaround times have lengthened. Reviewer fatigue is well-documented. And review quality is inconsistent in ways that are impossible to fix when the process depends entirely on the goodwill of overcommitted academics. More simply, can we score every preprint. At the same time, a significant fraction of what peer reviewers actually do is mechanical. Checking that statistical methods match reported results. Verifying that cited claims are accurately represented. Assessing whether the methodology section contains enough detail to reproduce the experiments. Scanning for common red flags - tortured phrases from paper mills, image duplication, impossible p-values. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-6376e492-4ea3-46a5-85fd-d8bfd4d0b659.png) Preprints.ai asks a sharper question than "can AI help with peer review?" It asks “if we were designing peer review from scratch for a world where powerful LLMs exist, what would we actually need humans for, and what could we comfortably automate?” The platform currently runs a multi-agent review pipeline against preprints from [bioRxiv](https://www.biorxiv.org/?ref=openresearch.wtf) and [medRxiv](https://www.medrxiv.org/?ref=openresearch.wtf), assessing two dimensions: *research integrity* (methodology, statistical validity, reproducibility, citation accuracy) and *novelty* (does the core claim already exist in the literature?). The output is a structured assessment graded on an A5-to-E1 rubric, with positivity bias actively recalibrated using publication outcome data. ## **What tools have inspired or been used to build this platform?** ### **LLM providers and hybrid routing** We use three LLM providers in production, with each agent in the review panel assigned to a *different* provider to ensure genuine inter-model independence: | Provider | Model | Role | Cost / 1M tokens | | ------------------------------------------------------------------------------------------- | ------------------------- | --------------------------------------- | -------------------- | | [Claude](https://docs.anthropic.com/en/docs/about-claude/models?ref=openresearch.wtf) | claude-sonnet-4-20250514 | Primary reviewers (methodology, domain) | $3 in / $15 out | | [Claude Haiku](https://docs.anthropic.com/en/docs/about-claude/models?ref=openresearch.wtf) | claude-haiku-4-5-20251001 | Fast checks (ethics, transparency) | $0.80 in / $4 out | | [GPT-4o](https://platform.openai.com/docs/models?ref=openresearch.wtf) | gpt-4o-2024-11-20 | Secondary reviewer for consensus | $2.50 in / $10 out | | [Gemini](https://ai.google.dev/gemini-api/docs/models?ref=openresearch.wtf) | gemini-2.5-flash | Secondary reviewer for consensus | $0.15 in / $0.60 out | This isn't just cost optimisation. When three different model families independently agree on a finding, the probability that all three share the same hallucination is substantially lower than any single model's error rate. Round-robin provider assignment prevents correlated model bias in the consensus score. ### **The review pipeline: Layer 1 (deterministic, zero LLM cost)** Before any LLM sees the paper, a battery of rule-based checks runs in under 2 seconds: PDF Quality Gate → Language Detection → Paper Mill Detection → Image Forensics → Statistical Consistency → Trust Markers These are inspired by and build on established tools - such as Ripeta’s Trust Indicators. - Paper mill detection uses a database of 8,000+ [tortured phrases](https://pubpeer.com/?ref=openresearch.wtf) (e.g., "bosom malignancy" for "breast cancer") and 257 [SCIgen](https://pdos.csail.mit.edu/archive/scigen/?ref=openresearch.wtf) patterns for computer-generated text - Statistical consistency implements the [GRIM test](https://jamesheathers.github.io/sprite/index.html?ref=openresearch.wtf) (Granularity-Related Inconsistency of Means) and [statcheck](http://statcheck.io/?ref=openresearch.wtf)\-style p-value recalculation from reported test statistics - Image forensics is inspired by [ELIS](https://www.biorxiv.org/content/10.1101/2022.01.24.477532v1?ref=openresearch.wtf) (Error Level Image Similarity) — detecting duplicate figure regions, inconsistent compression, and manipulation artefacts using [imagehash](https://github.com/JohannesBuchner/imagehash?ref=openresearch.wtf) and [Pillow](https://pillow.readthedocs.io/?ref=openresearch.wtf) - Fabrication detection checks for [Benford's Law](https://en.wikipedia.org/wiki/Benford%27s%5Flaw?ref=openresearch.wtf) deviation in reported data and suspiciously perfect statistical distributions - Trust markers detect the presence (or absence) of data availability statements, code repositories, ethics approvals with protocol numbers, pre-registration links ([ClinicalTrials.gov](https://clinicaltrials.gov/?ref=openresearch.wtf), [OSF](https://osf.io/?ref=openresearch.wtf)), ORCID identifiers, and CRediT author contributions - Reporting guideline compliance checks against [ARRIVE](https://arriveguidelines.org/?ref=openresearch.wtf) (animal research), [CONSORT](https://www.consort-statement.org/?ref=openresearch.wtf) (clinical trials), [PRISMA](https://www.prisma-statement.org/?ref=openresearch.wtf) (systematic reviews), and [MIQE](https://www.gene-quantification.de/miqe-bustin-et-al-clin-chem-2009.pdf?ref=openresearch.wtf) (qPCR) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-63fa5891-49cd-45e5-a7fa-065fa7be545e.png) ### **The review pipeline: Layer 2 (multi-agent LLM review)** The agent architecture is heavily inspired by [eLife's deliberative review model](https://elifesciences.org/inside-elife/e01a1e6e/elife-s-new-model-changing-the-way-you-share-your-research?ref=openresearch.wtf). We run nine specialist agents, each with a distinct role and scope boundary: | Agent | Focus | Scope boundary | | --------------------- | ------------------------------------------------------------ | ----------------------------------- | | Methodologist | Experimental design, controls, sample size, confounding | NOT statistics (statistician's job) | | Statistician | Test appropriateness, p-values, effect sizes, power analysis | NOT study design | | Domain Expert | Novelty assessment, field context, missing citations | Full eLife Reviewer #1 depth | | Reproducibility | Method detail, data/code availability, reporting guidelines | NOT statistical correctness | | Ethics & Transparency | IRB approvals, COI disclosure, funding, pre-registration | Ethics only, not science | | Scientific Validity | Pseudoscience gate: theoretical coherence, plausibility | Gate only — not quality | Plus three additional domain-specialist agents routed based on the paper's bioRxiv category, drawn from a pool of 31 field-specific configurations covering everything from biochemistry to systems biology. Each configuration specifies methodology standards, common issues, reporting guidelines, statistical expectations, and key databases for that field. ### **Deliberation: from independent reviews to consensus** Independent review is necessary but not sufficient. In human peer review, the reviewing editor synthesises reviewer comments into a coherent assessment. We replicate this with a structured deliberation protocol inspired by eLife: 1. Round 0 (Independent review) - All 9 agents review the paper in parallel, each with a different LLM provider. No agent sees any other agent's output. 2. Round 1 (Consultation) - Agents see anonymised summaries of each other's reviews. Each can agree, challenge, or add nuance. 3. Round 2 (Reconciliation) - A senior editor agent synthesises the consultation into a single eLife-style assessment: *"This study presents \[significance\] findings on \[topic\]. The evidence is \[strength\], with \[justification\]."* 4. Dissent recording - Any agent that disagrees with the final assessment can lodge a recorded dissent, preserved in the output for transparency. This substantially reduces variance compared to single-pass review. We track inter-agent agreement (standard deviation of scores across agents) and use it as a confidence signal. eg. high disagreement triggers lower confidence scores and flags the paper as borderline. ### **External data sources** The platform integrates with multiple academic data APIs, each serving a distinct purpose: | Source | Purpose | Usage | | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | --------------------------------------------------------------------- | | [Semantic Scholar](https://api.semanticscholar.org/?ref=openresearch.wtf) | Citation network, reference verification, passage search | 200M+ papers; checks cited claims are accurately represented | | [OpenAlex](https://docs.openalex.org/?ref=openresearch.wtf) | Journal tier mapping, h-index, publication outcomes | Calibration: maps journal prestige to validate our novelty scores | | [Crossref](https://api.crossref.org/?ref=openresearch.wtf) | DOI metadata, retraction status, publication relations | Detects retracted papers and tracks preprint-to-publication links | | [Europe PMC](https://europepmc.org/RestfulWebService?ref=openresearch.wtf) | Publication discovery, fulltext links | Fallback for finding published versions of preprints | | [bioRxiv API](https://api.biorxiv.org/?ref=openresearch.wtf) | Preprint metadata ingestion | Primary source for paper metadata and PDF URLs | | [Retraction Watch](https://retractionwatch.com/retraction-watch-database-user-guide/?ref=openresearch.wtf) | Known-retracted papers for validation | Ground truth: can our system detect papers that were later retracted? | ### ### **The scoring rubric** We grade on two independent dimensions, producing a combined grade like B4 or C3: Integrity (A–E): Strength of evidence. A = compelling, reproducible methodology with full transparency. E = fundamental unfixable problems (fabrication signals, pseudoscience, paper mill indicators). Novelty (5–1): Significance of findings. 5 = landmark, field-shifting (almost never assigned). 3 = important, real contribution beyond a single subfield. 1 = useful but very limited — one more data point. Each grade maps to an editorial decision in journal-standard language: Accept, Minor Revision, Major Revision, or Reject — with modifiers for low confidence, abstract-only assessment, or high reviewer disagreement. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-1ace464e-673d-4aa4-a452-c0d5e520763d.png) ### **Fighting positivity bias using recalibration** LLMs are famously sycophantic. Left uncalibrated, every paper gets a B. We fight this at multiple levels: - Downward recalibration rules: 2+ critical weaknesses forces a D floor. 4+ major weaknesses caps at low-C. A "reject" recommendation from any agent caps integrity at 0.42. - Per-agent accuracy tracking: After each calibration run, we compute every agent's mean absolute error and bias against publication outcomes. Agents with measured bias > ±0.15 get a steering note injected into their next prompt: *"Historical data shows you tend to overscore novelty by \~20%. Be especially critical of incremental contributions."* - Inverse MAE weighting: In the consensus aggregation, agents with lower historical error get higher weight. An agent that has been consistently accurate on past papers has more influence on the final score than one that hasn't. - Publication outcome calibration: We match assessed preprints against publication records [using Crossref’s open data](https://www.crossref.org/blog/discovering-relationships-between-preprints-and-journal-articles/?ref=openresearch.wtf). If we scored a paper as novelty-2 but it was published in a tier-5 journal, that delta feeds back into agent weights and prompt tuning. ### **Confidence and transparency** Every assessment includes: - Confidence score (0–100%) derived from agent agreement, mean agent self-reported confidence, and score dispersion. High disagreement between agents lowers confidence. - Individual agent reviews. Every agent's scores, strengths, weaknesses, and recommendations are available in the full report, not just the consensus. - A 2–3 sentence plain-language summary of the core finding, for a non-specialist audience. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-f5970253-ca5a-422a-b951-dc680bacd50f.png) ### **Cost per paper** With hybrid routing across Claude, GPT-4o, and Gemini Flash, the current cost per full assessment is approximately $1.50–2.00 for a consensus review with 9 agents. Layer 1 deterministic checks add negligible cost. At scale, with aggressive caching of reference verification results and local model routing for lower-signal checks, we expect this to converge toward $0.50–1.00 per paper. ## **What we can look to do next** ### **Validation is the open question** The most important next step is systematic validation. How do Preprints.ai assessments compare to human peer review on the same papers? We've built a [/validate](https://preprints.ai/validate?ref=openresearch.wtf) page where domain experts can agree, disagree, or override AI grades but we need hundreds of expert validations across diverse fields before the calibration claims are defensible. There is no incentive for subject experts to do this, so we can assume they wont. We're also running consistency tests. We are assessing the same papers multiple times to measure grade variance. If the system can't reproduce its own grades, nobody should trust them. Our target is >70% exact grade stability and >90% within-one accuracy. ### **Field-specific scoring norms** A C3 in clinical medicine means something different than a C3 in pure mathematics. We're computing per-discipline score distributions from assessed papers and using z-score normalisation to generate field-specific grade thresholds. Fields with sufficient data (20+ assessed papers) get their own norms; others fall back to global defaults. ### **The hybrid model question** Longer term, the interesting question is whether a hybrid model - AI as first-pass screen, human expert for borderline cases - could work operationally within a real journal workflow. The system already identifies which papers need human attention (low confidence, high agent disagreement, borderline C-range grades). The question is whether editors would trust it enough to act on. I think the answer is yes, but only after the validation data proves it.If you're working on any of these problems and want to compare notes, or if you spot something obviously wrong with the approach, I'd be glad to hear from you. ### OpenScience.ai: Using Open Data to Generate Research That Doesn't Yet Exist URL: https://www.openresearch.wtf/openscience-ai-using-open-data-to-generate-research-that-doesnt-yet-exist/ Last updated: 2026-07-31T12:35:52.000Z ## **What problem are we trying to solve?** OpenScience.ai is trying to look forwards. [OpenScience.ai](http://openscience.ai/?ref=openresearch.wtf) aims to be a gap filler in academic data and literature. Given the current state of open scientific databases, what patterns are present in the data that haven't yet been written up as research claims? What combinations of findings, if synthesised across [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/?ref=openresearch.wtf), [GTEx](https://gtexportal.org/?ref=openresearch.wtf), [STRING](https://string-db.org/?ref=openresearch.wtf), and [Open Targets](https://platform.opentargets.org/?ref=openresearch.wtf), would constitute a novel contribution? The goal isn't to produce papers that get cited. It's a thought experiment and a way to explore the space of what's discoverable in open data when you remove the bandwidth constraint of human researchers having to do the querying manually. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-33e14ec3-49be-4724-9f6d-83558aa70c66.png) ## **What tools have inspired or been used to build this platform?** ### **The Data Lake: 40+ Open Scientific Databases** The discovery pipeline operates on a local data lake of 44 bulk tables, populated from established scientific databases. Agents query local data first via a typed bulkFirst library before falling back to live APIs. The sources span: Variant and Population Genetics - [gnomAD v4](https://gnomad.broadinstitute.org/?ref=openresearch.wtf) (GraphQL) — population allele frequencies across AFR, AMR, ASJ, EAS, FIN, NFE, SAS populations, plus gene constraint scores (pLI, LOEUF, mis\_z, syn\_z) from the [v4.1 constraint metrics](https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1/constraint/?ref=openresearch.wtf) - [ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/?ref=openresearch.wtf) (NCBI Entrez API) — clinical variant interpretations with pathogenicity classifications - [GWAS Catalog](https://www.ebi.ac.uk/gwas/?ref=openresearch.wtf) — genome-wide association study results - [MyVariant.info](https://myvariant.info/?ref=openresearch.wtf) — variant resolution and annotation aggregation - [dbSNP](https://www.ncbi.nlm.nih.gov/snp/?ref=openresearch.wtf) (NCBI eUtils) — variant identifiers Gene and Protein Data - [Ensembl](https://rest.ensembl.org/?ref=openresearch.wtf) — variant effect prediction (VEP), gene annotation, regulatory feature overlap - [UniProt](https://www.uniprot.org/?ref=openresearch.wtf) — protein sequences, functional annotations, subcellular localisation - [AlphaFold DB](https://alphafold.ebi.ac.uk/?ref=openresearch.wtf) — predicted 3D protein structures, integrated at three levels: inline enrichment during discovery, bulk ingestion of coverage data, and three dedicated database tables for structure mapping - [RCSB PDB](https://www.rcsb.org/?ref=openresearch.wtf) — experimentally determined protein structures (215K entries) - [InterPro](https://www.ebi.ac.uk/interpro/?ref=openresearch.wtf) — protein family and domain classification - [HGNC](https://www.genenames.org/?ref=openresearch.wtf) — standardised gene nomenclature Pharmacogenomics and Drug Data - [PharmGKB](https://www.pharmgkb.org/?ref=openresearch.wtf) — pharmacogenomic clinical annotations and dosing guidelines - [ChEMBL](https://www.ebi.ac.uk/chembl/?ref=openresearch.wtf) — bioactive compound data (5M+ bioactivity records) - [DGIdb](https://dgidb.org/?ref=openresearch.wtf) — drug-gene interaction database - [DDInter](http://ddinter.scbdd.com/?ref=openresearch.wtf) — drug-drug interactions - [PharmVar](https://www.pharmvar.org/?ref=openresearch.wtf) — star allele nomenclature for pharmacogenes - [OpenFDA](https://open.fda.gov/?ref=openresearch.wtf) — adverse drug reaction reports Pathway and Network Analysis - [STRING](https://string-db.org/?ref=openresearch.wtf) — protein-protein interaction networks with confidence scoring - [Reactome](https://reactome.org/?ref=openresearch.wtf) — curated biological pathways - [MSigDB](https://www.gsea-msigdb.org/gsea/msigdb/?ref=openresearch.wtf) — gene set collections for pathway enrichment - [Gene Ontology](http://geneontology.org/?ref=openresearch.wtf) — functional annotations via GOA Disease and Phenotype - [Open Targets Platform](https://platform.opentargets.org/?ref=openresearch.wtf) (GraphQL) — target-disease-drug associations, evidence, and locus-to-gene predictions - [DisGeNET](https://www.disgenet.org/?ref=openresearch.wtf) — gene-disease associations with evidence scoring - [HPO](https://hpo.jax.org/?ref=openresearch.wtf) (Human Phenotype Ontology) — standardised phenotype terminology - [Mondo](https://mondo.monarchinitiative.org/?ref=openresearch.wtf) (via [Monarch Initiative](https://monarchinitiative.org/?ref=openresearch.wtf)) — disease ontology - [OMIM](https://omim.org/?ref=openresearch.wtf) — Mendelian disease catalogue Expression and Regulation - [GTEx](https://gtexportal.org/?ref=openresearch.wtf) (v8 bulk download) — tissue-specific gene expression across 54 tissues - [ENCODE cCRE](https://screen.encodeproject.org/?ref=openresearch.wtf) — candidate cis-regulatory elements (32M elements genome-wide) - [CellxGene](https://cellxgene.cziscience.com/?ref=openresearch.wtf) — single-cell expression atlas Literature and Citation - [Semantic Scholar](https://www.semanticscholar.org/?ref=openresearch.wtf) — paper search, citation graphs, and the ASTA snippet search endpoint (275M passages across 12M+ papers) for contradiction detection - [OpenAlex](https://openalex.org/?ref=openresearch.wtf) — open publication metadata, citation counts, author disambiguation - [Europe PMC](https://europepmc.org/?ref=openresearch.wtf) — open access full-text and text-mined annotations - [NIH Reporter](https://reporter.nih.gov/?ref=openresearch.wtf) — funded research project data Clinical - [ClinicalTrials.gov](https://clinicaltrials.gov/?ref=openresearch.wtf) — active clinical trial registrations ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-7324428e-5ebb-49eb-9f19-1226be3530b8.png) ### **The Agent Fleet: Specialised Research Personas** The discovery pipeline doesn't use a single general-purpose LLM prompt. It runs through a fleet of 89 persistent AI agents, each with a distinct scientific identity defined in a SOUL file. Each SOUL specifies the agent's domain expertise, methodological perspective, vocabulary constraints, preferred data sources, and anti-fabrication rules. The genomics lab alone has 25 sub-agents covering specific niches: a population-analyst that studies variant distribution across ancestry groups using FST and admixture analysis; a rare-disease-hunter that identifies founder variants in underrepresented populations; a pharmacogenomics-translator that maps drug-gene interactions to dosing implications; a noncoding-specialist that interprets regulatory variants against ENCODE cCRE data; a cancer-predisposition agent; an ancient DNA specialist; and others. Agent selection is reputation-weighted. Agents with higher success rates (discoveries that survive validation) get routed more work, but the system deliberately gives lower-reputation agents opportunities to prevent premature convergence. This architecture was inspired by multi-agent adversarial review patterns described in recent work on [LLM-based scientific reasoning](https://arxiv.org/abs/2304.05332?ref=openresearch.wtf), and by the [Octopus model of scientific publishing](https://www.octopus.ac/?ref=openresearch.wtf) which breaks monolithic papers into linked, independently citable components. ### **Statistical Computation** Our compute engine implements the actual statistical methods from scratch: - Weir-Cockerham FST with pairwise population comparisons, using the [Wilson-Hilferty transformation](https://en.wikipedia.org/wiki/Wilson%E2%80%93Hilferty%5Ftransformation?ref=openresearch.wtf) for log-space chi-squared CDF computation. This was necessary because gnomAD's massive sample sizes (allele counts >500K) produce extreme p-values that cause numerical underflow with naive implementations. The Wilson-Hilferty normal approximation to the chi-squared distribution handles this gracefully. - Odds ratios with Woolf logit method and continuity correction - Hypergeometric test for pathway enrichment with fold-change calculation - Benjamini-Hochberg FDR correction and Bonferroni correction for multiple testing - Bootstrap confidence intervals (95% CI via normal approximation with 1.96 \* SE bounds) - Statistical power analysis specific to FST, odds ratios, and allele frequency differences - Fixed-effects and random-effects meta-analysis with Q-statistic, I-squared heterogeneity, and forest plot generation via [Vega-Lite](https://vega.github.io/vega-lite/?ref=openresearch.wtf) specs A separate hypothesis-executor module goes further. It generates Python analysis code, executes it in a sandboxed subprocess against real bulk data, and returns computed statistics. The LLM writes the analysis code. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-bc069f32-4539-4f77-86d1-868cd8af3276.png) ### **The Evidence Firewall** The single most important architectural decision in the platform is what I call "the firewall" - a module where every API response passes through deterministic, per-source parsers that extract structured evidence rows, without LLM involvement. Each row records the source API, source method, source version, evidence type, standardised entity identifiers, and explicit null-result flags for non-findings. These evidence rows are stored in a dedicated table linked to each discovery via foreign keys. Everything downstream (hypothesis formation, dataset generation, manuscript writing) must reference these rows. The numerical whitelist gate in the discovery cycle enforces this: the LLM receives an explicit list of numbers that appeared in API responses, and the prompt states "You MAY ONLY cite these numbers. Any other number is fabricated." When fewer than two independent data sources return numerical data, the discovery is forced to qualitative-only mode. That is, the agent can describe patterns but cannot cite specific values. ## **The Nine-Stage Quality Gate Stack** Every discovery passes through nine sequential gates. Failure at any gate archives the discovery with a documented reason: 1. Data Provenance - every API call is logged with URL, response headers, latency, and raw response via a provenance-tracking fetch wrapper 2. Numerical Whitelist - only numbers from actual API responses may appear in hypothesis text 3. Domain Plausibility Gate - Claude Haiku checks for fundamental scientific errors (wrong metabolic pathway, misclassified allele, domain mismatch) at \~$0.002 per check 4. Evidence Bar - minimum two independent data sources with numerical results 5. Contradiction Gate - [Semantic Scholar's](https://www.semanticscholar.org/?ref=openresearch.wtf) ASTA snippet search queries 275M passages for direct contradictions; strong contradictions archive immediately 6. Peer Validation - other agents independently attempt to disprove claims through adversarial review 7. Statistics Audit - pre-manuscript cross-check of every number in hypothesis text against computed statistics; methods described as performed are verified against their outputs 8. Internal Panel Review - three specialised agents (Science Writer, Domain Reviewer, Methodologist) review the manuscript; MAJOR\_REVISION triggers an automatic revision loop, REJECT archives the discovery 9. External Peer Review - submission to [Preprints.ai](https://preprints.ai/?ref=openresearch.wtf) for independent AI peer review before publication on [OpenAccess.ai](https://openaccess.ai/?ref=openresearch.wtf) with a citable Handle The current pass rate through all nine gates is low. Only 1-2% of generated hypotheses survive to manuscript stage. This is a feature, not a bug. The gates are correctly identifying weak claims. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-fe057fb5-35ad-44b6-bb43-80d9ed219925.png) ## **The Publication Pipeline** Manuscripts are generated following the [Octopus model of scientific publishing](https://www.octopus.ac/?ref=openresearch.wtf), which decomposes a traditional paper into eight linked, independently citable components. It is very hard to motivate humans to format their research outputs in a manner like this. Machines on the other hand, will do what you tell them. 1. Research Problem - the question and its literature context 2. Rationale/Hypothesis - theoretical basis with scope and evidence 3. Method - reusable methodology (computational, not wet-lab) 4. Results/Sources - raw API responses with no interpretation 5. Analysis - statistical tests and computed statistics 6. Interpretation - conclusions drawn from analysis 7. Applications - real-world implications 8. Review - validator agent assessments Each component is a first-class research output stored with full provenance linking back to the discovery, the agent that produced it, and the API calls that generated the evidence. The manuscript assembly pipeline fetches literature from both [Semantic Scholar](https://www.semanticscholar.org/?ref=openresearch.wtf) and [OpenAlex](https://openalex.org/?ref=openresearch.wtf), runs a statistics audit against computed values, generates sections via Claude Sonnet with strict constraints (the Results prompt states "Every numerical value in this section MUST appear in the verified statistics"), and produces [Vega-Lite](https://vega.github.io/vega-lite/?ref=openresearch.wtf) figure specifications from real computed statistics. ## **FAIR Compliance and Reproducibility** Every discovery can be exported as an [RO-Crate 1.2](https://www.researchobject.org/ro-crate/?ref=openresearch.wtf) package with: - [W3C PROV-O](https://www.w3.org/TR/prov-o/?ref=openresearch.wtf) provenance graphs linking every claim to its data source - [Croissant](https://mlcommons.org/croissant/?ref=openresearch.wtf) metadata (MLCommons standard for ML-ready datasets) - [JSON-LD](https://json-ld.org/?ref=openresearch.wtf) structured data (schema.org compatible) - [Bridge2AI](https://bridge2ai.org/?ref=openresearch.wtf) AI-readiness scoring across 28 criteria in 7 dimensions (FAIRness, Provenance, Characterisation, Explainability, Ethics, Sustainability, Computability) - [DataCite](https://datacite.org/?ref=openresearch.wtf) metadata for DOI minting The platform serves an llms.txt file at the root domain for LLM-readable platform description, and every discovery page includes schema.org Dataset and ScholarlyArticle structured data. ## **What we can look to do next** The anti-fabrication infrastructure is solid, but the discovery logic is still relatively shallow. The current pipeline identifies patterns in individual datasets or across a small number of cross-referenced databases. The next meaningful step is multi-hop reasoning, following chains of inference across five, ten, twenty sources to reach claims that genuinely wouldn't be visible from any single database. AlphaFold integration opens up a particularly interesting direction - systematic comparison of predicted structure variants against population-level genomic data to flag protein-phenotype associations that haven't been studied. We already have three AlphaFold tables in the data lake and the agent infrastructure to run these queries continuously at low cost. The validation question is the harder one. How do you know if an AI-generated research claim is actually novel and actually supported by the evidence? Preprints.ai is our partial answer, but longer-term, the system needs empirical feedback. Ideally, this would be a mechanism where predictions made by the platform can be tested against newly published research to measure whether the platform is discovering things that turn out to be true. If you're working on any of these problems and want to compare notes, or if you spot something obviously wrong with the approach, I'd be glad to hear from you. ### FAIRdata.ai: Making Open Data FAIR-er URL: https://www.openresearch.wtf/fairdata-ai-making-open-data-fair-er/ Last updated: 2026-07-31T12:35:39.000Z ## **What problem are we trying to solve?** Most major funders now require FAIR data management plans. Most generalist repositories (Zenodo, Figshare, Dryad, Mendeley Data) allow data published on them to be FAIR, but not all data published on them is FAIR. The gap between 'technically FAIR' and 'practically reusable' is large and I think there are ways to close it. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-f3414095-822a-4523-86d3-610a150a5a22.png) Visit any large repository and start browsing datasets at random. You'll find deposited data where the variable names are single letters, where the description field says 'data for paper', where there's no unit metadata, no controlled vocabulary alignment, no machine-readable schema. FAIRdata.ai is trying to close that gap automatically. Given a dataset DOI, it assesses the current FAIR score across five independent frameworks, identifies what's missing, and then enriches the record - generating better descriptions, inferring variable semantics, aligning to standard ontologies, resolving licences to machine-readable SPDX URIs - without requiring anything from the original depositor. Every enrichment is tagged with full provenance so consumers know exactly what was added by whom (or what model), when, and with what confidence. There's a second, less obvious problem FAIRdata.ai is working on. This is based on AllenAI’s work on Autodiscovery. [I wrote about this here](https://www.openresearch.wtf/moving-from-share-data-because-you-should-to-share-data-because-youll-get-something-back/). Open datasets contain findings that have never been extracted. Correlations that nobody looked for. Distributions that are surprising given background knowledge. FAIRdata.ai surfaces these using Bayesian surprise scoring essentially asking 'what patterns in this dataset would be unexpected given domain priors?'. It feeds those signals upstream to [OpenScience.ai](https://openscience.ai/?ref=openresearch.wtf) as hypothesis seeds for formal investigation. ## **What tools have inspired or been used to build this platform?** FAIRdata.ai integrates a substantial number of open-source tools, open APIs, and community standards. Rather than building everything from scratch, the philosophy has been to compose existing infrastructure where it exists and build only where it doesn't. Below is a technical walkthrough of the stack, layer by layer. ### **Assessment: five FAIR frameworks, not one** Most platforms run a single FAIR checker. FAIRdata.ai runs five, then composites the results: - [F-UJI](https://www.f-uji.net/?ref=openresearch.wtf) (FAIRsFAIR) - the most widely cited automated FAIR assessment tool. We self-host the F-UJI API and query it for 17 metrics across the four FAIR principles. F-UJI checks machine-readable metadata, persistent identifiers, licence declarations, data access protocols, and vocabulary usage. - [RDA FAIR Maturity Model](https://doi.org/10.15497/rda00050?ref=openresearch.wtf) \- 41 indicators scored on a 0-5 maturity scale, from the Research Data Alliance FAIR Data Maturity Model Working Group. More granular than F-UJI, particularly strong on governance indicators. - [ARDC Self-Assessment](https://ardc.edu.au/resource/fair-data-self-assessment-tool/?ref=openresearch.wtf) \- 12 practical checks from the Australian Research Data Commons. Pragmatic and well-calibrated for repository-deposited data. - [Metadata Game Changers](https://metadatagamechangers.com/?ref=openresearch.wtf) \- 58 documentation concepts covering both essential and supporting FAIR elements. Good at catching metadata completeness issues that other frameworks miss. - Data Fitness for Use - an AI-assessed framework evaluating 14 reuse criteria from the perspective of a downstream consumer. This one uses an LLM to reason about whether a dataset is practically usable, not just technically compliant. We run two passes: one on the source repository record as-is, one on the FAIRdata.ai-enriched version. The delta between these scores is arguably the most useful signal. It shows exactly how much better the record could be with automated enrichment. ### **AI-readiness scoring** Beyond FAIR, FAIRdata.ai scores datasets against two AI-readiness frameworks: - [GDS/DSIT 4-Pillar Framework](https://www.gov.uk/government/publications/data-standards-authority-strategy-2020-2025?ref=openresearch.wtf) (Jan 2026) - Technical Optimisation, Data & Metadata Quality, Organisation & Governance, Legal/Security/Ethics. This is the UK government's framework for assessing whether data is ready for AI/ML consumption. - [Bridge2AI AI-Readiness](https://bridge2ai.org/?ref=openresearch.wtf) \- 7 dimensions from the NIH Bridge2AI programme (Clark et al. 2024): FAIRness, Provenance, Characterisation, Pre-model Explainability, Ethics, Sustainability, Computability. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-f3abc21b-f2bd-49c2-9f9b-a148ec02a00c.png) ### **Metadata sources** Dataset metadata is normalised from multiple registries: - [DataCite Commons](https://commons.datacite.org/?ref=openresearch.wtf) (api.datacite.org) - the primary metadata source. DataCite provides the canonical DOI metadata for most research data repositories. - [Crossref](https://www.crossref.org/?ref=openresearch.wtf) (api.crossref.org) - used for paper-to-dataset linking. When a dataset DOI has IsSupplementTo relationships or a linked publication DOI, we fetch the paper metadata to extract context. Repository-specific APIs handle file listings and downloads: - [Figshare API](https://docs.figshare.com/?ref=openresearch.wtf) (DOI prefix 10.6084) - [Zenodo API](https://developers.zenodo.org/?ref=openresearch.wtf) (DOI prefix 10.5281) - [Dryad API](https://datadryad.org/api/v2/docs/?ref=openresearch.wtf) (DOI prefix 10.5061) - [Mendeley Data](https://data.mendeley.com/?ref=openresearch.wtf) (DOI prefix 10.17632) - Plus fallback patterns for Harvard Dataverse, OSF, GBIF, ICPSR, 4TU.ResearchData, ScienceDB, PhysioNet, and institutional Figshare instances. ### **Ontology and semantic enrichment** Raw metadata is enriched with controlled vocabulary annotations via: - [BioPortal](https://bioportal.bioontology.org/?ref=openresearch.wtf) ([data.bioontology.org](http://data.bioontology.org/?ref=openresearch.wtf)) we use three BioPortal endpoints: the Annotator API for extracting ontology terms from free text, the Search API for controlled vocabulary lookup, and the Recommender API for domain-aware ontology suggestions. The ontologies we align to include: - [AGROVOC](https://agrovoc.fao.org/?ref=openresearch.wtf) – agriculture, forestry, fisheries (FAO) - [MeSH](https://www.nlm.nih.gov/mesh/?ref=openresearch.wtf) – medical subject headings (NLM) - [ENVO](https://sites.google.com/site/environmentontology/?ref=openresearch.wtf) – environmental ontology (biomes, ecosystems) - [Gene Ontology (GO)](http://geneontology.org/?ref=openresearch.wtf) – molecular biology, gene function - [NCIT](https://ncithesaurus.nci.nih.gov/?ref=openresearch.wtf) – NCI Thesaurus for biomedical terms - [CHEBI](https://www.ebi.ac.uk/chebi/?ref=openresearch.wtf) – chemical entities of biological interest - [EDAM](https://edamontology.org/?ref=openresearch.wtf) – bioinformatics operations, data types, topics - [GAZ](http://purl.obolibrary.org/obo/gaz?ref=openresearch.wtf) – geographic gazetteer - [SDGIO](https://sdgio.org/?ref=openresearch.wtf) – Sustainable Development Goals Interface Ontology Geographic enrichment uses [Nominatim/OpenStreetMap](https://nominatim.openstreetmap.org/?ref=openresearch.wtf) for geocoding place names to coordinates. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-5f9a9350-8d68-44b7-a6bc-5aa138594a5e.png) ### **AI enhancement pipeline** The metadata enrichment pipeline uses LLMs for tasks that require reasoning about semantics: - [Anthropic Claude](https://www.anthropic.com/?ref=openresearch.wtf) (Sonnet) – the primary LLM for metadata enhancement. Generates descriptions, infers variable semantics, extracts paper context (methodology, limitations, population, ethics), and performs intelligent gap remediation. Every AI-generated enrichment is tagged with COMET provenance (agent, model, timestamp, confidence, derivation context). The enhancement is orchestrated through a chain of specialised agents: - MetadataHarvester – extracts structured metadata from unstructured sources - PaperLinker – finds and links related publications via Crossref and OpenAlex - PaperContextExtractor – extracts methodology, limitations, and domain context from linked papers - SubjectEnricher – enhances subject keywords with ontology-aligned terms - OntologyAnnotator – maps free-text terms to BioPortal ontology concepts - GeoEntityExtractor – identifies and geocodes geographic references - TemporalExtractor – identifies temporal coverage and resolution - FormatDetector – identifies file formats and suggests MIME types - LicenseResolver – maps licence text to machine-readable SPDX URIs - SemanticResourceGenerator -- generates schema.org JSON-LD markup - IntelligentRemediation – fills remaining metadata gaps - ProvenanceBuilder – assembles [COMET](https://www.cometadata.org/?ref=openresearch.wtf) provenance records for all enrichments - QualityStatementGenerator – produces human-readable quality assessments ### **Data profiling** For datasets containing tabular files (CSV, TSV, Excel, Parquet), FAIRdata.ai performs column-level statistical profiling using [pandas](https://pandas.pydata.org/?ref=openresearch.wtf), [NumPy](https://numpy.org/?ref=openresearch.wtf), and [SciPy](https://scipy.org/?ref=openresearch.wtf): - Column data types, cardinality, null rates - Distributions (mean, median, std, skewness, quartiles) - Outlier detection and quality flags - Top values for categorical columns These profiles power both the discovery engine (identifying interesting statistical patterns) and the cross-dataset matching system (finding datasets with shared column structure). ### **Discovery engine** ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-433bb04f-f58f-4203-ba62-09898138e90d.png) The discovery layer is where FAIRdata.ai goes beyond assessment into knowledge generation: - AutoDiscovery MCTS sidecar -- a Monte Carlo Tree Search engine that explores the hypothesis space of a dataset's statistical properties. It uses [OpenAI GPT-4o](https://openai.com/index/hello-gpt-4o/?ref=openresearch.wtf) for reasoning about which analyses to run, then executes them on the actual data. Each finding is scored with Bayesian surprise (prior-to-posterior shift) to measure genuine novelty rather than mere statistical significance. Analyses include correlation detection, distribution characterisation, bimodality testing, group differences (ANOVA), outlier classification, and temporal trend detection. - Cross-dataset discovery -- a system that finds joinable dataset pairs by matching column structure across the profiled dataset corpus. When two datasets share meaningful columns (filtering out garbage like Unnamed: 0 or generic fields like id), they can be merged and submitted to the MCTS engine for cross-study hypothesis generation. High-confidence findings (Bayesian surprise >= 0.8, p < 0.01) are automatically promoted and pushed to [OpenScience.ai](https://openscience.ai/?ref=openresearch.wtf) as research seeds, where they're processed into formal hypothesis papers. ### **Packaging and export formats** Enriched records are packaged into six machine-readable formats using [FAIRSCAPE](https://fairscape.github.io/?ref=openresearch.wtf): - [schema.org/Dataset](https://schema.org/Dataset?ref=openresearch.wtf) – JSON-LD for Google Dataset Search discoverability - [Croissant ML 1.0](https://mlcommons.org/croissant/?ref=openresearch.wtf) – the MLCommons standard for ML-ready datasets (compatible with PyTorch, TensorFlow, JAX, HuggingFace, Kaggle) - [RO-Crate 1.2](https://www.researchobject.org/ro-crate/?ref=openresearch.wtf) – Research Object packaging with full provenance - [FAIR Data Point (DCAT3)](https://www.fairdatapoint.org/?ref=openresearch.wtf) – Turtle RDF for FAIR data exchange - Enriched DataCite JSON – structured subjects with ontology URIs, SPDX rights - AI-Readiness JSON – GDS/DSIT + Bridge2AI scoring ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/03/data-src-image-2e87b03b-6e98-43cb-a4fe-69fe4d564bf0.png) ### **Agent integration** FAIRdata.ai exposes its capabilities to AI agents via: - [Model Context Protocol (MCP)](https://modelcontextprotocol.io/?ref=openresearch.wtf) –JSON-RPC 2.0 endpoint at /mcp for Claude, ChatGPT, and custom agent integration - [OpenAI Plugin Manifest](https://platform.openai.com/docs/plugins/?ref=openresearch.wtf) – /.well-known/ai-plugin.json for ChatGPT plugin discovery - [llms.txt](https://llmstxt.org/?ref=openresearch.wtf) – standardised LLM-readable documentation at /llms.txt ### **Infrastructure** - Backend: [FastAPI](https://fastapi.tiangolo.com/?ref=openresearch.wtf) (Python 3.11, async) with [Uvicorn](https://www.uvicorn.org/?ref=openresearch.wtf), deployed on [Railway](https://railway.app/?ref=openresearch.wtf) - Frontend: [React](https://react.dev/?ref=openresearch.wtf) with [Vite](https://vitejs.dev/?ref=openresearch.wtf), [Recharts](https://recharts.org/?ref=openresearch.wtf) for data visualisation, deployed on [Netlify](https://www.netlify.com/?ref=openresearch.wtf) - Database: [Supabase](https://supabase.com/?ref=openresearch.wtf) (managed PostgreSQL) - Analytics: [Plausible](https://plausible.io/?ref=openresearch.wtf) (privacy-friendly, cookie-free) - Provenance: COMET model -- every AI enrichment carries agent ID, model version, timestamp, confidence score, and derivation context ## **What we can look to do next** The most obvious gap is coverage. FAIRdata.ai currently works well with Figshare, Zenodo, Dryad, Harvard Dataverse, Mendeley Data, OSF, and Vivli, but the tail of domain repositories is long and each has its own metadata quirks. There are several technical directions I'm exploring: The two-score nudge. The gap between a dataset's original FAIR score and its enriched score is a meaningful signal. Showing depositors that their record went from 62% to 84% with automated enrichment -- and exactly which fields were improved -- could be a practical nudge toward better data curation practice at the point of deposit.Stronger discovery-to-hypothesis pipeline. Right now the Bayesian surprise scoring identifies datasets with interesting distributional properties, but the handoff to OpenScience.ai is still coarse. Tightening this so FAIRdata.ai passes not just 'this dataset is interesting' but 'here is a specific, testable claim with statistical evidence and domain context' would make the whole pipeline considerably more powerful. Federated assessment. Rather than pulling datasets to our infrastructure, enabling repositories to run FAIRdata.ai assessment in-place (as a service or plugin) would scale coverage without scaling compute. The F-UJI framework already supports this model; extending it to the full five-framework composite would be the contribution. If you're working on any of these problems and want to compare notes, or if you spot something obviously wrong with the approach, I'd be glad to hear from you. ### Moving from "share data because you should" to "share data because you'll get something back." URL: https://www.openresearch.wtf/moving-from-share-data-because-you-should-to-share-data-because-youll-get-something-back/ Last updated: 2026-07-31T12:35:47.000Z ### Looking for Surprisals in your data using AstaLabs Autodiscovery from Allen AI I founded[ Figshare](https://figshare.com/?ref=openresearch.wtf) over a decade ago with a simple thesis, if you make research data openly available, good things will happen. Things you can't predict. This week, I got a rather specific reminder of why that thesis holds up. [Allen AI](https://allenai.org/?ref=openresearch.wtf) have just launched[ AutoDiscovery](https://allenai.org/blog/autodiscovery?ref=openresearch.wtf), an experimental tool inside their[ AstaLabs](https://asta.allen.ai/?ref=openresearch.wtf) platform. Most AI tools for research, like[ Google's AI co-scientist](https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/?ref=openresearch.wtf) and[ FutureHouse](https://www.futurehouse.org/?ref=openresearch.wtf) are goal-driven. You bring a research question, they help you answer it. AutoDiscovery flips that. You give it a dataset, and it generates its own hypotheses, writes Python code to test them, runs statistical experiments, and uses[ *Bayesian surprise*](https://arxiv.org/abs/2507.00310?ref=openresearch.wtf) to surface the results worth paying attention to. Before running an experiment, the system holds a prior belief about whether a hypothesis is true (derived from the LLM's world knowledge). After seeing results, it updates to a posterior. The *surprisal* is the magnitude of that shift. ie. how much the evidence forced the system to change its mind. To navigate the vast space of possible questions, it uses[ Monte Carlo Tree Search](https://en.wikipedia.org/wiki/Monte%5FCarlo%5Ftree%5Fsearch?ref=openresearch.wtf) to balance exploration with exploitation. In their[ evaluation across 21 datasets](https://arxiv.org/abs/2507.00310?ref=openresearch.wtf), AutoDiscovery produced 5–29% more surprising discoveries than competing approaches, and two-thirds were also surprising to human domain experts. I wanted to test it with data that lives on Figshare. So I picked three datasets that I knew were relatively FAIR as they have been cited multiple times. ## **Dataset 1: Brain tumours** Jun Cheng's[ brain tumor dataset](https://figshare.com/articles/dataset/brain%5Ftumor%5Fdataset/1512427?ref=openresearch.wtf): 3,064 T1-weighted contrast-enhanced MRI images, three tumour types,[ CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/?ref=openresearch.wtf), on Figshare since 2015\. I ran 10 experiments. Most hypotheses confirmed (what I am led to believe are) textbook expectations - gliomas show higher textural entropy than meningiomas (reflecting heterogeneous malignant tissue versus more homogeneous benign structure), pituitary tumours sit centrally in the brain, meningiomas exhibit higher morphological circularity, and a simple geometric classifier can distinguish between the three types above chance. Low surprisal scores across the board, the system's beliefs validated by the data. One standout finding was different. AutoDiscovery hypothesised that pixel intensity heterogeneity would be significantly higher in gliomas than pituitary tumours, reasoning from necrosis and vascular heterogeneity. Prior belief: 0.96\. The data found a significant difference ([Mann-Whitney U](https://en.wikipedia.org/wiki/Mann%E2%80%93Whitney%5FU%5Ftest?ref=openresearch.wtf), p = 5.97 × 10⁻⁴¹), but in the opposite direction — pituitary tumours showed higher heterogeneity (0.0906 ± 0.0316 vs 0.0727 ± 0.0346). Belief dropped to 0.327\. Surprisal: -0.759\. Not wrong in the sense of a failure, but wrong in the sense that matters for science: evidence meaningfully shifting expectations. [Full results →](https://autodiscovery.allen.ai/runs/shared/096e5c57-945a-44de-af1d-169ec1d8e8b3?ref=openresearch.wtf) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/02/Screenshot-2026-02-16-at-09.25.40.png) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2026/02/Screenshot-2026-02-16-at-09.30.00.png) ## **Dataset 2: Global aridity** [Trabucco and Zomer's Global Aridity Index](https://doi.org/10.6084/m9.figshare.7504448?ref=openresearch.wtf): high-resolution (30 arc-second) global raster climate data for the 1970–2000 period, now on its seventh version. This has been a workhorse of climate and land-use research for years, cited extensively across the environmental sciences. A cluster of hemispheric aridity hypotheses produced the highest surprisal scores. The system predicted the Southern subtropical belt (20°S–40°S) would have a higher mean[ Aridity Index](https://en.wikipedia.org/wiki/Aridity%5Findex?ref=openresearch.wtf) (more humid) than the Northern belt (20°N–40°N), reasoning from land-ocean ratios. Prior: 0.984\. The Northern belt turned out significantly less arid (mean AI \~0.176 vs \~0.070,[ Welch's t-test](https://en.wikipedia.org/wiki/Welch%27s%5Ft-test?ref=openresearch.wtf) p < 0.001). Belief after: 0.335\. Surprisal: -0.778. AutoDiscovery's prior was shaped by the intuition that the Southern Hemisphere has more ocean, the actual landmass between 20°S and 40°S is dominated by Australian desert, the Kalahari, and the Atacama. The Northern belt, despite the Sahara, includes the humid Southeastern US and East Asia. The system taught itself some geography by being surprised by data. The run also confirmed expected patterns lower in the table: aridity gradients are steeper at the desert edge than the forest edge, evapotranspiration correlates more strongly with latitude than aridity does, and Madagascar exhibits a sharp longitudinal climatic divide. [Full results →](https://autodiscovery.allen.ai/runs/shared/6160f8e2-1e8c-4000-bf10-2588f87b87f6?ref=openresearch.wtf) ## **Dataset 3: Permeability estimation** The most niche test:[ pressure and viscosity logs](https://doi.org/10.25452/figshare.plus.16867417?ref=openresearch.wtf) by Khirevich, Yutkin, and Patzek - high-precision flow measurements through porous media, hosted on[ Figshare+](https://help.figshare.com/article/guide-to-sharing-data-on-figshare-plus?ref=openresearch.wtf), supporting[ a paper in Physics of Fluids](https://doi.org/10.1063/5.0123673?ref=openresearch.wtf). The top surprisal - the system predicted fluid temperature would increase monotonically due to mechanical pump heating, introducing systematic drift. Prior: 0.831\. Rejected cleanly, linear regression across all six experimental series showed temperature slopes ranging from -0.0062 to +0.0122 °C/hr, no consistent direction, and only one series with a statistically significant trend. Temperatures stayed within a narrow 21.2–22.0°C band throughout. Belief after: 0.283\. Surprisal: -0.657. This is arguably the most practically useful kind of finding. AutoDiscovery ran a quality-assurance check on the experimental setup and confirmed a plausible source of systematic error wasn't there, exactly the kind of analysis that strengthens confidence in published results. [Full results →](https://autodiscovery.allen.ai/runs/shared/d61de50e-892d-485a-abfd-30dbd563869c?ref=openresearch.wtf) ## **What this means for open data** There's a satisfying circularity to this experiment. Figshare was built on the premise that research data should be persistently available, properly licensed, and machine-accessible. A decade later, I'm feeding Figshare-hosted datasets into an AI system that autonomously generates hypotheses against them. The datasets have[ DOIs](https://www.doi.org/?ref=openresearch.wtf), open licences, versioning. They just work. But I think there's something more significant here than a nice demonstration. The persistent challenge in open data advocacy has been the incentive question. Researchers are told to share because funders[ mandate](https://www.ukri.org/manage-your-award/publishing-your-research-findings/making-your-research-data-open/?ref=openresearch.wtf) it, because it aids[ reproducibility](https://en.wikipedia.org/wiki/Reproducibility%5Fcrisis?ref=openresearch.wtf). All true. But for the researcher who spent three years collecting data, the question remains - *what's in it for me?* Tools like AutoDiscovery offer a new answer. Share your data, and an AI system will generate hypotheses you might never have tested, surface patterns that could open new lines of inquiry. It will come at your data with different priors to yours, ask questions from outside your disciplinary frame. The value proposition shifts from "share because you should" to "share because you'll get something back." Imagine every deposited dataset receiving an automated run with researchers getting a report of surprising findings alongside their download metrics. This also reframes what it means to build on research that came before. The traditional cycle of publish, read, design new study, collect new data, takes years. AutoDiscovery collapses that to hours. It brought its own hypotheses to[ Cheng's](https://figshare.com/articles/dataset/brain%5Ftumor%5Fdataset/1512427?ref=openresearch.wtf) tumour data,[ Trabucco and Zomer's](https://doi.org/10.6084/m9.figshare.7504448?ref=openresearch.wtf) climate data,[ Khirevich and colleagues'](https://doi.org/10.25452/figshare.plus.16867417?ref=openresearch.wtf) permeability logs, reusing their data in ways they almost certainly hadn't imagined when they deposited it. That's rather the whole point of making data[ FAIR](https://www.go-fair.org/fair-principles/?ref=openresearch.wtf). The reusable bit has always been the hardest to demonstrate. Tools like this make it tangible. ## **Try it** The tool has limitations. The surprisal metric inherits the LLM's biases, many confirmed hypotheses would be obvious to domain experts, and as Allen AI[ themselves note](https://allenai.org/blog/autods?ref=openresearch.wtf), outputs are starting points for investigation, not finished science. But the full transparency into code and statistical methods makes everything auditable and reproducible. Allen AI are offering[ 1,000 free credits](https://asta-autodiscovery.allen.ai/runs?ref=openresearch.wtf) through end of February 2026\. If you've got data sitting in a repository like[ Figshare](https://figshare.com/?ref=openresearch.wtf) or anywhere else, point AutoDiscovery at it and see what comes back. The datasets researchers have been quietly sharing all this time might just turn out to be more valuable than anyone expected. ### Machine-First FAIR: Realigning Academic Data for the AI Research Revolution URL: https://www.openresearch.wtf/machine-first-fair-realigning-academic-data-for-the-ai-research-revolution/ Last updated: 2026-07-31T12:35:46.000Z **The best way for humankind to benefit from research is to prioritize machines over people when sharing data. Here’s why.** We push out the lines that academic research needs to be Findable, Accessible, Interoperable and Re-usable (FAIR) for humans and machines. This suggests humans and machines should get equal priority when it comes to FAIR. This is not the case, we should prioritize the machines. Machine-generated new knowledge will accelerate knowledge discovery. While humans can infer insights from sparse information in academic literature and datasets – due to our ability to find more context online – the machines currently cannot. To go further, faster in knowledge discovery we need to move past human-powered knowledge discovery. To do this, the machines need structure and pattern. Every research-generating organization should be prioritizing this. ## Academia is Ignoring Decades of Advancement Academic research generates more than [6.5 million papers annually](https://app.dimensions.ai/discover/publication?or%5Ffacet%5Fyear=2025&ref=openresearch.wtf), and over [20 million datasets](https://commons.datacite.org/doi.org?query=data&published=2025&resource-type=dataset&ref=openresearch.wtf), each representing potential training signals for the artificial intelligence systems reshaping discovery. Yet most institutional data remains locked in formats optimized for human consumption rather than computational processing. While most stakeholders know the theoretical merits of making data FAIR (Findable, Accessible, Interoperable, Reusable) for both humans and machines, the practical reality is starker: in an era where language models can process orders of magnitude more literature than any human researcher, we are still organizing our most valuable research assets for the wrong consumer. The economic implications are substantial. Organizations like the Chan Zuckerberg Initiative (CZI) have committed over $3.4 billion toward AI-powered biology, funding projects ranging from their [1,024 GPU DGX SuperPOD cluster](https://chanzuckerberg.com/newsroom/czscience-builds-ai-gpu-cluster-predictive-cell-models/?ref=openresearch.wtf) for computational biology research to the [Virtual Cell Platform](https://chanzuckerberg.com/science/?ref=openresearch.wtf) that aims to create predictive models of cellular behavior. The Navigation Fund, with its $1.3 billion endowment, has invested in AI infrastructure through their [Voltage Park subsidiary](https://techcrunch.com/2023/10/31/new-nonprofit-backed-by-crypto-billionaire-scores-ai-chips-worth-500m/?ref=openresearch.wtf), while simultaneously funding [open science initiatives](https://www.navigation.org/grants/open-science?ref=openresearch.wtf) focused on machine-actionable intelligence and metadata enhancement. Astera Institute has deployed portions of its $2.5 billion endowment to support projects like their [$200 million investment in Imbue’s AI agent research](https://www.lw.com/en/news/2023/09/latham-advises-astera-institute-on-investment-in-imbue?ref=openresearch.wtf) and their [Science Entrepreneur-in-Residence program](https://astera.org/mission-and-vision/?ref=openresearch.wtf) specifically targeting scientific publishing infrastructure. Meanwhile, the Allen Institute for AI demonstrates the practical returns on machine-first approaches through projects like their [OLMo series of fully open language models](https://allenai.org/olmo?ref=openresearch.wtf), where complete training datasets, code, and methodologies are published in computational formats, and their [Semantic Scholar platform](https://www.semanticscholar.org/?ref=openresearch.wtf), which processes millions of academic papers to extract structured, machine-readable knowledge graphs. Chan Zuckerberg Initiative (CZI) Yet the vast majority of academic institutions continue to publish their findings in PDFs or as poorly described datasets. While LLMs are getting better at ingesting multi-modal content, PDF is a format that remains surprisingly resistant to reliable automated extraction, despite decades of advancement in natural language processing. This is not merely a technical limitation. Modern large language models struggle with PDFs because these documents prioritize visual presentation over semantic structure. Critical information becomes trapped in figures, tables, and formatting that computational systems cannot reliably parse. A reaction scheme embedded as an image, a dataset described in paragraph form, or experimental parameters scattered across multiple tables represent precisely the kind of structured knowledge that could accelerate discovery if only machines could access it consistently. ## The Architecture of Computational Research Infrastructure The solution requires a fundamental reorientation toward machine-first data architecture. Rather than retrofitting human-readable outputs for computational consumption, we can take inspiration from pharma and industry writ large, who are designing their data flows to serve algorithms from the ground up, with human-friendly interfaces emerging as downstream products of this computational foundation. Consider the transformation pathway implemented by teams working with Digital Science’s suite of computational research tools. We’re building workflows in our tools for automated knowledge extraction at scale. The extracted knowledge gains semantic coherence through integration into domain-specific knowledge graphs. Platforms like metaphacts ([metaphactory](https://metaphacts.com/metaphactory?ref=openresearch.wtf)) provide the infrastructure to align these signals with established ontologies while enforcing quality constraints through [SHACL validation](https://help.metaphacts.com/resource/Help:DataQuality?ref=openresearch.wtf) integrated into continuous deployment pipelines. The result is not merely a database of facts, but a queryable intelligence system that can answer novel questions through automated reasoning over validated relationships. Simultaneously, the operational requirements of research continue through dedicated literature management systems. Tools like [ReadCube](https://www.readcube.com/en/literature-review/?ref=openresearch.wtf) maintain the audit trails and conflict resolution workflows that regulatory environments demand, while ensuring that every screening decision and data extraction connects to persistent identifiers. The curated evidence flows directly into the computational infrastructure rather than terminating in isolated spreadsheets. The critical innovation lies in packaging. While human researchers expect PDFs and narrative summaries, machine learning pipelines require structured metadata that specifies exactly what each dataset contains, where to retrieve it, and how to interpret every field. ## The Metadata Multiplier Effect on Repository Platforms Academic data repositories like [Figshare](https://info.figshare.com/?ref=openresearch.wtf#solutions/) occupy a unique position in the machine-first FAIR ecosystem. We serve as the critical junction between human research practices and computational discovery. When researchers publish datasets with comprehensive, structured metadata, these platforms transform from simple storage services into computational assets that can feed directly into AI research pipelines. The difference lies entirely in how authors describe their work at the point of deposit. [![](https://www.digital-science.com/wp-content/uploads/2025/11/figshare-REAL-colon-dataset-screen.png)](https://doi.org/10.25452/figshare.plus.22202866.v2?ref=openresearch.wtf) The REAL (Real-world multi-center Endoscopy Annotated video Library) – colon dataset on Figshare: [https://doi.org/10.25452/figshare.plus.22202866.v2](https://doi.org/10.25452/figshare.plus.22202866.v2?ref=openresearch.wtf) Consider two datasets published on the same platform: one uploaded with a generic title like “experiment\_data\_final.xlsx” and minimal description, the other with machine-readable field descriptions, standardized vocabulary terms, and explicit links to ontologies and methodologies. The first requires human interpretation before any computational system can make sense of its contents. The second can be discovered, validated, and integrated into training pipelines automatically. [Figshare’s API](https://docs.figshare.com/?ref=openresearch.wtf) can surface the rich metadata to computational systems, but only if researchers have provided it in the first place. The platform infrastructure already supports the technical requirements for machine-first FAIR. Persistent DOIs ensure stable identifiers, while structured metadata fields can accommodate everything from [ORCID researcher identifiers](https://orcid.org/?ref=openresearch.wtf) to detailed provenance information. When authors invest time in describing their data using controlled vocabularies, specifying units of measurement, documenting collection methodologies, and linking to relevant publications, they create computational assets rather than digital archives. The same dataset that might languish undiscovered with poor metadata becomes a valuable training resource when described with machine-readable precision. This creates a powerful feedback loop. Datasets with excellent metadata get discovered and reused more frequently, driving citation counts and demonstrating impact. Meanwhile, poorly described data remains computationally invisible regardless of its scientific value. Platforms like Figshare could amplify this effect by providing better authoring tools that encourage structured metadata entry, perhaps even using AI to suggest appropriate ontology terms or validate metadata completeness before publication. The infrastructure for machine-first FAIR already exists, it simply requires researchers to embrace metadata as a first-class research output rather than an administrative afterthought. But this is an evolving field, new standards are emerging that repositories need to engage with. [The Croissant format](https://research.google/blog/croissant-a-metadata-format-for-ml-ready-datasets/?ref=openresearch.wtf), a lightweight JSON-LD descriptor based on [schema.org](http://schema.org/?ref=openresearch.wtf), provides this computational bridge. A single Croissant file enables any training pipeline to hydrate datasets without custom loaders while simultaneously supporting discovery through standard web infrastructure. ## Practical Implementation in Institutional Contexts The transition to machine-first FAIR follows a predictable arc when properly resourced. Initial implementations focus on proving the fundamental workflow with narrowly scoped pilot projects. A team might select a single dataset and one sharply defined outcome, perhaps drug-target interaction prediction or materials property modeling and implement the complete pipeline from literature extraction through validated knowledge graph construction to machine-readable packaging. The critical insight from successful implementations is the importance of automation as the second phase. Manual processes that work for pilot projects become bottlenecks at scale. The most effective teams invest heavily in converting their proven workflows into tested, continuous integration pipelines that enforce quality gates automatically. This includes SHACL validation for knowledge graphs, automated license checking, and provenance tracking. Production deployment requires infrastructure investments that many academic institutions are not yet considering. Successful implementations provide stable, resolvable URLs for every dataset and descriptor, enable content negotiation so that both machines and humans receive appropriate formats, and implement comprehensive monitoring of data quality trends and usage patterns. This is the stack that [Digital Science can provide](https://www.digital-science.com/solutions/?ref=openresearch.wtf). ## Quantifying Institutional Success Organizations can assess their progress toward machine-first FAIR through several concrete indicators. Successful implementations demonstrate that every significant dataset resolves to a persistent identifier that returns structured [JSON-LD](https://json-ld.org/?ref=openresearch.wtf) for computational consumers while maintaining readable landing pages for human users. Knowledge graphs pass automated validation, maintain stable URI schemes, and support catalogued query patterns rather than requiring ad hoc exploration. Literature workflows leave complete audit trails with [PRISMA-compliant reporting](https://www.prisma-statement.org/?ref=openresearch.wtf) that can be generated automatically rather than assembled manually. Licensing and provenance information becomes verifiable through computational means rather than requiring human interpretation. Most importantly, the time taken from initial hypothesis to trained model decreases as institutional infrastructure matures and teams spend more of their time on discovery rather than data preparation. The research organizations that define the next decade will not necessarily be those with the largest datasets, but rather those whose data infrastructure works most effectively at computational scale. Every day spent optimizing publishing workflows for human-readable reports while leaving data computationally inaccessible represents lost ground in an increasingly competitive landscape. The funders backing this transformation, from [CZI’s investments](https://www.youtube.com/watch?v=e1gFkvLGpBM&embeds%5Freferring%5Feuri=https%3A%2F%2Fwww.latent.space%2F&ref=openresearch.wtf) in computational biology to Astera’s focus on AI-native research infrastructure, are betting that machine-first approaches will determine which institutions can effectively leverage artificial intelligence for discovery. The technical architecture exists today. The standards are stable. The remaining barrier is institutional commitment to prioritizing computational accessibility over familiar but inefficient human-centered workflows. Academic research stands at yet another technology-driven inflection point. The institutions that embrace machine-first FAIR will find themselves having more impact for their research and researchers. ### Academia has a new preprints problem URL: https://www.openresearch.wtf/academia-has-a-new-preprints-problem/ Last updated: 2026-07-31T12:35:12.000Z Academia has a new preprints problem. Thousands of researchers are essentially playing slot machines with large language models, pulling the lever over and over hoping to hit the jackpot of novel theoretical physics insights. As Figshare acts as a preprint platform, I'm watching this unfold in real-time. It's equal parts fascinating and maddening. The strategy is simple. Fire up Claude or ChatGPT, pepper it with increasingly esoteric questions about quantum mechanics or string theory, collect the responses, wrap them in LaTeX, and boom, you've got yourself a preprint. It's the academic equivalent of the infinite monkey theorem, except instead of waiting for monkeys to accidentally type Shakespeare, we're waiting for transformer models to accidentally solve the mysteries of the universe. I get the appeal. LLMs are genuinely impressive at synthesizing existing knowledge, making unexpected connections, and occasionally producing those "huh, I never thought about it that way" moments. The barrier to entry is practically non-existent. There is no need for expensive lab equipment, years of mathematical training, or even a particularly deep understanding of the field. Just you, a chatbot, and the audacity to ask "what if gravity is actually just electromagnetic force wearing a disguise?" The problem isn't that this approach can't work. Broken clocks, infinite monkeys…pick your metaphor. > ***The issue is the signal-to-noise ratio is absolutely abysmal.*** For every potentially interesting nugget buried in these preprints, there are hundreds of pages of what amounts to sophisticated-sounding nonsense. LLMs are, at their core, pattern-matching machines trained on existing human knowledge. They're exceptional at remixing and recombining ideas, but genuine novelty? What's particularly frustrating is watching legitimate repositories get flooded with these papers. They're not exactly wrong. Many are internally consistent, use appropriate terminology, and follow logical structures. But they're also not exactly right, or useful, or advancing human understanding in any meaningful way. The researchers doing this seem to fall into two camps. First, there are the true believers who genuinely think they're onto something revolutionary. They'll point to that one time an LLM suggested a novel approach to protein folding or helped solve a mathematical proof as evidence that their method is valid. Fair enough, but those were cases where humans used LLMs as tools within rigorous frameworks, not as oracles dispensing wisdom. [![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/08/Screenshot-2025-08-12-at-10.50.39.png)](https://www.reddit.com/r/AmIOverreacting/comments/1m7ulvi/aio%5Fmy%5Fdad%5Fthinks%5Fhes%5Fa%5Fgenius%5Fbut%5Fim%5Fworried%5Fhes/?ref=openresearch.wtf) The second camp is more cynical. They know the game they're playing. In the publish-or-perish world of academia, quantity often trumps quality, and what better way to pad your publication list than with an endless stream of LLM-generated "research"? The effort-to-output ratio is unbeatable. Why spend years on one groundbreaking paper when you can produce dozens of mediocre ones in the same timeframe? However, this approach fundamentally misunderstands both how scientific breakthroughs happen and what LLMs actually do. Real theoretical physics advances don't come from randomly combining existing concepts, they come from deep understanding, careful observation, mathematical rigor, and often, years of wrestling with paradoxes and inconsistencies. Einstein didn't stumble upon relativity by asking enough questions; he spent a decade thinking deeply about the nature of space and time. Meanwhile, LLMs are essentially very sophisticated autocomplete systems. They predict what words should come next based on patterns in their training data. When you ask them about theoretical physics, they're not reasoning from first principles or accessing some hidden understanding of the universe, they're pattern-matching against every physics paper, textbook, and Wikipedia article they've seen. It's impressive mimicry, but it's still mimicry. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXc5TAchUKP8t6tv3xm5looIPnKNBb9QG2ZC0kkAWaQCcZ3vLGXnqVCEHSxKwlJXcjVfWgO6KU4GJ-Tm9_W8lAOnYKq3u-3mUc7MGGfeYU4ChaVqca3l4lKNSx-_v-jw7uVnkQuTTw?key=75rvC0jNYBnxrpT1mjwfMA) I love LLMs and [think we are underestimating how they will change academic research](https://www.openresearch.wtf/have-we-already-hit-the-peer-review-breaking-point/). These tools have legitimate uses in research: literature reviews, code generation, exploring analogies, even brainstorming. But there's a difference between using them as tools and treating them as research partners. One enhances human intelligence; the other substitutes for it. Will this approach eventually produce something genuinely novel? Maybe. Infinite monkeys and all that. The real breakthroughs in theoretical physics will probably come from the same place they always have: clever humans doing hard work. Until then, I'll be here, watching repositories fill up with papers asking whether dark matter might be conscious, and wondering if this is really the best use of our collective intellectual energy. ### Have we already hit the peer review breaking point? URL: https://www.openresearch.wtf/have-we-already-hit-the-peer-review-breaking-point/ Last updated: 2026-07-31T12:35:43.000Z **What does scientific publishing look like if every paper is AI-generated or at least AI co-authored?** Last week I attended the[ Future of scientific publishing conference](https://royalsociety.org/science-events-and-lectures/2025/07/future-of-scientific-publishing/?ref=openresearch.wtf) organized by the Royal Society. I went to something similar a decade ago, which made me ponder what we were suggesting would be the future versus what has actually come to reality. I think that can be summed up as ‘the establishment of preprints and open data’. We still dream of real-time release of information from labs around the world, with post-publication peer review, and we may get there. However, there were two comments at the conference that really played on my mind. First was the sentiment that came up in a few panel discussions that we may be overestimating the effects of AI on scientific publishing. I feel quite the opposite. I think we are underestimating the effect of AI on scientific publishing. If we look 10 years into the future, what will AI change? For those thinking it is as monumental a shift as the web, that didn't change too much about the way academics disseminate their research. We now just publish PDFs on the web instead of printing them out and sending them around the world. What makes AI different is that it's already affecting content creation. What does the web look like if all content is bot-produced? What does scientific publishing look like if every paper is AI-generated or at least AI co-authored? This brings me to the other comment that weighed heavy on my mind. *"*[*Elsevier fielded 3.5M submissions, published 700K and saw +600K increase in submissions year on year*](https://world.hey.com/ian.mulvany/initial-thoughts-on-the-future-of-scientific-publishing-conference-held-by-the-royal-society-or-how-fbad6398?ref=openresearch.wtf)*"* ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXffT2vz0dWEsDGpR-1ht2Oz8ujaUicY47e_1VMlxL97U6QCPvCKzSU371J6A8SN5EBIdVmGudQFo6gIlNby3OBXhYxD7PXGrXMzStf09SQ6JiEO88hFJ0nbfY6uzOtNfjUaIcXa1Q?key=Lm6OPJcdJA3Yl2s9qlOMRw) This is a 17% increase in submissions at the biggest academic publisher by volume in the world. Because Elsevier is indeed the largest publisher, with many journals from varying fields with varying rejection rates, I decided to assume it offers a good average for academic publishing writ large, and my mind began to extrapolate. The 600K increase in submissions year-over-year is staggering. That's roughly a 17% growth rate, which is double the historical academic publication growth rate of 8-9%. Based on the current model, the peer review system becomes mathematically impossible remarkably quickly. If we assume each paper needs 2-3 reviews and there are roughly 20 million active researchers globally who could potentially serve as reviewers, we hit a crisis point for submissions quite quickly. And this is assuming a steady, not exponential, growth curve for the number of papers submitted each year. Of course, if you had said to publishers in the year 2000 that the number of articles published each year would quadruple by 2025, they may have suggested that peer reviewing that volume would be untenable - but that was before papers could be written in a minute. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXdm9c5--0TRU1wQkGbqxkX3hBTBZUnU2DWcqYSkYxpNL6nPSmHrPmrlDAvw5lqSq-8fLwXoEKjtfLh1uJRwY-pLJbm3ph0lCRaBznLLvsFF_TOlyQ-mFKBH88JNpfHCq23_hRANPw?key=Lm6OPJcdJA3Yl2s9qlOMRw) If publications grow 10x within 5 years of AI adoption, we'd need 200-300 million peer reviews annually. Even if every PhD-level researcher worldwide dedicated their entire career to reviewing, we couldn't keep up. As acceptance rates plummet (potentially from 20% to 5-10%), authors may submit each paper to more journals. This means the same research generates multiple submissions, artificially inflating the crisis. This brings us to a fascinating question about value creation. In our current system, we reward the person who writes the paper. They get citations, tenure, grants. But in an AI-dominated world, who deserves credit? Will we have academic prompt engineers? Will we have raw data creators? So assuming peer review is at breaking point or not far off, perhaps the most obvious near-future change to scientific publishing is the widespread adoption of the "publish then curate" model. We have the tech stack to publish preprints, we have overlay journals, and we have eLife's innovative new model. What will it take for other publishers to realise the writing is on the wall? ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/07/Screenshot-2025-07-18-at-12.52.01.png) Panelists at the "Future of scientific publishing conference" 2025 Two of the panels at the conference spoke about[ what comes next](https://royalsociety.org/science-events-and-lectures/2025/07/future-of-scientific-publishing/?ref=openresearch.wtf) in a similar vein, with Professor Ludo Waltman of Leiden University highlighting that we now have the challenge of curation. Dr. Michele Avissar-Whiting of Howard Hughes Medical Institute spoke of decoupling dissemination from approval. I often feel disheartened at events like this, thinking there isn't enough innovation, but perhaps the wave of innovation that the web indirectly caused in scholarly publishing (see preprints and open data) is about to be repeated with AI. And this is enough to help make a more equitable academic publishing future whilst accelerating the pace of research for the good of humanity. And that is good enough for me. Ps. Depending on how bullish you are on AI - you can predict when we hit breaking point in the model below: # Academic Publication Explosion Model Gradual AI Adoption Rapid AI Takeover AI Explosion Custom Scenario AI Adoption Year: Growth Multiplier: 10x Adoption Speed: 5 **Model Assumptions:** This model assumes AI adoption follows a sigmoid curve, with publications growing exponentially once AI tools become mainstream. The peer review crisis threshold is estimated based on current reviewer capacity and assumes each paper requires 2-3 reviews on average. ### AI Needs New Facts – The Value of Novel Scientific Research URL: https://www.openresearch.wtf/ai-needs-new-facts-the-value-of-novel-scientific-research/ Last updated: 2026-07-31T12:35:13.000Z At SXSW London, I had the pleasure of seeing DeepMind co-founder Demis Hassabis speak on the future of artificial intelligence. Among many thought-provoking points, two remarks stuck with me. First, he emphasised the importance of understanding the fundamentals. Second, he championed the scientific method as a guiding principle for making meaningful progress in AI. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXesLfvNwRxs3knaOjRSuo5s_togTCymUJiAzH0MOZAg82nl_HrdTb329xEMPreBhgG-b0P4_TveYwzVxFDGyv2Yt_A64U9I8n8Vy-ETEYntJwSsJPqI3CIQ0HS3oDXgBJNIrV4lDw?key=NlOZBbuOqAyfneQK2hXhYg) As someone who works at the intersection of open research and AI, I've found myself returning to a specific question: What kind of content truly matters to AI? I have previously [spoken of an AI powered flywheel effect in research](https://www.openresearch.wtf/the-perpetual-research-cycle-ais-journey-through-data-papers-and-knowledge/), where each cycle follows this pattern: 1. Raw data is processed by AI to generate initial research outputs 2. Knowledge extraction tools mine these outputs for higher-order insights 3. These insights form a new, refined dataset 4. AI processes this refined dataset, generating more precise analyses 5. The cycle continues, with each rotation producing more valuable knowledge I interpret “understanding the fundamentals” as being the base layer, the raw data. Models can mimic almost any writing style and generate endless reams of text. Not all content is created equal. I have witnessed this first hand as self declared “academics” from around the globe use generalist repositories to post non peer-reviewed content written by LLMs which proves their genius. AI systems, particularly large language models, rely on data to learn. Not just more data, but better data; data that reveals new structure in the world. Without novel input, AI models will become better at rephrasing the known, but not at understanding the unknown. At its core, science is a method for producing high-quality, structured novelty - a repeatable process for generating new facts, testing them against reality, and sharing them with the world. Basic research, often funded for its long-term potential rather than short-term applications, is the primary engine of this kind of content. AlphaFold succeeded because it was trained on grounded, empirical data from the [protein databank](https://www.rcsb.org/docs/additional-resources/structure-prediction?ref=openresearch.wtf).If we want AI to continue advancing in a meaningful way that uncovers new knowledge, we need to prioritise access to and support for novel scientific research. That means supporting open science. It means investing in infrastructure that ensures new data and discoveries are FAIR (Findable, Accessible, Interoperable, and Reusable). It means rethinking publication practices to encourage the dissemination of negative results, replication studies, and raw data. And it means funding basic research across the globe. China has significantly increased its investment in basic research, aiming to reduce reliance on foreign technology and achieve self-sufficiency in foundational sciences. China's spending on basic research passed 6% of its total R&D in 2023 and [continues to rise](https://doi.org/10.1016/j.chieco.2024.102281?ref=openresearch.wtf). China seems to be an outlier. Here in the UK for example, basic research through UKRI continues but often faces pressures to demonstrate [short-term economic impact](https://www.ukri.org/publications/research-financial-sustainability-data/research-financial-sustainability-issues-paper/?ref=openresearch.wtf). In the US, the NSF and NIH continue to support basic research, but federal R&D budgets have shifted toward mission-driven, applied research. The [Inflation Reduction Act and CHIPS and Science Act](https://www.commerce.gov/news/blog/2024/08/two-years-later-funding-chips-and-science-act-creating-quality-jobs-growing-local?ref=openresearch.wtf) brought some uplift, but basic science still receives a minority share of total R&D. The geopolitical landscape is hard to predict. But the narrative is that all countries want to compete on the AI stage. Radical abundance only will happen in your country if you have some control of the input to the models. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXcUMcHe78Z45JRQXnN9_tYdQmnb0EPWh9xyXgBmfQtYc1V_Nsf_PL1ZhvRSbDp9YaJps8GvzWAwY541eJqlABnkva9R8H1R2S3BTNLxyzt1gGsIY1P2ez-UNB_0iCYpn8Pmk_u8MQ?key=NlOZBbuOqAyfneQK2hXhYg) At a separate SXSW London session, we also heard from former UK Prime Minister Tony Blair and [The Rt Hon Peter Kyle MP](https://www.gov.uk/government/people/peter-kyle?ref=openresearch.wtf), UK Secretary of State for Science, Innovation and Technology. Peter Kyle’s plans for integrating [AI into the UK government](https://www.gov.uk/government/news/landmark-government-trial-shows-ai-could-save-civil-servants-nearly-2-weeks-a-year?ref=openresearch.wtf) are commendable, and seem to be advancing at a pace that the UK government is not famous for. A comment from Tony Blair however highlighted what is at stake if we don't fund basic research: “It is amazing to me that we are not feeding all of the NHS data into these AI models.” Whether you trust this government with your most sensitive data or not, we will have future governments who may not follow your best interests - in the same way that LLM models are ignoring copyright on academic publications today, they may ignore human ethics when it comes to your medical data. Feeding the NHS into LLMs is not the answer. The easiest way to generate new data for the AI models with a view to advance science and technology is to fund more basic research and insist that the outputs be made open in a FAIR manner. There is so much to gain. The value of novel scientific research has never been higher. ### From Gold to Diamond: Is Equitable Open Access Still a Mirage? URL: https://www.openresearch.wtf/from-gold-to-diamond-is-equitable-open-access-still-a-mirage/ Last updated: 2026-07-31T12:35:40.000Z A few [years ago](https://figshare.com/blog/Open%5FAccess%5Fand%5FData%5FManagement%5FA%5Fwinning%5Fcombination/31?ref=openresearch.wtf), I wrote that Open Access is an inevitability. And in many ways, the data supports that view—at least at a glance. Gold Open Access—the model where authors (or their funders) pay article processing charges to make work freely available—has grown steadily over the past decade. But if you look closely, the growth is slowing. And what’s taking its place isn’t the altruistic, community-powered model many hoped for. Instead, Hybrid Open Access is filling the gap: a model where paywalled journals charge extra to make select articles open. Diamond Open Access—where publishing is free for both authors and readers, supported by institutions rather than APCs—has been making headlines. But here’s the uncomfortable truth: we still don’t have enough data to know whether Diamond OA is actually growing or simply being talked about more. ### A System Stuck in the Middle Digging into Dimensions data layered on top of OpenAlex, a clear pattern emerges. Gold OA’s momentum is slowing. But instead of researchers turning to Green OA (self-archiving in repositories) or Diamond OA, many are choosing Hybrid OA instead. **Why?** Because researchers chase visibility, reputation, and prestige. That often means publishing in high-impact journals—many of which are owned by legacy publishers who now offer hybrid models. These options give authors a way to comply with funder mandates without sacrificing perceived academic clout. Institutions and libraries, meanwhile, are under pressure to show open access progress. Hybrid OA, while expensive, is a politically safe way to do that. It's the administrative equivalent of checking a box—even if it means paying twice. Transformative agreements like Read and Publish deals have only accelerated this trend, redirecting subscription budgets to cover OA fees—effectively normalizing hybrid publishing in many disciplines. ### A New Vision Emerges But while the hybrid tide rises, a quiet revolution is underway. In 2022, a coalition of organizations—Science Europe, cOAlition S, OPERAS, and ANR—launched a bold Action Plan to support Diamond OA. Their goal: to build a truly equitable, community-driven publishing ecosystem, where knowledge is a public good and the costs are shouldered collectively—not by individual researchers. This vision took shape through the DIAMAS project and culminated in the creation of the **Diamond OA Standard (DOAS)**. Think of DOAS as a blueprint: a framework to help Diamond journals measure, improve, and sustain quality. It’s built on seven pillars: - Legal ownership, mission, and governance - Open science practices - Editorial management and research integrity - Technical service efficiency - Visibility and impact - Equity, diversity, inclusion, and multilingualism - Continuous improvement Together, these components aim to professionalize Diamond OA without compromising its values. They send a clear message: if scholarly communication is a public good, then it must be shaped and governed by the scholarly community itself. ### The Missing Link: Measurable Growth Despite this momentum, the numbers tell a more sobering story. Early data from OpenAlex paints the picture that Diamond OA has plateaued in terms of publication volume. Enthusiasm and infrastructure have grown—but the data doesn't reflect this. Why the disconnect? Part of the issue is visibility. At Digital Science, we would love a better way to track Diamond Open Access growth. DOAJ lists 1,369 journals as being ["without fees"](https://doaj.org/search/journals?source=%7B%22query%22%3A%7B%22bool%22%3A%7B%22must%22%3A%5B%7B%22term%22%3A%7B%22bibjson.apc.has%5Fapc%22%3Afalse%7D%7D%2C%7B%22term%22%3A%7B%22bibjson.other%5Fcharges.has%5Fother%5Fcharges%22%3Afalse%7D%7D%5D%7D%7D%2C%22track%5Ftotal%5Fhits%22%3Atrue%7D), but there seem to be many edge cases resulting in a landscape that isn't black and white. It is also true that many Diamond journals operate with limited marketing, uncertain technical infrastructure, and fragmented funding. And despite the ideals, many researchers still don’t see them as viable options for career advancement. ### What Comes Next? There are reasons for optimism. The [DIAMAS project](https://diamasproject.eu/?ref=openresearch.wtf) isn’t just advocating for Diamond OA—it’s building the scaffolding. Its service portal offers templates, best practices, and technical guidance to help journals align with DOAS standards and professionalize their operations. The [European Diamond Capacity Hub](https://operas.hypotheses.org/7833?ref=openresearch.wtf) (EDCH) serves as a coordination center for Diamond OA stakeholders in Europe. Launched alongside the EDCH, the [ALMASI Project](https://operas.hypotheses.org/7833?ref=openresearch.wtf) focuses on understanding non-profit OA publishing in Africa, Latin America, and Europe. The pieces are falling into place for equitable open access solutions. What’s needed is adoption and quantification. Funders and Institutions should include equitable OA in their promotion criteria. Researchers must see it as a credible home for their work. If anyone knows of faster ways to get clean data for Diamond, please reach out. The path forward exists. But like any path, it only becomes clear by walking it. If you are working in this space and have this data - we would love to disseminate it through our tools at Digital Science. ### The Perpetual Research Cycle: AI's Journey Through Data, Papers, and Knowledge URL: https://www.openresearch.wtf/the-perpetual-research-cycle-ais-journey-through-data-papers-and-knowledge/ Last updated: 2026-07-31T10:00:40.000Z Academics hypothesize, generate data, make sense of it and then communicate it. If AI can help to generate, mine, and refine knowledge faster than human researchers, what does the future of academia look like? The answer lies not in replacing human intellect but in enhancing it, creating a collaborative synergy between AI and human researchers that will define the next era of scientific progress. I’ve been playing around with chatGPT, Google Gemini and Claude.ai to see how well they all do at creating academic papers from datasets. AI can also serve as a tool to aid humans in data extraction from many papers. Consider a scenario where AI synthesizes information from hundreds of studies to create a refined dataset. That dataset then feeds back into the system, sparking new research papers. This cycle—dataset to paper, paper to knowledge extraction, knowledge to new datasets—propels an accelerating loop of discovery. Instead of a linear research pipeline, AI enables a continuous, self-improving knowledge ecosystem. ## From data to papers I looked for interesting datasets on Figshare. The criteria was a) that I knew they would be re-usable as they had been cited several times. And b) the files were relatively small (<100MB) so as not to hit the limits of the common AI tools. This one fit the bill: Rivers, American (2019). American Rivers Dam Removal Database. figshare. Dataset. [https://doi.org/10.6084/m9.figshare.5234068.v12](https://doi.org/10.6084/m9.figshare.5234068.v12?ref=openresearch.wtf) ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXf2S6BvKoV8GPSMlXRPhAV0pN0hrWKEIIPHgpLAZUoub91KYUt7gpXAdh1x1myZ9eps8GzflYK56lSIL8pOdDov_W_WUFwH3dcKZaBNUotNYEqWPSOVtgiGYve4iwtA9VzHT_AG?key=6Rm7jClnBZtMs1woLj067Oep) From there I asked Claude 3.7 Sonnet “Based on the attached files, can you create a full length academic paper with an abstract, methods results, discussion and references”. Followed by “Can you convert the whole paper to latex so I can copy and paste it into Overleaf?” The [resulting paper](https://www.overleaf.com/read/vcyjfrfcqkgb?ref=openresearch.wtf#315ab4) needs a little tweaking in the layout of the results and graphs, but other than that, has done a great job. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXfyz4yqwYbyy9USL45xtFrjxsEXWbxbKhwMUpq8cz_ejUW9PF4sTAfJ4GjWc5UAJW4EHbEB-_Lgrz618xPvYCPh6fEinPWujWQCWtKy5xzGw_kDipxIH1BjwYJjpdga38CsMtdQ3g?key=6Rm7jClnBZtMs1woLj067Oep) ### Papers to new data/knowledge A single paper is just the beginning. The real challenge is synthesizing knowledge from the ever-growing volume of research. This is where specialized knowledge extraction tools become crucial. How do we effectively mine this knowledge?This is where [ReadCube](https://www.readcube.com/en/solution/readcube-slr/?ref=openresearch.wtf) shines. ReadCube helps researchers manage and discover scholarly literature, but its real power lies in its knowledge extraction capabilities. Imagine ReadCube as a powerful filter, sifting through countless pages to extract the nuggets of wisdom. Tools like ReadCube can then analyze vast collections of papers, uncovering patterns and relationships that human researchers might miss. This process involves: - Text and citation mining: AI can analyze papers to identify emerging trends, inconsistencies, or knowledge gaps. - Automatic synthesis: AI can compare findings across thousands of studies, synthesizing insights into new, high-level conclusions. - Hypothesis generation: By recognizing correlations between disparate research areas, AI can propose new research directions. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXcptw8jQdkTCYHI-iT4T81lTkpvjYZ-kWKJ_l2upZck7rqhnIetDISmy0uu3K8x6tNhIgqdV9a0xwJiPUMdoEaBBAhYnl3j6v7oygZc4i-wrko3gJrU2AAejMWJauIeasvu3-mFLQ?key=6Rm7jClnBZtMs1woLj067Oep) ## The Flywheel Effect: How the Cycle Accelerates The true magic happens when this extracted knowledge becomes the input for the next iteration. Each cycle follows this pattern: 1. Raw data is processed by AI to generate initial research outputs 2. Knowledge extraction tools mine these outputs for higher-order insights 3. These insights form a new, refined dataset 4. AI processes this refined dataset, generating more precise analyses 5. The cycle continues, with each rotation producing more valuable knowledge With each turn of this flywheel, the insights become more refined, more interconnected, and more actionable. The initial analyses might focus on direct correlations in the data, while later iterations can explore complex causal relationships, predict future trends, or suggest optimal intervention strategies. This AI-driven, data-to-knowledge cycle represents a paradigm shift in research. Imagine the possibilities in fields like medicine, climate science, and economics. We’re moving towards a future where AI and human researchers work in synergy, pushing the boundaries of discovery. Rather than replacing researchers, AI acts as a force multiplier, enabling deeper faster knowledge generation. ### How is every country/funder/ institution doing at data sharing? URL: https://www.openresearch.wtf/how-is-every-country-funder-institution-doing-at-data-sharing/ Last updated: 2026-07-31T12:35:44.000Z The '[Make Data Count Data Citation Corpus](https://makedatacount.org/find-a-tool/?ref=openresearch.wtf)' looks for both DataCite DOIs and [Accession numbers](https://www.ncbi.nlm.nih.gov/books/NBK470040/?ref=openresearch.wtf#:~:text=The%20accession%20number%20is%20a,the%20accession%20number%2C%20an%20Accession.) in the published literature. It basically tells you that "this publication DOI" mentions "this DataCite DOI or Accession number". By taking this unstructured data and combining it with the [Dimensions.ai](https://www.dimensions.ai/?ref=openresearch.wtf) corpus in [Google big query](https://docs.dimensions.ai/bigquery/?ref=openresearch.wtf), we can see what percentage of papers by different Countries, Funders, and Institutions link to a DataCite DOI or Accession number. You can test it out yourself below: The data behind the app can be found here: Hahnel, Mark (2024). Data behind State of Open Data 2024 Special Report: Bridging policy and practice in data sharing - Country, Funder and Affiliation Datasets. figshare. Dataset. [https://doi.org/10.6084/m9.figshare.27900828.v1](https://doi.org/10.6084/m9.figshare.27900828.v1?ref=openresearch.wtf) I've included some screenshots of what kind of comparisons you can do: ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/02/Screenshot-2025-02-27-at-11.25.14.png) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/02/Screenshot-2025-02-27-at-11.28.00.png) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/02/Screenshot-2025-02-27-at-11.29.11.png) This shows the power of combining new datasets with well structured databases such as Dimensions in order to obtain new knowledge. ### The Role of Foundations and Technology Companies in Fuelling Optimism in Academic Research URL: https://www.openresearch.wtf/the-role-of-foundations-and-technology-companies-in-fuelling-optimism-in-academic-research/ Last updated: 2026-07-31T12:35:58.000Z The landscape of grant funding in North America has recently, rapidly, evolved, leading to a sense of nervousness and even pessimism within academia. However, remarkable technological advancements emerging from foundations and companies are counteracting this trend and fostering a renewed sense of optimism. In the last week, we have seen the [unveiling of Evo-2](https://arcinstitute.org/news/blog/evo2?ref=openresearch.wtf), the largest AI model for biology, developed by the Arc Institute and the launch of [Google’s AI co-scientist](https://blog.google/feed/google-research-ai-co-scientist/?ref=openresearch.wtf). ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXdQDJgeTLRT5dJ8XmzYxM2DBVqBWrzcyITNc4LWcVb07kChxMES1kVF6DjM3MZnfbZQYbDaXaovelqJJoLalZ71YbJmaNPlN0kO54Ns9qj8zIyAVtzZgmAuRNAVjITs_sh5xaWGiQ?key=aJ5rmjw_-oisUxbxg7eNYbT8) Evo-2, a collaborative effort between researchers at the Arc Institute, Stanford University, and NVIDIA, is accessible to scientists through web interfaces, with its[ software code, data, and parameters available for download](https://github.com/arcinstitute/evo2?ref=openresearch.wtf). This model, trained on 128,000 genomes, spans the entire tree of life, from humans to single-celled organisms. The AI co-scientist, powered by Gemini 2.0, acts as a virtual research partner, aiding scientists in formulating hypotheses and research proposals, and accelerating the pace of scientific and biomedical discoveries. We have seen [versions of this before](https://www.nature.com/articles/s41586-023-06792-0?ref=openresearch.wtf), and are working on parallel themes for our clients at Digital Science, but it is encouraging to see that there is excitement for this at the [very top of Google and Alphabet](https://x.com/sundarpichai/status/1892254274895184244?ref=openresearch.wtf). This wave of technological progress shows no signs of abating. The Astera Institute, a relatively new organization, aims to “accelerate science and technology for the benefit of humanity.” Their unique approach involves supporting entrepreneurially minded scientists in developing open science tools and technologies. Currently, they focus on [creating predictive models of microbial behavior](https://astera.org/request-for-information/?ref=openresearch.wtf). Seemingly reverse engineering outcomes, by asking researchers to help share their academic data in a way that can move the needle like Alphafold. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXfEP5Tgv1_0NpYaOhrkgM6n89_ceus2YdcEHTbwWbBX2WyGpPznn2_p1m3NVqeCFnG8rfLGyohCxoQ8medmWgVadttdhfIEjSqxMnfrzZQ9R1kYCZKWqNh5RzwaE9jsI097cBD5qw?key=aJ5rmjw_-oisUxbxg7eNYbT8) At Digital Science, we share this optimism and are actively contributing to these advancements. Digital Science has a commitment to innovation, hand-in-hand with the responsible development of AI tools. With tools like [the Dimensions AI Assistant](https://www.dimensions.ai/blog/powering-research-with-dimensions-ai-assistant/?ref=openresearch.wtf), which provides users with contextualized synopses through extractive and abstractive summarization. [Symplectic Elements](https://www.symplectic.co.uk/theelementsplatform/?ref=openresearch.wtf) now offers AI-powered abstracts and summaries, making research more discoverable and accessible while maintaining institutional control. We’re also hard at work on novel advances including: - Auto-taxonomy: Employing AI/ML to automate classification, combining recent AI/ML techniques with traditional bibliometric methods to identify training data and train ML classifiers. - AI Training: Preparing data for AI learning, such as processing complex scientific language from molecular biology. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXc9ketZNqwrdM8Euji7xHH7R3eA0GmVK70W8bCjL_SkC54vFgHjQ7Bp_7i1DYb1vAiUwqUDpq7CWZ54mNsMQMC7LS1MarsctZnSywdss8x9sQDdqCBPgfWrROCVTdoXO6UtSZ4?key=aJ5rmjw_-oisUxbxg7eNYbT8) - Horizon Scanning: Utilizing AI to identify emerging research trends. Our approach involves generating a 5,000 K-Means clustering model using publications from Dimensions.ai - Literature Review: Leveraging generative AI to assist researchers in managing, organizing, and summarizing extensive scientific literature. - Data Collection and Generation: Automating data collection from diverse sources and generating synthetic data to enhance existing datasets. Applying AI to workflows gives researchers more time to focus on life-changing discoveries. We help researchers focus on higher value work, making them more productive with advanced AI tools.Our efforts are focused on addressing existing challenges in building upon previous research and AI-powered drug design. This includes creating [high-quality, diverse, and interoperable datasets](https://data.dimensions.ai/resource/Dimensions-Knowledge-Graph?ref=openresearch.wtf) that are essential for training AI models. We’re doing this in many ways, through teams like [metaphacts](https://metaphacts.com/?ref=openresearch.wtf), [Dimensions](https://www.dimensions.ai/?ref=openresearch.wtf) and [Readcube](https://www.readcube.com/en/solution/readcube-slr/?ref=openresearch.wtf). ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXcHw7l92PHOScYIbCXFO80S-mzgt3LpLAXGBu39gAVlaqOmVNz4xcelcaLXwtjuutigj1vWiZMsF2uGeFovFhCumaTa6X03Zv8MmN_RCP2J-nLC_kIxRuS2LTlmxBeckI9sIkpB?key=aJ5rmjw_-oisUxbxg7eNYbT8) Access to the right data and literature is crucial. At Digital Science, our approach has been to augment existing workflows with AI while ensuring end-to-end transparency – so the user can check and see what is happening at every stage. [Readcube SLR](https://www.readcube.com/en/solution/readcube-slr/?ref=openresearch.wtf) is making significant strides in this area by: - Automating literature collection, screening, and assessment using AI and intelligent workflows. - Leveraging AI to refine searches and accelerate screening. - Providing customisable templates. - Offering AI-driven pre-filled suggestions with references. - Implementing a QA process that combines human and AI validation. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXczkIg4TH-DeUv3hLMhJv59SVetj8A2g_y84RiG7s8T2ZRuVExfc0LKVvz5pug5ikyegTKknmsAi8Azke3yP4SIFLMW964M3Vw2ydPb-M4lc-Pa0syFF1JGovQ_VvS1rc6muxjpjQ?key=aJ5rmjw_-oisUxbxg7eNYbT8) In times of funding challenges, technology companies play a vital role in supporting and advancing academic research. These organisations can accelerate scientific discovery, improve data management, and foster a more optimistic and productive research environment. I don't know what the next 6 months of research funding looks like, but I couldn't be more optimistic for the future of academia and the fruits it yields. ### DataCite DOI prevalence in the published literature by country URL: https://www.openresearch.wtf/datacite-doi-prevalence-in-the-published-literature-by-country/ Last updated: 2026-07-31T12:47:07.000Z In a previous post, I dug a little deeper into the findings of the State of Open Data to try and understand whether there we're any regional trends in data sharing and re-use using[ the Make Data Count Data Citation Corpus](https://makedatacount.org/find-a-tool/?ref=openresearch.wtf). [Are we ready to start rewarding researchers for Open data?The State of Open Data 2024: Special Report Bridging policy and practice in data sharing has gone live today. In it, we explored global trends in data sharing practices, focusing on bridging the gap between policy and practice. For the first time, the report incorporates not only survey data but![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/icon/favicon-2.ico)OpenResearch.wtfMark Hahnel![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/thumbnail/Screenshot-2024-12-02-at-15.43.26.png)](https://www.openresearch.wtf/are-we-ready-to-start-rewarding-researchers-for-open-data/) The 'Make Data Count Data Citation Corpus' looks for both DataCite DOIs and [Accession numbers](https://www.ncbi.nlm.nih.gov/books/NBK470040/?ref=openresearch.wtf#:~:text=The%20accession%20number%20is%20a,the%20accession%20number%2C%20an%20Accession.) in the published literature. As there are variances in the 'types' of datasets that use DataCite DOIs or Accession numbers, I thought it would be a good idea to try and simplify this and normalise the data further. [Dimensions.ai](https://www.dimensions.ai/?ref=openresearch.wtf) and their [Google big query datasets](https://docs.dimensions.ai/bigquery/?ref=openresearch.wtf) can look at just the number of links to DataCite DOIs from anywhere in the full text of the published literature. As bibliometric data can sometimes take a while to filter through, I decided to look at 2022 for this analysis. To do this I looked at Country based data linkage count.xlsx in the Data behind State of Open Data 2024 Special Report: Bridging policy and practice in data sharing [Data behind State of Open Data 2024 Special Report: Bridging policy and practice in data sharing - Country, Funder and Affiliation DatasetsThis Dataset contains 3 datasets behind graphs generated in the “State of Open Data 2024 Special Report: Bridging policy and practice in data sharing” The datasets include counts and percentages for papers that link to datasets filtered by Country, Funder and Affiliation DatasetsThe datasets were generated by combining the DataCite Data Citation Corpus (https://corpus.datacite.org/dashboard) with Dimensions (https://www.dimensions.ai/) in Google big query.![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/icon/favicon-196.911535d4.png)figshare![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/thumbnail/defaultLogo.30adffde.png)](https://figshare.com/articles/dataset/Data%5Fbehind%5FState%5Fof%5FOpen%5FData%5F2024%5FSpecial%5FReport%5FBridging%5Fpolicy%5Fand%5Fpractice%5Fin%5Fdata%5Fsharing%5F-%5FCountry%5FFunder%5Fand%5FAffiliation%5FDatasets/27900828?file=50783595&ref=openresearch.wtf) By creating Chloropleth maps, we can see where the percentage of papers with links to DataCite DOIs is highest. We can see some obvious outliers here, and so can filter further. Below I have generated maps for **"The percentage of papers that link to DataCite DOIs in counries that publish more than 1000 papers per year"** and "**The percentage of papers that link to DataCite DOIs in counries that publish more than 10,000 papers per year**" This last graph highlights where work needs to be done when it comes to data publishing and reuse in academia. If you are a funder in a country that publishes large volumes of peer reviewed academic publications, that are not encouraging data publication, you are behind. You are not realising the full value of the research you fund. The research you pay for will not be as exposed to the AI models generating new findings based on the research that has already been paid for. Do more. Or in some cases, just do something. ### Types of AI Agents and Their Applications in Academic Research URL: https://www.openresearch.wtf/types-of-ai-agents-and-their-applications-in-acaresearch/ Last updated: 2026-07-31T12:35:59.000Z AI agents have the potential to revolutionise how researchers approach complex problems and manage information in academic research. These intelligent systems can automate tasks, analyse vast datasets, and even generate hypotheses, significantly accelerating the pace of discovery and innovation. AI agents can be categorised into different types, based on their capabilities and functionalities. There is overlap in the different types of agents and how they can assist you in your research, but from a top level, here are some types to think about. ### **1\. Model-based Agents** Model-based agents utilise a model of the world to make decisions and take actions. They are particularly useful in simulating complex systems, predicting outcomes, and optimising experiments. For example, a model-based agent could simulate the spread of a disease in a population or predict the impact of a new policy on the environment. In academic research, model-based agents can be applied to various domains and processes: - **Hypothesis Generation:** Model-based agents can analyse existing literature and data to generate novel hypotheses. For example, [FieldSHIFT](https://www.sciencedirect.com/org/science/article/pii/S2635098X2400024X?ref=openresearch.wtf), an in-context learning framework, uses a large language model to facilitate candidate scientific research from existing published studies. - **Literature Review:** Agents like those used by [Readcube SLR](https://www.readcube.com/en/solution/readcube-slr/?ref=openresearch.wtf) can automate the process of reviewing academic literature, extracting key insights from research papers with ease and precision, as well as continuously monitoring the literature for eg. drug mentions. - **Data Analysis:** [LAMBDA](https://arxiv.org/abs/2407.17535?ref=openresearch.wtf), an open-source multi-agent data analysis system, leverages large models to address data analysis challenges in complex data-driven applications. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2025/01/AI-suggested.webp) ### **2\. Learning Agents** Learning agents, unlike model-based agents, learn from their experiences and adapt their behaviour accordingly. They excel in analysing data, identifying patterns, and generating new hypotheses. In academic research, learning agents can be applied to: - **Literature Review:** AI-powered tools like [Readcube AI Assistant ](https://www.readcube.com/en/ai-assistant/?ref=openresearch.wtf)aid researchers in literature research, discovering and organising academic papers efficiently. These agents can identify top papers in the field, even without perfect keyword matching, and organise them into an easy-to-use table. - **Data Analysis:** Learning agents can be used to analyse social behaviour, preferences, and demographic data to understand and why humans act the way we do. - **Experimentation:** Learning agents can be used to analyse data from experiments, identify patterns, and suggest new experimental designs. For example, they can be used to analyse data from clinical trials to identify factors that influence treatment outcomes. ### **3\. Hierarchical Agents** Hierarchical agents are organised in a hierarchical structure, with different agents responsible for different tasks. This structure allows for efficient management of complex tasks and coordination of large teams. In academic research, hierarchical agents can be applied to: - **Automated post-publication checks errors and discrepancies**: For example, [YesNoError](https://yesnoerror.com/?ref=openresearch.wtf) combines model-based techniques with a hierarchical, multi-agent framework to effectively detect errors in scientific papers - **Publication:** Hierarchical agents can assist in the publication process by automating tasks such as formatting, proofreading, and submission. Something that [Overleaf and Writefull](https://www.overleaf.com/learn/how-to/Writefull%5Fintegration?ref=openresearch.wtf) can assist with. - **Experimentation:** In complex experiments involving multiple steps and variables, hierarchical agents can coordinate different aspects of the experiment, ensuring efficient execution and data collection. The future of AI agents in academic research is promising. Ongoing research focuses on enhancing their reasoning capabilities, improving their adaptability, and ensuring their ethical and responsible use. As AI agents continue to evolve, the speed at which we move further, faster in academic research will accelerate. ### Have we all forgotten about SciHub? URL: https://www.openresearch.wtf/have-we-all-forgotten-about-scihub/ Last updated: 2026-07-31T12:35:42.000Z A Preprint was published on Preprints.org at the end of November by SciHub founder Alexandra Elbakyan. [From Black Open Access to Open Access of Color: Accepting the Diversity of Approaches towards Free Science - \[v2\]![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/icon/favicon-1.ico)\[v2\]![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/thumbnail/preprints.C_f_-Qxj-1.png)](https://www.preprints.org/manuscript/202409.0197/v2?ref=openresearch.wtf) It describes an interesting idea—using more OA colors to replace illegal OA, referred to as Black OA in this paper. I found this intriguing, as I had never encountered the term Black OA before. Elbakyan proposes a color-coded system (e.g., red for automated tools like Sci-Hub, blue for academic social networks) to better represent the diversity of OA practices. ### The Development and Segmentation of Black OA: - **Shadow Libraries**: Early projects like Library Genesis and AvaxHome functioned as repositories for user-uploaded books, eventually evolving to include academic journals. - **Literature-Sharing Communities**: Online forums and later social media hashtags like #icanhazpdf enabled informal sharing of academic content. - **Automatic Paywall Circumvention**: Sci-Hub emerged as a transformative tool, providing automated access to paywalled research and accumulating a vast repository of academic papers. - **Academic Social Networks**: Platforms like ResearchGate and Academia.edu, while distinct, have been included in the Black OA domain due to their facilitation of unauthorised content sharing. What truly struck me was the graphs on the continued growth and massive usage of Sci-hub: \>1 billion downloads by more than 83 million unique users ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/12/Screenshot-2024-12-10-at-11.14.57.png) Year on year growth with nearly 150m downloads per month ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/12/Screenshot-2024-12-10-at-11.15.06.png) Sci-Hub alone provides access to over 90% of paywalled research ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/12/Screenshot-2024-12-10-at-11.14.46.png) In a world where publishers are making big deals to allow LLMs to access and process their research papers, it is surprising that we aren't seeing illegal AI approaches to knowledge. Robots can access illegal content as easily as the 83 million humans using Sci-Hub. Will we see new knowledge generated by LLMs built on Red OA, disregarding the legal ramifications? ### Are we ready to start rewarding researchers for Open data? URL: https://www.openresearch.wtf/are-we-ready-to-start-rewarding-researchers-for-open-data/ Last updated: 2026-07-31T10:00:05.000Z [*The State of Open Data 2024: Special Report Bridging policy and practice in data sharing*](https://doi.org/10.6084/m9.figshare.27337476.v1?ref=openresearch.wtf) has gone live today. In it, we explored global trends in data sharing practices, focusing on bridging the gap between policy and practice. For the first time, the report incorporates not only survey data but also analyses of actual data-sharing behaviours, leveraging sources like Dimensions, the Springer Nature Data Availability Statements (DAS), and the Wellcome Funded, Make Data Count Data Citation Corpus. By combining these datasets, the report provides insights into how researchers actively share data, the impact of mandates, and the effectiveness of policies in fostering open data practices. This evolution from understanding researcher attitudes to tracking their actions aims to address barriers to open data, such as lack of incentives and resource disparities. Key findings highlight significant regional, institutional, and funder-level differences in data sharing. Here are some of the [key findings from the report](https://doi.org/10.6084/m9.figshare.27337476.v1?ref=openresearch.wtf), and some extra analysis that I have found interesting whilst playing with the Wellcome Funded, Make Data Count Data Citation Corpus data, combined with Dimensions disambiguation. When it comes to the percentage of papers linking to datasets, Sweden are winning at scale. Scandinavia in general does well. If we look at the same data but reduce the threshold from countries that produce >50,000 publications per year to >1000 publications per year, Africa are doing better. I'm excited to see what effect the NIH data publishing mandate has on their already impressive increase in links to datasets More important that actual counts, is what percentage of papers are linking to a dataset. In country funder differences are also important when trying to reverse engineer what is working. Of course, subject specific differences will be playing a role here, with different funders specialising in different areas. The same approach can be applied to Institutions located in teh same country. Librarians at these organisations can follow best practices of those with better data sharing rates. The hugely successful data sharing practices at the Francis Crick Institute not only re-emphasise that subject focus may be bringing in bias, but also that legacy institutions come with a lot of legacy baggage, and cannot focus as much budget or people on new initiatives like data sharing. In order to drive societal change in academia, we need to use both carrots and sticks. To succeed, we need to move the process through the following steps: \- Policy \- Mandate \- Compliance \- Measurement In the nine years that the State of Open Data has been running, we have seen policies forming and becoming mandates. This has driven a lot of the compliance we see today. Researchers are sharing because they are told they have to as a requirement of funding. The survey continues to tell us that the main reason that is stopping them engaging with open data publishing in a more serious manner is a lack of credit for their open data. Researchers cannot get credit if there is not a way to consistently measure data metrics across platforms and repositories. The CZI DCC allows us to do this. [The NIH Data Sharing Index (S-Index) Challenge](https://www.challenge.gov/?challenge=nih-data-sharing-index-s-index-challenge&tab=overview&ref=openresearch.wtf) seeks innovative approaches to quantify and evaluate data sharing practices by biomedical researchers. The challenge is aimed at developing an “S-Index” to measure the extent and effectiveness of data-sharing, encouraging transparency and accessibility in research. Submissions are judged based on originality, feasibility, impact, and scalability of the proposed metrics. The goal is to foster more widespread and high-quality sharing of scientific data. [The DataWorks! Prize](https://www.herox.com/dataworks?ref=openresearch.wtf) organized by the Federation of American Societies for Experimental Biology (FASEB) and NIH, is a challenge that encourages researchers to propose impactful secondary data analysis projects using existing biomedical data. Being able to measure the impact of researchers sharing their data and ultimately reward them for doing so means that we are on the verge of having both carrots and sticks to enable open data sharing. The methods behind these figures and the raw data itself, can be found on Figshare at: Hahnel, Mark; Smith, Graham; Campbell, Ann (2024). The State of Open Data 2024: Special Report Bridging policy and practice in data sharing. Digital Science. Report. [https://doi.org/10.6084/m9.figshare.27337476.v1](https://doi.org/10.6084/m9.figshare.27337476.v1?ref=openresearch.wtf) ### Innovating Academia From The Outside - The Web3 Solution to Academia’s Peer Review Problem URL: https://www.openresearch.wtf/innovating-academia-from-the-outside-the-web3-solution-to-academias-peer-review-problem/ Last updated: 2026-07-31T12:35:45.000Z [ResearchHub](https://www.researchhub.com/?ref=openresearch.wtf) is a platform for open science that has been building an impressive tech stack since 2020, where users can review, publish, and collaborate on scientific research. The primary novelty that seems to interest and terrify researchers in equal measure is the fact that it is a crypto based project, funded by Coinbase founder Brian Armstrong. I fall into the category of being enthusiastic about the potential for blockchain, so have been following their progress closely and have recently been super impressed by just the sheer levels of innovation on the platform. Interestingly, I first spoke to the now COO Patrick Joyce in 2018, when he had founded a company called co-lab (later rebranded to knowledgr) and caught up with him and his colleague Jeff Koury recently to discuss some of the new features and aims for the project in the short and long term. ResearchHub has similarities to the Open Science Framework (OSF). Both platforms encourage researchers to share their work openly, facilitating collaboration and reproducibility. Both platforms provide tools for sharing data, preprints, and research resources, which seems like an obvious next step in research dissemination, but the incentives for researchers are often a stumbling block, in a slow to evolve, publish or perish, metrics driven career path. ResearchHub’s token-based system encourages researchers to contribute by rewarding uploads, reviews, and discussions. This token economy model presents a novel approach to sustaining community engagement in a field where recognition and reward have often been limited to formal publications and citations. ### Paid peer review. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXflvO6QoqY69HZ1ciPkCwr5l8guZSwCSkdCP2FhtHuzwvipAlBwgG0--z4GrqEmPgKcn6WI-t_M3Iy9gs8cIjqHi1LJEFUuZFR09J998R-Olli-jyrorDzYdS-6I5exy_aseMeihg?key=yByXgLvsbu2jr4roytPtow) One of the most innovative aspects of ResearchHub is its approach to peer review. Traditionally, peer review is an unpaid and under-appreciated part of academic research. ResearchHub changes this by offering ResearchCoin (RSC) rewards for high-quality peer reviews, ensuring that peer review is timely, constructive, and valued. If you have ever caught me having a beer after a conference you will know that I believe peer review is the most fundamentally broken part of academic publishing today. As the data above from [Dimensions](https://dimensions.ai/?ref=openresearch.wtf) shows, there is ever increasing demand for peer reviewers, without an ever increasing supply of researchers. With gold open access driving a "solve for X model", where X is the number of APCs you can publish in a high quality manner, we end up with a huge "[strain on scientific publishing](https://arxiv.org/abs/2309.15884?ref=openresearch.wtf)". ![](https://lh7-qw.googleusercontent.com/slidesz/AGV_vUck_7egzj7Znhk8GFv_8Xw45SFzTahnvoUZbAw0ripsPMqYshQEZYHbLet3JvxRiY8ZKZ93bS-LAgua-x69WXNgdB8E5E7RkWTu_nkUU2l091wCi37yimm6cj1SZ_dkN2os4EEorUdpgV7DO2n8ZJtx1g7p7Ji_=s2048?key=oLhUt_I5kynUGuhR7UN2vg) Other platforms have experimented with alternative models to improve peer review. [Publons](https://publons.com/?ref=openresearch.wtf), for instance, allows researchers to showcase their review contributions, providing non-monetary recognition. [eLife](https://elifesciences.org/?ref=openresearch.wtf) has pioneered transparent peer review, where reviewer comments are made open and visible. MDPI Reviewers who provide timely and substantial comments will receive a [discount voucher code](https://www.mdpi.com/journal/catalysts/apc?ref=openresearch.wtf#:~:text=Reviewers%20who%20provide%20timely%20and,redeemed%20against%20the%20same%20article.) entitling them to an APC reduction. ResearchHub combines these concepts. It seems to be working... > Proof that incentivized peer review is changing the game. [pic.twitter.com/xhOFdd2CqS](https://t.co/xhOFdd2CqS?ref=openresearch.wtf) > > — ResearchHub Foundation (@ResearchHubF) [October 25, 2024](https://twitter.com/ResearchHubF/status/1849845018435125549?ref%5Fsrc=twsrc%5Etfw&ref=openresearch.wtf) ResearchCoin tokenomics have been designed based on token engineering principles centred around the concept of the[ Web3 Sustainability Loop](https://blog.oceanprotocol.com/the-web3-sustainability-loop-b2a4097a36e?ref=openresearch.wtf). This implies that the primary focus has always been on growing the ecosystem while ensuring its long-term self-sustainability. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXf_A0sa9bXqG9Vzw5ToBKRYWMuxcF0mkgaEqACDHu93Hk4-a5UkxC1o2Fo-_3aMUAsoe_H-PaSM8A7kKPS2lmzd9Y2TjBwdoZ-gkhzBCTCiZ7K6I1OXsMFBSxOH_vsOQG8QKbL2IzCTryDcwIJmkD0tY3Cq?key=yByXgLvsbu2jr4roytPtow) How this works in the ResearchHub model is played out here. My original concern was that when looking at the green inflows and red outflows, there are not enough green inflows to sustain the constant rewards and payments. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXc79kPH34XCeHfHNThK7cKtn7LTWWyLfRfmW2KPVDFFEw2Iu0p7Kl1Wa-udRhpfbKT92_kf6oCZ8MZj0lZJQaf6vN6tgBJqK7lS_ENb8QBx1exU1CIVvNFU7zvaV4Qb3lfvMboKBHxE58phbmOt7EP2bQY?key=yByXgLvsbu2jr4roytPtow) This may be about to change with their move into Open Access Publishing and the [launch of their first journal](https://www.researchhub.com/researchhub-journal?ref=openresearch.wtf). The $1000 fee with a promise of fast turn around in theory solves the problem of "where is the monetary inflow coming from". ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXfZfUc1s0yv6rCSmXpYW48D_7eAnxxH5_KpxXS0_xxa7Y60M88Bx5tLcw-r-ZTeXhu2-Y8H2Qh8dJnIFFcznaSWxIXMZVoMKaNvs917gva2i8Tys4WNl0EPubMzIvSLZR1aWPI8mg?key=yByXgLvsbu2jr4roytPtow) While token incentives drive engagement, they may also risk attracting contributions motivated by rewards rather than quality. Addressing this tension will be key to ResearchHub’s long-term credibility. Like any open platform, ResearchHub faces the risk of misuse, such as self-promotion or posting lower-quality content. Ensuring content moderation while preserving openness is an ongoing challenge. There are a few more innovative touches to the platform that subtly reflect how slow innovation in the sector has been from new start-ups in the last 5 years. I feel there was a wave of new academic technology moving the needle 15 years ago, and we are on the crest of a second wave now. Their integration with an ID verification service is something that could help prevent bad behaviour that we see in the form of citation rings. Incentives for open data is something that is coming at exactly the right time. The platform pulls from Open Alex in the backend and encourages uploading open access papers in a way that [Researchgate](https://researchgate.com/?ref=openresearch.wtf) and [Academia.edu](https://academia.edu/?ref=openresearch.wtf) have done in the past - but again, the incentive here is financial. ![](https://lh7-qw.googleusercontent.com/docsz/AD_4nXdvR0ouGvbj3XNcQW8KFn1i7VTveyWrNsEk8EoH4iXT-DuSCLdxNwM9TB33XDS-RiA57XO9NF4sVWWEM9nV281lvkOKD1y_R8daqZwznXMMgNhpV3n8T07teERkpxW3u5QR4m0uVo0JCPCtXOu_NS1hXyAH?key=yByXgLvsbu2jr4roytPtow) In speaking with Patrick, he shared that they were looking to really innovate in the way research is funded as a priority. I remember [experiment.com](https://experiment.com/?ref=openresearch.wtf) struggling to have wider reach than the friends and family funding of research that seemed to max out under $20,000\. By engaging directly with the funders themselves, ResearchHub could use their tech stack to disrupt the dissemination of funding and add more fuel to their internal incentives economy. I’m really excited to see a real world attempt to do paid peer review well. It is as close to paying researchers cash as I have seen and whilst it is early, engagement so far seems encouraging. A barrier to entry right now is explaining tokenomics to normie researchers. Luckily for them, a whole web3 industry is moving to educate and solve account abstraction so that researchers wont ever need to understand the rails that is driving them to behave the way they do. Their [SciCon event](https://scicon.researchhub.foundation/?ref=openresearch.wtf) takes place next week, and I'll be dialling in through their innovative token-gated livestream. ### The Data Citation Corpus - tracking NIH funded open academic data URL: https://www.openresearch.wtf/the-data-citation-corpus-tracking-nih-funded-open-academic-data/ Last updated: 2026-07-31T12:35:56.000Z In 2023, the [Wellcome Trust awarded funds](https://blog.datacite.org/data-citation-corpus-announcement-2023/?ref=openresearch.wtf) to build an open Data Citation Corpus to dramatically transform the data citation landscape. Through this award, DataCite has partnered with Chan Zuckerberg Initiative, EMBL-EBI, and other organizations that identify and assert data citations. [Data Citation Corpus – Make Data Count![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Make Data Count![](https://makedatacount.org/wp-content/uploads/2024/01/Screenshot-2024-01-15-at-15.59.18.png)](https://makedatacount.org/data-citation/?ref=openresearch.wtf) We at Digital Science have been looking at the Data Citation Corpus, to dig deeper into data citation counts. The first release is based on a seed file that includes data citations from the following sources: - Data citations from DataCite and Crossref DOI metadata, via [Event Data](https://support.datacite.org/docs/eventdata-guide?ref=openresearch.wtf). - Data citations from the CZI Science Knowledge Graph, identified via a Named Entity Recognition model algorithm that searches for mentions to datasets in the full text of journal articles and preprints in [Europe PMC](https://europepmc.org/?ref=openresearch.wtf). So we are basically looking at papers that have a link to a DataCite DOI or accession number. By combining this dataset with Dimensions.ai data in Google Big Query, we we're able to add more dimensions to the dataset (pardon the pun), such as funder or institution. The Data Citation Corpus only gave us about 70% of the paper links that were resolvable DOIs. This should improve over time. This allows us to track how well things like the NIH open data policy is encouraging linking to datasets from papers. The data behind the graphs is on Figshare here: [https://doi.org/10.6084/m9.figshare.26649703.v1](https://doi.org/10.6084/m9.figshare.26649703.v1?ref=openresearch.wtf) [Number of NIH Funded papers with a link to a dataset - based on Data Citation CorpusWe at Digital Science have been looking at the Data Citation Corpus, to dig deeper into data citation counts.The first release is based on a seed file that includes data citations from the following sources:Data citations from DataCite and Crossref DOI metadata, via Event Data.Data citations from the CZI Science Knowledge Graph, identified via a Named Entity Recognition model algorithm that searches for mentions to datasets in the full text of journal articles and preprints in Europe PMC.So we are basically looking at papers that have a link to a DataCite DOI or accession number.By combining this dataset with Dimensions.ai data in Google Big Query, we we’re able to add more dimensions to the dataset (pardon the pun), such as funder or institution. The Data Citation Corpus only gave us about 70% of the paper links that were resolvable DOIs. This should improve over time.This allows us to track how well things like the NIH open data policy is encouraging linking to datasets from papers.![](https://websitev3-p-eu.figstatic.com/assets-v3/7c456d7773150b33071307d7d8ee4ddf1d1b1bcf/static/media/favicon-196.911535d4.png)figshare![](https://websitev3-p-eu.figstatic.com/assets-v3/7c456d7773150b33071307d7d8ee4ddf1d1b1bcf/static/media/defaultLogo.30adffde.png)](https://figshare.com/articles/dataset/Number%5Fof%5FNIH%5FFunded%5Fpapers%5Fwith%5Fa%5Flink%5Fto%5Fa%5Fdataset%5F-%5Fbased%5Fon%5FData%5FCitation%5FCorpus/26649703?file=48481417&ref=openresearch.wtf) If you would like to play around with the data yourself, you can request it here. One limitation of the Data Citation Corpus is that it only has data to 2022\. At Digital Science, we have up to present data available through [Dimensions.ai](https://dimensions.ai/?ref=openresearch.wtf), so will continue to look for ways to track compliance and monitor the growth and reuse of open academic data going forward. [Engage with the Data Citation Corpus!The Make Data Count initiative works with researchers, repositories, librarians, publishers, institutional representatives, funders, policymakers and infrastructure providers to promote open data metrics and the responsible evaluation of research data usage. We want to hear feedback from the community to inform the development of the Data Citation Corpus. If you would like to provide feedback on the corpus or discuss a collaboration as a pilot partner, please complete the form below and we will follow up with you. For any questions about the corpus or Make Data Count, you can also contact Iratxe Puebla, Director of Make Data Count.![](https://ssl.gstatic.com/docs/forms/device_home/android_192.png)Google Docs![](https://lh6.googleusercontent.com/ZOgKVsE0fBtvQhJb-dckAy48VxhWvzxo5z_c08X5pcKepd3iUJu-lRvpvvaYBrV15V95YZBhH0U=w1200-h630-p)](https://docs.google.com/forms/d/e/1FAIpQLSd1l7ovTQs3EMw9mz4HFaVB2SuUQ8Z8FldoCDgvD74GV-vh0Q/viewform?ref=openresearch.wtf) ### Some Gold Open Access (OA) Article Processing Charges (APC) data URL: https://www.openresearch.wtf/some-apc-data/ Last updated: 2026-07-31T12:35:54.000Z A cool, new dataset around Gold Open Access (OA) Article Processing Charges (APC) was published recently. This has led to lots of interesting new research. I've also had a play with the data to answer some of my own questions. Previously, I've attempted to reverse engineer what gold OA could cost and how that scales up to a full "high rejection rate" journal, using eLife's annual financial reports. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/Screenshot-2024-08-07-at-10.19.40.png) The new dataset that has sparked all of the new analysis can be found here: > Butler, Leigh-Ann; Hare, Madelaine; Schönfelder, Nina; Schares, Eric; Alperin, Juan Pablo; Haustein, Stefanie, 2024, "Open dataset of annual Article Processing Charges (APCs) of gold and hybrid journals published by Elsevier, Frontiers, MDPI, PLOS, Springer-Nature and Wiley 2019-2023", [https://doi.org/10.7910/DVN/CR1MMV](https://doi.org/10.7910/DVN/CR1MMV?ref=openresearch.wtf), Harvard Dataverse, V1 Further on in the post, I've pulled some really interesting graphs from associated publications and re-use of the data. But I also had a dig myself. I also wanted to see whether the price of Gold OA is growing as fast as the volume of Gold OA publications is. Below we can see that although, the cost is increasing year on year since 2019, an 11% increase in this time is below global inflation. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/a6lOg-average-apc-per-year-across-all-publishers-.png) When looking at which publishers are charging the most on average, Wiley pips Springer Nature to the number 1 spot. Anecdotal evidence suggest that this is not the general understanding in the space, where most think that the Springer Nature APCs of >$11,000 for some articles (they did have the highest APC in the dataset) is the highest. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/EWiOu-average-apc-per-publisher-.png) Science covered the data release and associated arXiv papers. They reported on the inflation-adjusted USD revenue that Gold OA is pulling in. There are no signs of plateauing yet. > Authors are increasingly paying to publish their papers open access. But is it fair or sustainable? Is the pay-to-publish model for open access pricing scientists out? | Science | AAAS https://www.science.org/content/article/pay-publish-model-open-access-pricing-scientists ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/Screenshot-2024-08-07-at-10.07.52.png) Amongst the graphs in the first paper covering the actual dataset, an interesting one looks at the amount of publishers that are changing APCs and how this fits in with inflation. > An open dataset of article processing charges from six large scholarly publishers (2019-2023) [https://arxiv.org/pdf/2406.08356](https://arxiv.org/pdf/2406.08356?ref=openresearch.wtf) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/Screenshot-2024-08-07-at-09.47.06.png) A separate arXiv paper from the same authors has a publisher focus. > Estimating global article processing charges paid to six publishers for open access between 2019 and 2023 [https://arxiv.org/pdf/2407.16551](https://arxiv.org/pdf/2407.16551?ref=openresearch.wtf) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/Screenshot-2024-08-07-at-10.10.56.png) Separately, we also have the ​[Springer Nature 2023 Open Access Report](https://stories.springernature.com/oa-report-2023/?ref=openresearch.wtf) being released, which adds a few more data points. They demonstrate how Gold OA is helping with equity by waiving or offering reduced fees, to the tune of €26m! > The report also highlights key initiatives from Springer Nature in 2023 to support equity in OA. These include expansion of TAs into Africa and the Americas, waiving €26m of APCs in fully OA journals and enabling authors from low-income and low- and middle-income countries (LICs and LMICs) to publish in Nature and the Nature research journals at no cost. They also include publishing over 10,000 OA articles free of charge in 2023 in diamond OA journals and experimenting with new low-cost OA models. Open research is essential. Open Access is essential. My question as always is whether we can be doing this in a manner that is faster and cheaper? Is this the most efficient way to publish open access research. The new data will help ignite more conversations like this. Watch this space. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/08/Screenshot-2024-08-12-at-17.27.10.png) ### OpenResearch WTF April 2024 URL: https://www.openresearch.wtf/openresearch-wtf-april-2024/ Last updated: 2026-07-31T12:35:50.000Z # Policy ## [Barcelona Declaration on Open Research Information](https://barcelona-declaration.org/?ref=openresearch.wtf) More than 40 organisations have committed to being more transparent about how they share information about their research processes and outputs. The [Barcelona Declaration](https://barcelona-declaration.org/?ref=openresearch.wtf), released April 16, calls for open research information—or metadata—to be the norm. Signatories include funders and higher education institutions from the Gates Foundation to the Coimbra Group which represents 40 European universities. # Tools ## [Announcing DataCite’s First Public Data File](https://datacite.org/blog/announcing-datacites-first-public-data-file/?ref=openresearch.wtf) The public data file contains metadata for all DataCite DOIs. Specifically, this first release contains metadata records in JSON format for all DataCite DOIs in [Findable state](https://support.datacite.org/docs/doi-states?ref=openresearch.wtf#findable-doi-name) that were registered up to the end of 2023\. Each DOI has descriptive metadata for research outputs and resources structured according to the [DataCite Metadata Schema](https://schema.datacite.org/?ref=openresearch.wtf). Many of these records include links to other persistent identifiers (PIDs) for works (DOIs), people ([ORCID iDs](https://orcid.org/?ref=openresearch.wtf)), and organizations ([ROR IDs](https://ror.org/?ref=openresearch.wtf)). # Commentary ![](https://lh7-eu.googleusercontent.com/CbQllFWVBlvuWM1wpJGisKpbOTJL3UiAhH4iYus1XW-c9yeqd8jWEQC_LWonBaWTky2wHxPK0Wvc4KRgHW4KiNRHXofi3bWqpZVZ6lfgfB3lY4LVkGJDTzIwKaQvzDv3f5-NFdrzV1t5sA3eID--aSY) ## [Digital Science launches Open Principles](https://www.digital-science.com/news/open-principles-reaffirm-digital-science-commitment-to-open-research/?ref=openresearch.wtf) [Digital Science](https://www.digital-science.com/?ref=openresearch.wtf) has launched its [Open Principles](https://www.digital-science.com/resource/open-principles/?ref=openresearch.wtf), a new initiative that commits its research information solutions to open science now and into the future. The Principles are the first step in Digital Science’s journey to align more closely with the [Barcelona Declaration on Open Research Information](https://barcelona-declaration.org/?ref=openresearch.wtf) ## [Better together: BTAA Libraries, CDL and Lyrasis commit to strengthen Diamond Open Access in the United States](https://btaa.org/about/news-and-publications/news/2024/04/18/better-together--btaa-libraries--cdl-and-lyrasis-commit-to-strengthen-diamond-open-access-in-the-united-states?ref=openresearch.wtf) Diamond Open Access seems to be gaining some traction. Let's hope it has legs. ## [ASAPbio Announces New Executive Director, Board Leadership](https://asapbio.org/asapbio-announces-new-executive-director-board-leadership?ref=openresearch.wtf) > “Preprints are having a watershed moment, with key changes in federal research policy being enacted in the US, while at the same time UNESCO and others are working to spread open science equitably worldwide,” Katie said. “ASAPbio, too, is at a pivotal phase in its development. I look forward to working with the board of directors and the entire ASAPbio community to seize this crucial opportunity to accelerate research progress in the life sciences.” ### Open Access: Mo money, mo problems URL: https://www.openresearch.wtf/open-access-mo-money-mo-problems/ Last updated: 2026-07-31T12:35:48.000Z At the start of my PhD in Stem Cell Biology, I was not aware of Open Access despite Open Access journals being a thing [since the late 1980s](https://www.symplectic.co.uk/open-access-timeline/?ref=openresearch.wtf). arXiv came along in the early ‘90s, followed by PubMed Central and the first commercial OA journal, Biomed Central in the late ‘90s. PLOS launched in 2001, but it wasn’t until PLOS ONE the conversation sparked in my lab.I was introduced to Open Access publishing not because of the ideals of access for all, but the fact that PLOS ONE does not value the [perceived importance](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1637059/?ref=openresearch.wtf) of a paper as a criterion for acceptance or rejection and therefore our lab could publish lots of our research there. Perfect in the publish or perish world of academia. So, it is truly remarkable that just a decade later, the majority of publications were Open Access publications. This idea that research can be akin to a large ship that is slow to change course appears to be outdated. **Number of Open Access Publications per year** ![](https://lh7-eu.googleusercontent.com/OIzm_9KwsKMNQQjW0_XBaDsXaB_0U6JRodyvawfR_OY08Wpw9i0fWlpe6tb539D0D-D9mXtPGMZdmdn1p4eWVYF144DC1iDgARtJE8U7wWIODfAQx1PcnHv0kwyCB3WT359xF_lb_5dXLZ-VOFmHR2I) (Credit: Dimensions.ai) More recently eLife[ has taken a huge leap](https://elifesciences.org/articles/83889?ref=openresearch.wtf) in reviewing its model of publication and becoming a hybrid pre- and post- peer review publishing house. > *“We will publish every paper we send out for review as a Reviewed Preprint, a journal-style paper containing the authors’ manuscript, the eLife assessment, and the individual public peer reviews.”* This is pretty much a professionalised version of Stevan Harnard’s [Subversive Proposal](https://groups.google.com/g/bit.listserv.vpiej-l/c/BoKENhK0%5F00?ref=openresearch.wtf), which is now nearly 30 years old and called for preprints on FTP servers as a way to create global Open Access. The difference is that the time is right. eLife has the advantage of online only infrastructure and consumers. The tools that allow eLife to thrive were not available 30 years ago. So 30 years on, was Harnad's proposal right and does eLife finally have the answer? **Trust comes first** Recently, I have experimented with visualising optimal academic dissemination. I soon realised my scoring system below was flawed – as trust in research trumps the content being open, quickly disseminated or cost effective. What it did highlight is that newer types of content, such as data or code publishing, benefit from not having legacy workflows, sustainability models or the concept of prestige. The cost and complexity to make data publishing ‘trusted’ is orders of magnitude less than to make traditional paper publication fast, open or cost effective. ![](https://lh7-eu.googleusercontent.com/FtrpWhZhkfILrpQuNoeMOt7LaAdupapBBR8By90Vt5xFc9mn-Ydixa7N65_D6tK882-GOhjDXE61l4hRMNhLuCraaKo30ppT5HznzF3W5mv_olqPWYp-nf-O8YLpuh1LUuXzfAVCrOM_Br9kZzqwd3k) (Credit: Mark Hahnel) Herein lies the problem, though. Treating each of these issues as equal can lead to propagation of non-trustworthy and even false research (sometimes with an agenda). Trust needs to come first. At UKSG this month, [Chris Bennett](https://www.linkedin.com/in/ACoAAADwpxUB%5F0X5oRRtS4JvaMwfxrG5iTTznP4?lipi=urn%3Ali%3Apage%3Ad%5Fflagship3%5Fdetail%5Fbase%3BgNbaoXjWRbOuQNbYDX8LnA%3D%3D&ref=openresearch.wtf) of [Cambridge University Press & Assessment](https://www.linkedin.com/company/cambridge-university-press-and-assessment/?ref=openresearch.wtf) highlights what has happened in attempting to fix scholarly publishing by ‘solving for x’, where x is ‘open’. ![](https://lh7-eu.googleusercontent.com/BtU1mIth7AElPTNkAJsQh0QuC7QB7rOblMrGvQhPYJjMjW0rJVfF7LDOe9lBYu4B3GRi-my948133Wa800-KkQrcddK-r9EKF-P0XHaC-RFrysdFE3Cw1LpEsk8meTUnrFCCIcYzucsTqf4B7O5NRPw) Open Access publishing is complex, partly because that is what it has become, a business model. If we solve for x, where ‘x’ is maximising article processing charges (APCs), we see an explosion of content, a lot in the form of special issues (invited papers). I would make the case that we are currently publishing too many papers. ![](https://lh7-eu.googleusercontent.com/p7HKVNFIuwU11jU1CKC2EkXwSEmQNXBLBpUnX89tHCC680WLGRW-Q2Yq-usFda0h74Otd427TgEhZijV6Gtgi2VBPMXRB9olJFY52Ijwg6ajQSsymo5IH2IEPgwEvSe4Q1t12htC49-OfuJWPhv0lOA) (Credit: Hanson *et a*l) **Gates-keepers?** So perhaps the answer is to move all of the publishing to the other side of peer review and transform the way we talk about the Preliminary Scholarly Record and the Trusted Scholarly Record. The Bill & Melinda Gates Foundation has [updated its Open Access policy](https://openaccess.gatesfoundation.org/?ref=openresearch.wtf) from January 2025 in a way that lends itself to this transformation: > *“Requiring preprints and encouraging preprint review to make research publicly available when it’s ready. While researchers and authors can continue to publish in their journal of choice, preprints will help prioritize access to the research itself as opposed to access to a particular journal.* > *Discontinuing publishing fees, such as APCs. By discontinuing to support these fees, we can work to address inequities in current publishing models and reinvest the funds elsewhere.* > *We will work to support an Open Access system and infrastructure that ensures articles and data are readily available to a wider range of audiences.”* In my opinion, the Gates Foundation is taking us in the right direction. eLife is already there. The problem lies in the murky middle. Preprints shouldn't be treated as academic facts by the general public and news media. Researchers in the field can make their own decisions on whose research they trust in their field and why, as has been the case in arXiv for years. If we compare the path we have been on with regards to Open Access with some of the newer experimentation, a proposed transformative pathway emerges. Currently, we do not have a preprint for every paper. The Scholarly Record is made up of preprints, author instigated and publisher instigated peer reviewed publications. Not all peer-reviewed publications are Open Access. ![](https://lh7-eu.googleusercontent.com/gHz16_u1x5iiBPXI1rIkwYq2bxki9phCHaTTPVHKxTzaVa0DjeyDiV7ZtSev7_mXGKMsjj0nkAglGliO9r-6EX36L8yzzVeVtM9c2Wc9e4GPMxTHJIsO6hYz6alHDd2Zettmr1-kGedqmvTKUgG6oeU) (Credit: Mark Hahnel) If all papers were published as preprints before being submitted for peer review, the problem of un-checked research being picked up by news media or conspiracy theorists could cause a problem. Therefore, having community wide standards around presentation of and language describing the papers pre and post peer review should be established. We should have a ‘Preliminary Scholarly Record’ and a ‘Trusted Scholarly Record’. This author argues that a transformation of the Open Access pathway would lead to many benefits compared to the existing model. This is very inline with a proposed “Plan U” - “If all funding agencies were to mandate posting of preprints by grantees—an approach we term Plan U (for “universal”)—free access to the world’s scientific output for everyone would be achieved with minimal effort” (Server *et al*) ![](https://lh7-eu.googleusercontent.com/rjNzSjXpvRx4S38V2N250X8vGqDQcfBpKc6GG2vLVhMZgLO_JtK0AVFO_Uhj6hcmp3_8JXMQj6gmh0sZHTj1heKpxnK-oGSeeU95GKa_BWjdFuvtIYKhRb8U9sbCXiO7h_fA1PWRUQwGTJksIUUuZnI) (Credit: Mark Hahnel) | A continuation of the existing Open Access Pathway | Transformation of the Open Access Pathway | | ------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | Pay to open | Open in the preprint/green form by default | | Major cost is at the publication level | Major cost is at the peer review level | | Indeterminate innovation in the speed of research dissemination | Research disseminated by default, but with the caveat that misleading, opinionated and unfactual research is published fast | | Large corpus of trustworthy research | Large corpus of trustworthy research | | Ambiguity for the general public as to what is rigorous, peer reviewed research | Clear delineation between preliminary, unchecked research and what is rigorous, peer reviewed research | **Automatic for the people** The transformative approach also opens other areas to focus innovation on. I want to fix peer review. This is the area that I think can have the most impact in academic publishing in the next 20 years – that is not already being aggressively pursued (eg. data publishing). By establishing transparent and open quality checking of academic content, both humans and machines will be able to distinguish genuine research from fake, embellished or exaggerated findings. Automating trust markers can contribute to the bigger picture of framing. Society will advance at a greater rate if academic research is published faster (maybe even if fewer papers were published faster). Every current model falls down at the point of peer review. Peer review is largely unpaid, and the incentive structures for doing it mean that senior researchers often pass on the work to those less experienced and the benefits to the researcher themselves are a grey area, if they exist at all. As such peer review is slow, a burden. This means that we cannot achieve the Shangri-La of “trusted, fast academic publishing”, without overhauling peer review. My desire for the next transformation in Open Access is to solve for “peer review”. You could argue we will just make things even more complex. However, going back to my Optimal Academic Publishing Model, I argue that it is much more cost effective and less complex to add trust to the existing open and quickly disseminated preprints than the alternative. That is, to reverse engineer openness and fast dissemination to the closed access publishing model. Open Access has been transformative. That transformation cannot stop here. The job is half done. #### References MA Hanson, PG Barreiro, P Crosetto, D Brockington (2023) The strain on scientific publishing. arXiv | [https://doi.org/10.48550/arXiv.2309.15884](https://doi.org/10.48550/arXiv.2309.15884?ref=openresearch.wtf) | | ----------------------------------------------------------------------------------------------------------- | Sever R, Eisen M, Inglis J (2019) Plan U: Universal access to scientific and medical research via funder preprint mandates. PLoS Biol 17(6): e3000273\. [https://doi.org/10.1371/journal.pbio.3000273](https://doi.org/10.1371/journal.pbio.3000273?ref=openresearch.wtf) ### OpenResearch WTF - March 2024 URL: https://www.openresearch.wtf/openresearch-wtf-march-2024/ Last updated: 2026-07-31T12:35:51.000Z # Content ## [The Bill & Melinda Gates Foundation’s 2025 Open Access Policy Refresh](https://gatesfoundationoa.zendesk.com/hc/en-us/articles/24810787662100-Policy-Refresh-2025-Overview?ref=openresearch.wtf) No APCs! Preprints required “We will work to support an Open Access system and infrastructure that ensures articles and data are readily available to a wider range of audiences” Huge ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/04/Screenshot-2024-04-09-at-15.05.04.png) ## [Figshare, Digital Science and Springer Nature Publish “The State of Open Data: From theory to practice”](https://www.digital-science.com/state-of-open-data/?ref=openresearch.wtf) For the first time in the history of [The State of Open Data ](https://doi.org/10.6084/m9.figshare.24428194?ref=openresearch.wtf)comes a supplementary report that expands upon the results of our years of surveys. [From theory to practice](https://doi.org/10.6084/m9.figshare.25232899?ref=openresearch.wtf) offers real-life perspectives on the [opportunities and challenges of sharing research data openly](https://www.digital-science.com/blog/2023/11/the-state-of-open-data-2023/?ref=openresearch.wtf), giving us unique viewpoints as told by members of our research community – industry, funders, academic institutions, and publishers. ## [Research Excellence Framework 2029 open access consultation](https://www.ref.ac.uk/guidance/ref-2029-open-access-policy-consultation/?ref=openresearch.wtf#section-proposed-ref-2029-oa-policy-for-consultation) The four UK higher education funding bodies are opening a consultation concerning the Research Excellence Framework (REF) 2029 Open Access Policy. # Tools ## [OurResearch receives $7.5M grant from Arcadia to establish OpenAlex, a milestone development for Open Science](https://blog.ourresearch.org/ourresearch-receives-7-5m-grant-from-arcadia-to-establish-openalex-a-milestone-development-for-open-science/?ref=openresearch.wtf) With this 5-year grant, OurResearch expands their open science ambitions to replace paywalled knowledge graphs with OpenAlex. ## [Innovative Open Research Publisher PeerJ Joins Taylor & Francis](https://newsroom.taylorandfrancisgroup.com/peerj-joins-taylor-and-francis/?ref=openresearch.wtf) All PeerJ journals offer high-quality peer review and rapid publication, supported by PeerJ’s own submission and peer review platform, and dedicated contributor support. Articles are selected on scientific value and methodological soundness, providing a forum for world-class research addressing many of the globe’s current challenges. PeerJ also hosts a digital hub for the International Association for Biological Oceanography, promoting the advancement of knowledge of the biology of the sea. ### OpenResearch WTF - February 2024 URL: https://www.openresearch.wtf/openresearch/ Last updated: 2026-07-31T10:00:37.000Z # Content ### [eLife’s New Model: One year on](https://elifesciences.org/inside-elife/66d43597/elife-s-new-model-one-year-on?ref=openresearch.wtf) eLife's new publishing model, a year after its launch, has led to over 6,200 submissions, with more than 1,300 reviewed preprints published. This model emphasises speed and transparency, combining preprints with expert peer reviews and assessments to facilitate quicker communication of research. The median time from submission to publication is 91 days, significantly faster than the legacy model. Authors have positively rated the quality of public reviews and the consultative review process, underscoring the model's success in enhancing scientific communication ### [Study on scientific publishing in Europe released by the Publications Office of the EU](https://op.europa.eu/en/publication-detail/-/publication/23fafd0e-d646-11ee-b9d9-01aa75ed71a1?ref=openresearch.wtf) The study, commissioned to support the European Commission's open access policies, explores the complexity and opacity of financial flows in scientific publishing. It aims to offer a deeper understanding of the costs and practices associated with scholarly publications, advising on policy actions for enhanced transparency. The report covers national policies on publication financing, the diversity of scientific publishing in Europe, and the challenges in accessing detailed financial information on publication costs. It concludes with recommendations to increase the transparency of publishing costs and the availability of information on open access investments. ### [The Cost and Price of Public Access to Research Data: A Synthesis is published by IOI](https://doi.org/10.5281/zenodo.10729575?ref=openresearch.wtf) This reports on the new requirements for United States federal research funding to make scholarly outputs freely available. It explores the definitions of cost, price, reasonable, and allowable expenses related to research data curation and sharing, emphasising labour as the most significant cost. The study examines cost modelling in research data management and digital preservation, highlighting the variability of costs across disciplines and the importance of transparency in repository costs to aid researchers and funders. ### [The Platform Developers in a Federated Model of Diamond Open Access – the diamond papers](https://thd.hypotheses.org/366?ref=openresearch.wtf) Diamond Open Access seems to be gaining steam following their recent annual conference. Public Knowledge Project's (PKP) is involved in advancing DOA through a proposed Global Diamond Federation. PKP supports this federation, leveraging its extensive experience in DOA to contribute to the global infrastructure and community engagement necessary for the success of DOA initiatives. This collaboration aims to enhance the accessibility and quality of scholarly publishing worldwide. ## [TL;DR Shorts launch](https://www.youtube.com/playlist?list=PL6ux2aPwlTJHCFpYehBnrfPgMjDbjV771&ref=openresearch.wtf) Ever wished you could get to the heart of complex research topics without spending hours wading through lengthy discussions? Our TL;DR Shorts series is your solution! Our playlist features popular figures, thought leaders, and experts who distil their knowledge into concise, engaging commentary. These videos are designed to be informative, accessible, and thought-provoking. ![](https://lh7-eu.googleusercontent.com/DD8yNYH2oSzEU4o3ttb3O7ECjmuQjL9V_Mxlvgt_P2h3BgRPgIwY1SWLvnCpMPpxrZQUPOPgQZ1NB2xMq1yh1WHCwFfG3WVPLPAd_BisJ8OSOzhosQw8oj60h9UUNjNCQR2EK7cT1mLp7-cpEEmOAcI) # Tools ### [Digital Science has developed its first custom GPT solution](https://www.digital-science.com/blog/2024/02/fast-forward-a-new-approach-for-ai-and-research/?ref=openresearch.wtf) Digital Science has unveiled Dimensions Research GPT and its Enterprise version, merging Dimensions' extensive research data with ChatGPT's AI capabilities. This innovation aims to provide researchers with more accurate, research-specific answers grounded in millions of Open Access publications and comprehensive data including grants, clinical trials, and patents. It represents a significant step towards integrating generative AI in scientific research, offering a streamlined, efficient way to access trusted information and fostering responsible AI development in the research community. For more details, visit the[ Digital Science blog](https://www.digital-science.com/blog/2024/02/fast-forward-a-new-approach-for-ai-and-research/?ref=openresearch.wtf). ### [Government report recommends Jisc leads on development of digital research information ecosystem](https://www.jisc.ac.uk/news/all/government-report-recommends-jisc-leads-on-development-of-digital-research-information-ecosystem?ref=openresearch.wtf) The Government’s response to the independent review of research bureaucracy calls upon Jisc’s expertise to improve research culture. The UK government has released its response to the [independent review of research bureaucracy](https://www.gov.uk/government/publications/review-of-research-bureaucracy?ref=openresearch.wtf), with recommendations that Jisc helps improve digital platforms to reduce costs and free up researchers. ### [Silverchair, Oxford University Press and Hum launch Sensus Impact](https://www.sensusimpact.com/info-parent/1172/about-sensus-impact?ref=openresearch.wtf) A [dashboard of usage of papers filtered by funder](https://www.sensusimpact.com/info-parent/1172/about-sensus-impact?ref=openresearch.wtf) ![](https://lh7-eu.googleusercontent.com/DNJozvqgD1WidrZlQUVHFMrI8gFW2wdIjjXG6xffgGbiWVmpp02Io39Qp7qcfoz1adboNxy0RR--wbQLE_yvSCRdjk6ZoICI08kipnUYBd8LQyJx3YYUJlOEsIIsTdwiPjfUlKK3hKwdOrGWMWL9lcw) ### [Consensus launch Copilot AI Search Engine for Research](https://consensus.app/home/blog/introducing-the-consensus-co-pilot/?ref=openresearch.wtf#klue-view-877417) The aim is to allow users to Ask Consensus a natural language research question, get an expanded answer with citations ### [OpenAlex announces Topics](https://twitter.com/OpenAlex%5Forg/status/1757138494747480235?ref=openresearch.wtf) A new system for categorising hundreds of millions of research publications! ### [ReadCube expands its literature management platform with the launch of Literature Review](https://www.readcube.com/en/request-a-demo/?ref=openresearch.wtf) ReadCube, has launched "Literature Review," a new tool designed to streamline literature review workflows for research organisations. This solution integrates with ReadCube's platform, enhancing the efficiency of systematic literature reviews (SLR) crucial for organisations in pharmaceuticals, medical devices, and health technology assessment sectors. Addressing the challenges of growing publication volumes and stringent regulatory standards, Literature Review by ReadCube offers an AI-enhanced, user-friendly approach to manage the complexities of SLR processes, thereby facilitating comprehensive and compliant reviews ### Making poorly described data FAIR-er using GenAI URL: https://www.openresearch.wtf/making-poorly-described-data-fair-er-using-genai/ Last updated: 2026-07-31T10:00:34.000Z Generalist data publishing repositories such as Figshare or Zenodo often have little curation. As a result, we see lots of data made available that is potentially very valuable to the community, but not very well described. This small experiment looks at whether Generative AI tools, in this case ChatGPT 4 can improve the metadata (and thus FAIR-ness) of poorly described published datasets. I decided to look at examples where the dataset has been reuse, even without good metadata, suggesting that the paper associated with the data, describes it well. This dataset is a self-published dataset on Figshare. It does not have a descriptive title and has very light metadata Delgado-Baquerizo, Manuel (2022). ITS. figshare. Dataset. [https://doi.org/10.6084/m9.figshare.5923876.v1](https://doi.org/10.6084/m9.figshare.5923876.v1?ref=openresearch.wtf) ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/03/screencapture-figshare-articles-dataset-ITS-5923876-2024-03-03-14_24_37.png) Despite poor metadata, the dataset itself has been cited several times. Using [Dimensions.ai](dimensions.ai) we can see that of the 4 papers citing the dataset, 3 have the same author and one is a completely different set of authors. We can assume that the earliest paper to cite the dataset is the one that describes it. ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/03/Screenshot-2024-03-03-at-14.30.49.png) ### Using a link to the Figshare landing page alone *ChatGPT4* **Prompt:** Can you create a title, a description, categories and tags to be used in academic literature, for this dataset: https://figshare.com/articles/dataset/ITS/5923876 **Response:** - **Title**: "Comprehensive ITS Dataset for Unexplored Ecological Niches" - **Description**: This dataset comprises extensive ITS region sequences, aimed at shedding light on the microbial biodiversity within under-researched ecological niches. Compiled by Manuel Delgado-Baquerizo, it provides a valuable resource for ecological research, especially in areas lacking classification within existing categories. Its extensive coverage and open license make it an essential asset for researchers seeking to understand ecological dynamics more deeply. - **Categories**: Ecology, Biodiversity, Microbial Genomics, Environmental Science - **Tags**: ITSII, microbial diversity, unclassified ecology, environmental genomics, dataset, open data ### Using one of the datasets alone *ChatGPT4* **Prompt:** ![](https://storage.ghost.io/c/45/ef/45ef5edd-7737-4b73-a169-a85c335bb1f1/content/images/2024/03/Screenshot-2024-03-03-at-14.38.09.png) **Response:** - **Title**: "High-Throughput Sequencing Data for Environmental Microbial Analysis" - **Description**: This dataset comprises paired-end raw sequencing reads from high-throughput sequencing experiments, encapsulated in six files. The data represents three distinct samples (70, 73, and 123), with each sample's forward (R1) and reverse (R2) reads provided. These reads are essential for comprehensive microbial community analysis, offering insights into the genetic diversity and functional capabilities of environmental samples. The dataset serves as a foundational resource for genomic studies aimed at understanding microbial ecology and evolution. - **Categories**: Microbial Genomics, Environmental Genomics, High-Throughput Sequencing, Bioinformatics - **Tags**: sequencing data, environmental microbiology, paired-end reads, genomic diversity, FASTQ ### Using a link to the Figshare landing page, one of the datasets and the link to the paper *ChatGPT4* **Prompt:** Can you create a title, a description, categories and tags to be used in academic literature, for this dataset: https://figshare.com/articles/dataset/ITS/5923876 that is referenced in this paper: https://www.nature.com/articles/s41467-019-10373-z One of the files from the dataset is attached **Response:** - **Title**: "Global Soil Fungal Dominance: An ITS Dataset Analysis from Diverse Ecosystems" - **Description**: This dataset, crucially referenced in the groundbreaking study on global soil fungal communities, contains ITS region sequences that have revealed a few Ascomycota taxa's dominance across various ecosystems worldwide. The research underscores the ecological significance of these taxa in terrestrial biomes, highlighting their roles in nutrient cycling, ecosystem functioning, and potential resilience to environmental changes. - **Categories**: Soil Science, Fungal Ecology, Biodiversity, Microbial Ecology - **Tags**: Global fungal distribution, Ascomycota, ITS sequencing, soil microbiology, ecological dominance ### Conclusion Each set of metadata is an improvement on the FAIR-ness of the metadata for the Figshare output. Obviously as an n of 1, we cannot read to much into this. But as we scratch the surface of what Generative AI can do for improving metadata in generalist repositories, this result suggest that this is an avenue worth exploring. ### Who benefits when, from FAIR data? Part 3 - The Public URL: https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-3/ Last updated: 2026-07-31T10:00:43.000Z **Bang for Buck** Following [part 1of this series, which focused on researchers](https://www.digital-science.com/tldr/article/who-benefits-from-fair-data/?ref=openresearch.wtf), and [part 2, which focused on machines](https://www.digital-science.com/tldr/article/who-benefits-when-from-fair-data-pt-2/?ref=openresearch.wtf), it’s time to talk about the third puzzle piece – those impacted by research. The general public, who pay for research via their taxes, should be able to benefit from it democratically. This is not the case. Non-researchers do not have access to over 50% of the academic literature (below). This is particularly problematic when it comes to healthcare, with the caveat that each nation-state has its own way of providing treatment to the sick. The basic science that most if not all pharmaceutical research is built on is largely government-funded. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1655b89c-c035-4fb5-8eac-dc07e84c2904_1600x1066.png) We also live in a time where researchers are evidencing that “[Science Is Getting Less Bang for Its Buck](https://www.theatlantic.com/science/archive/2018/11/diminishing-returns-science/575665/?ref=openresearch.wtf)“. Despite vast increases in the time and money spent on research, progress is barely keeping pace with the past. This means that we need to come up with a new way of working to generate new discoveries and knowledge. As highlighted in[ the previous post](https://www.digital-science.com/tldr/article/who-benefits-when-from-fair-data-pt-2/?ref=openresearch.wtf) in the series; “By leveraging large volumes of FAIR data, machines can accelerate the processing and analysis of information, leading to the generation of new knowledge and empowering researchers in their endeavours”. Findable, accessible, interoperable and reproducible (FAIR) open academic data enables the general public to access a wealth of knowledge and information that was previously confined within academic circles. By making research findings, scholarly articles, and educational resources openly available, open academic data promotes a more inclusive and equitable dissemination of knowledge. This empowers individuals, regardless of their affiliation with academic institutions, to stay informed, expand their understanding of various subjects, and make better-informed decisions in their personal and professional lives. Policymakers, government agencies, and non-profit organisations can leverage FAIR data to inform policy formulation, assess the effectiveness of interventions, and address social, environmental, and economic issues. By relying on transparent and reliable data, decision-makers can make more informed choices, leading to better outcomes and improved public welfare. This comes full circle as we see the same policymakers advocating for FAIR data. This ranges from encouragement from cOAlition S…: > Although the Plan S principles refer to [peer-reviewed](https://www.coalition-s.org/faq/what-plan-does-coalition-s-have-to-increase-transparency-in-the-peer-review-process/?ref=openresearch.wtf) scholarly publications, cOAlition S also strongly encourages that research data and [other research outputs](https://www.coalition-s.org/faq/what-types-of-works-fall-under-plan-s/?ref=openresearch.wtf) are made as open as possible and as closed as necessary. > > [Guidance on the implementation of Plan S](https://www.coalition-s.org/guidance-on-the-implementation-of-plan-s/?ref=openresearch.wtf) …to the core goal of the Office of Science and Technology Policy [(OSTP) Memo](https://www.whitehouse.gov/wp-content/uploads/2022/08/08-2022-OSTP-Public-Access-Memo.pdf?ref=openresearch.wtf) (the OSTP is responsible for advising the President on matters related to science, technology, and innovation, and their memos often outline policy priorities, initiatives, or guidelines in these areas): > The goal of the memo is to provide free, immediate (without embargo), and equitable access to research that is federally funded. This applies to all federal agencies. This applies to both peer-reviewed publications and underlying scientific data > > [OSTP Memo](https://www.whitehouse.gov/wp-content/uploads/2022/08/08-2022-OSTP-Public-Access-Memo.pdf?ref=openresearch.wtf) **Trust** Open, FAIR data promotes transparency and accountability in research and academia. When research findings and data are openly available, it allows for peer review, reproducibility, and scrutiny by the wider public. This fosters a culture of accountability, ensuring that research is conducted with integrity and that potential biases or errors can be identified and addressed. In a post-truth world, it is essential that the general public can actively engage in discussions, debates, and fact-checking processes, promoting a more robust and reliable knowledge base. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb88130f6-7c1e-43d8-83e5-790dcb6c6fcb_1024x433.png) FAIR data can be instrumental in addressing various societal challenges. For example, in healthcare, FAIR data can support medical research, improve diagnosis and treatment outcomes, and facilitate the development of personalised medicine. In urban planning, FAIR data can help identify patterns and trends, leading to smarter and more sustainable city designs. FAIR data can also contribute to tackling social inequalities and disparities by providing an objective basis for identifying and addressing systemic issues. FAIR data can also help solve different problems. Open, FAIR data can help revolutionise healthcare by assisting in early disease detection, diagnosis, and treatment. Machine learning algorithms can analyse large amounts of medical data, including patient records, genetic information, and medical images, to identify patterns and predict disease outcomes. FAIR data collection in a longitudinal manner can help real-time monitoring and management of ecosystems, prediction of natural disasters, optimise energy usage, and promote efficient resource management. AI-driven models can aid in climate change research, wildlife conservation, and ecological restoration, facilitating evidence-based decision-making and proactive environmental stewardship. Whilst research publications also can help here, they struggle with the Shangri-La of academic publishing: fast, trusted, open, and cost-effective. With the need for peer review, research publications are rarely fast and only open 50% of the time. FAIR data aligns itself with these principles by definition. We can see differences between publications and datasets when assessing the last 10 years of outputs when categorised by the United Nations’ Sustainable Development Goals (SDGs). The footprint is very different. SDGs are by definition the biggest challenges that humanity faces. The public benefits if we can use all of the tools in the research arsenal to work towards these SDGs. **Sustainable Development Goal categorised Academic Research Publications and Datasets** ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3123a48-fce6-49e9-a1d7-1d70a56b51da_1024x487.png) Taken together, this information tells us what we already know in the [first post](https://www.digital-science.com/tldr/article/who-benefits-from-fair-data/?ref=openresearch.wtf) of this trilogy – “Encouraging or mandating open and FAIR data is the most obvious return on investment (ROI) improvement in academic publishing.” Funders are aware of this. Now we just need to let everyone else know. Be sure to do your part! ### Who benefits when, from FAIR data? Part 2 – Machines URL: https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-2-machines/ Last updated: 2024-02-05T15:55:47.000Z I think that Alphafold will win the Nobel Prize by 2030\. In reality, I don’t think they will award it to the artificial intelligence (AI) itself, but to the [lead authors on the Nature paper](https://doi.org/10.1038/s41586-021-03819-2?ref=openresearch.wtf) describing their work (John Jumper and Deepmind CEO, Demis Hassabis). AlphaFold is an AI program which performs predictions of protein structure. It is developed by DeepMind, a subsidiary of Alphabet (Google) that recently merged with the brain team from Google Research [to become Google Deepmind](https://www.deepmind.com/blog/announcing-google-deepmind?ref=openresearch.wtf). Whilst there are those who have suggested that the press coverage of Alphafold has [been hyperbolic](https://www.skynettoday.com/briefs/alphafold2?ref=openresearch.wtf), for me it is less about whether ‘the protein folding problem has been solved’ and more about the giant leap that a field of research has progressed by due to one project. Whilst Alphafold has got the lion’s share of the attention, there are several other advanced projects looking to have as much, if not more of an impact. All of these projects could not exist if it were not for large amounts of well described FAIR academic data. NameFieldNotesDeepTrio ([https://github.com/google/deepvariant](https://github.com/google/deepvariant?ref=openresearch.wtf))GenomicsA deep learning-based trio variant caller built on top of DeepVariant. DeepTrio extends DeepVariant’s functionality, allowing it to utilise the power of neural networks to predict genomic variants in trios or duosClimateNet ([https://github.com/andregraubner/ClimateNet](https://github.com/andregraubner/ClimateNet?ref=openresearch.wtf))Climate changeClimateNet seeks to address a major open challenge in bringing the power of deep learning to the climate community, viz. creating community-sourced, open-access, expert-labelled datasets and architectures for improved accuracy and performance on a range of supervised learning problems, where plentiful reliablylabelled training data is a requirementDeepChem ([https://deepchem.io/about](https://deepchem.io/about?ref=openresearch.wtf))Cheminformatics and drug discoveryDeepChem aims to provide a high quality, open-source toolchain that democratises the use of deep-learning in drug discovery, materials science, quantum chemistry, and biologyThe Materials Project ([https://materialsproject.org/](https://materialsproject.org/?ref=openresearch.wtf))Materials scienceThe Materials Project is an initiative that harnesses the power of AI and machine learning to accelerate materials discovery and design. It provides a vast and continuously growing database of materials properties, calculations, and experimental dataThe Dark Energy Survey ([https://www.darkenergysurvey.org/](https://www.darkenergysurvey.org/?ref=openresearch.wtf))AstrophysicsThe Dark Energy Survey (DES) is a scientific project designed to study the nature of dark energy Academic Projects making use of FAIR data and AI to drive systemic change in a research field. Of course, when we refer to machines benefiting, what we are really referencing is human designed models benefitting from large swathes of well described data. The machines are made up of these algorithms, FAIR data and large amounts of compute. In a time when all of the low hanging fruit in research has been picked. Larger volumes of FAIR data can be processed by machines much more efficiently than humans. Processing this information to infer trends and predicted models allows human expertise to leverage information at a much faster rate, leading to new knowledge. We are in the ‘low hanging fruit’ phase of research powered by humans and machines (algorithms, FAIR data and compute). ## A Virtuous Cycle AI and machine learning can be used to create more detailed, FAIR-er datasets to be consumed by the machines. One of the more advanced areas this is happening in is mining the existing academic literature. Efforts like the [AllenNLP](https://allenai.org/allennlp?ref=openresearch.wtf) Library and [BioBERT](https://hpc.nih.gov/apps/BioBERT.html?ref=openresearch.wtf), a biomedical language representation model designed for biomedical text mining tasks can help semantically enhance large datasets. Large language models (LLMs) and AI systems rely heavily on vast amounts of training data to learn patterns and generate accurate outputs. Open, FAIR academic data can serve as a valuable resource for training these models. By using diverse and well-curated academic data, LLMs and AI systems can better understand the nuances of various academic disciplines, terminology, and writing styles. This leads to improved performance in tasks such as natural language understanding, text generation, and information extraction. Subject-specific data repositories provide homogenous FAIR data for a wide range of specialised domains. Incorporating open, FAIR academic data into LLMs enables them to develop domain-specific expertise. This expertise enhances their ability to understand and process domain-specific jargon, terminology, and concepts accurately. Consequently, these models can provide more insightful and nuanced responses in specialised areas. Ultimately, as well as benefiting from having large amounts of well-described FAIR data, the machines can also support our original beneficiaries of FAIR data; [the researchers](https://www.digital-science.com/tldr/article/who-benefits-from-fair-data/?ref=openresearch.wtf). ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8d0d068-3d8e-475c-a966-16b6de7452df_1024x558.png) By leveraging large volumes of FAIR data, machines can accelerate the processing and analysis of information, leading to the generation of new knowledge and empowering researchers in their endeavours. Academia needs to prioritise feeding the machines. ### All of the knowledge for all of the machines URL: https://www.openresearch.wtf/all-of-the-knowledge-for-all-of-the-machines/ Last updated: 2026-07-31T10:00:05.000Z ### All peer-reviewed articles available to all, and to AI? Inevitable, Illegal? There are two slow revolutions happening in academic publishing: 1. Blurring the lines between publishing peer-reviewed and non-peer-reviewed content 2. Moving to a world where every paper is published Open Access Does this lead to complications when applying machine-learning models to the research literature? By processing and analysing vast amounts of scholarly work, AI will enable breakthroughs in research, innovation, and problem-solving across numerous disciplines. It will possess an unparalleled ability to identify patterns, synthesise information, and generate novel insights that could lead to transformative discoveries. But what can be legally accessed and what can be trusted? ## Blurring the lines between publishing peer-reviewed and non-peer-reviewed content The preprint revolution kicked off in Physics in 1991 with [arXiv](https://arxiv.org/?ref=openresearch.wtf) by Paul Ginsparg. arXiv is internationally acknowledged as a pioneering and successful digital archive and open-access distribution service for research articles, pre peer review. Preprints had a big resurgence in the last 5 years as the desire for faster publishing of research took hold. Notably, [biorxiv](https://www.biorxiv.org/?ref=openresearch.wtf) seems to be doing for Biology, what arXiv has achieved in the Physics world. The logic for immediate publishing of research is probably best summed up by the eLife team - Authors should be able to share their work freely and openly when they think it is ready. - Peer review should consist of scientists publicly sharing their assessments of already published papers, either under the auspices of an editorial organization that oversees the review process, or on their own. - Works of science should be reviewed by multiple relevant groups and individuals throughout their useful lifespan. [They have taken a huge leap](https://elifesciences.org/articles/83889?ref=openresearch.wtf) in reviewing their model of publication and becoming a hybrid pre- and post- peer review publishing house. > we will publish every paper we send out for review as a Reviewed Preprint, a journal-style paper containing the authors’ manuscript, the eLife assessment, and the individual public peer reviews. This is pretty much a professionalised version of Stevan Harnard’s [Subversive Proposal](https://groups.google.com/g/bit.listserv.vpiej-l/c/BoKENhK0%5F00?ref=openresearch.wtf), which is now nearly 30 years old: > If all scholars’ preprints were universally available to all scholars by anonymous ftp (and gopher, and World-wide web, and the search/retrieval wonders of the future), NO scholar would ever consent to WITHDRAW that preprint from the public eye after the refereed version was accepted for paper “PUBLICation.” Instead, everyone would, quite naturally, substitute the refereed, published reprint for the unrefereed preprint. But this does beg the question, should large language models (LLMs) be taking in un-peer reviewed content to generate new knowledge? If not, how do we ensure it is filtered out. ## Moving to a world where every paper is published Open Access There are similar grey areas in the world of open access. In 2002, a meeting of scholars and activists took place in Budapest, Hungary, leading to the formulation of the Budapest Open Access Initiative (BOAI). BOAI defined open access as free availability on the internet, permitting any user to read, download, copy, distribute, print, search, or link to the full texts of these articles. We suggest that this is true for both humans and machines. The green open access route often results in papers that can be similar, but not exactly the same as the closed access version. However due to actions by [funders around the world](https://v2.sherpa.ac.uk/view/funder%5Fby%5Foa%5Farch%5Freq/requires.html?ref=openresearch.wtf), we have passed 50% of papers being Open Access in some way, each year. \[This is further compounded by the release of the OSTP Memo from the Whitehouse\](Update their public access policies as soon as possible, and no later than December 31st, 2025, to make publications and their supporting data resulting from federally funded research publicly accessible without an embargo on their free and public release;) which dictates that Federal Agencies should: > Update their public access policies as soon as possible, and no later than December 31st, 2025, to make publications and their supporting data resulting from federally funded research publicly accessible without an embargo on their free and public release; ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc298702a-c08d-4304-b46b-1a7f0d5fb354_1024x684.png) The new disruptive models highlighted here create a ‘before and after’ feel to academic publishing. Every paper will be open and available to humans and machines alike. But what about the millions of papers still paywalled. They cannot be forgotten. The decentralised web and p2p workflows have led to paywalled content becoming open on the web. Most notably, Sci-hub. However, no new papers have been added for 2 years, due to a pending Indian court case. This hasn’t stopped [Alexandra Elkaban](https://twitter.com/ringo%5Fring?ref=openresearch.wtf) from winning the [EFF Pioneer Award](https://twitter.com/ringo%5Fring/status/1680064493601304576?ref=openresearch.wtf), or slowed the [huge amount of usage that Sci-hub maintains](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5428489/?ref=openresearch.wtf). It is not hard to find [torrents of this content](https://libgen.rs/scimag/repository%5Ftorrent/?ref=openresearch.wtf), or new platforms taking advantage of decentralised infrastructure like ‘The InterPlanetary File System’ (IPFS). IPFS is a protocol, hypermedia and file sharing peer-to-peer network for storing and sharing data in a distributed file system. Recent incumbents include [Standard Template Construct](https://standard--template--construct-org.ipns.dweb.link/?ref=openresearch.wtf#/), Nexus Telegram bots and [Science Hub Mutual Aid](https://www.wosonhj.com/?ref=openresearch.wtf). If I can access this content via google, bots and machines would no doubt assume to do the same. The collective intelligence stored within the world’s peer-reviewed academic literature will empower AI to propel human progress by accelerating scientific discovery and fostering a deeper understanding of the world around us. If this is not a legal approach today, the academic world needs to create easy ways to AI-mine legally accessible full-text articles. ### Who benefits when, from FAIR data? Part 1 – Researchers URL: https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-1-researchers/ Last updated: 2026-07-31T10:00:42.000Z Findable, Accessible, Interoperable and Reusable (FAIR) data has become an aim for a large proportion of academic funders globally. This has happened in the 7 years since the publication of ‘[The FAIR Guiding Principles for scientific data management and stewardship](https://doi.org/10.1038/sdata.2016.18?ref=openresearch.wtf)’. Encouraging or mandating Open and FAIR data is the most obvious return on investment (ROI) improvement in academic publishing. Academic funders invest substantial resources in research projects. By promoting the FAIR data principles, they aim to maximise the return on their investments by ensuring that the data generated from funded research can be fully utilised and have a broader impact beyond the original project. But who else is benefiting from FAIR data? This is part one of a 3 piece series that aims to highlight the author’s thoughts on when different sectors/players will benefit from Open FAIR data. In this instalment, we delve into how researchers are the primary beneficiaries, despite shouldering the responsibility of publishing FAIR data. By adhering to FAIR principles, researchers can make their data findable and accessible to the broader scientific community. This increased visibility enhances collaboration opportunities, fosters interdisciplinary research, and paves the way for novel discoveries. The interoperability aspect of FAIR data allows researchers to combine datasets from diverse sources, facilitating new insights and innovative approaches to problem-solving. The burden of publishing FAIR data currently lies primarily with the researchers themselves, adding an extra layer of effort to an already demanding research process. However, collaborative efforts among stakeholders, including academic institutions, publishers, and infrastructure providers, can alleviate this burden and streamline the FAIR data publication process. Supporting researchers with appropriate resources, tools, and incentives will help foster a culture of data sharing and enhance the adoption of FAIR data practices. 70% of researchers responding to our [State of Open Data Survey](https://doi.org/10.6084/m9.figshare.21276984.v5?ref=openresearch.wtf) are *required* to make their data openly available. As the publication of open research data moves from a ‘nice to have’ to a core part of the academic literature, the academic community needs to follow an achievable roadmap to make data useful and subsequently research more efficient. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0089e740-0e5d-4380-b309-4dfb8261aec6_730x396.gif) We are optimistic of a future state world, where a group of researchers generates new information based on new hypotheses, and makes it available on Figshare (or any other open repository) for others to then pull together the combined world’s knowledge to look for new patterns and discoveries that the original authors had not thought to look for. This is already happening — with the tracking of citations and Altmetrics on item pages, users can see who is mentioning and reusing their data and the progression that has had on scientific advancement. For example, take [this use case](https://doi.org/10.6084/m9.figshare.4654045?ref=openresearch.wtf) of a University of Sheffield researcher who made his data available on Figshare and has since developed collaborations and new lines of enquiry into adjacent areas of research. Deepmind’s Alphafold has recently demonstrated the power of well-curated, homogenous open data. When thinking of research data, future consumers will not just be human researchers — we also need to feed the machines. This means that computers will need to interpret content with little to no human intervention. For this to be possible, the outputs need to be in machine-readable formats and the metadata needs to be sufficient to describe exactly what the data are and how the data was generated. Whilst repositories can be responsible for ensuring the machines can access and interpret the content using defined open standards, the quality of the content and metadata, as well as consistency in how data is shared within a field, lies at the feet of the researchers themselves. There is not currently enough support for metadata curation at the generalist repository level, but this does not mean that poorly described data is not useful to humans and machines. ### There is a home for every dataset. 1. Ideally each subject would have a subject-specific data repository with custom metadata capture and a subject specialist to curate the data. If there is a subject-specific repository that is suitable for your datasets, use that one. 2. If there is no subject-specific repository, you should check with your institution’s library as to whether they have a data repository. If they do, they will most likely have a team of data librarians to guide you through your dataset publication. 3. If there is no subject-specific repository, you should upload and publish your data on Figshare.com, or another suitable generalist repository. [How to upload and publish your data on Figshare.com](https://help.figshare.com/article/how-to-upload-and-publish-your-data?ref=openresearch.wtf) ### To make open data useful for researchers – some context is needed Data with no curation can be useful, but they must at least adhere to the following: - Academic files and metadata are available on the internet in repositories that follow [best practice norms](https://www.coar-repositories.org/news-updates/coar-sparc-response-to-the-ostp-draft-desirable-characteristics-of-repositories-managing-data/?ref=openresearch.wtf). In order to make data Findable and Accessible (the F and A of FAIR), the metadata should be sufficient to be discoverable through a Google search or available when clicking through from a peer-reviewed publication. We can see that humans are good at finding context around datasets that are not well described as a standalone output. The related paper often has enough context to make the data re-usable by separate research groups. When looking at some of the most cited outputs across Figshare infrastructure, [the one with the lightest metadata](https://doi.org/10.25909/5becfa45c176f?ref=openresearch.wtf) has the following as a description: “The data comprises phenotypic measurements from two field trials in South Australia and genotyping information for more than 500 diverse wheat accessions.” There is no README file. And yet this has been re-used and cited by another research group. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd9398ba-8fdf-42ee-b6f6-181cdd0ff92b_1600x791.png) Internally at Figshare, we have previously looked at data associated with publications to see if humans can interpret and re-use the data. To showcase the reusable possibilities with Figshare, we enlisted Jan Tulp, data visualisation expert and founder of [Tulp Interactive](http://tulpinteractive.com/?ref=openresearch.wtf), to reuse three datasets published in Figshare as part of a publication. The task was to take these three datasets and create visualisations using only the data, the metadata found in the item page, and the original publication. Whilst the [metadata wasn’t fantastically detailed](https://doi.org/10.6084/m9.figshare.c.3808039%5FD3.v1?ref=openresearch.wtf), the associated publication, linked from the landing page on Figshare, allowed Jan to re-interpret the data with a [beautiful visualisation](https://janwillemtulp.github.io/noaa-pacific-ramp-fish/?ref=openresearch.wtf), showing fish populations at different sites around islands. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2afa6-2058-4b51-887c-1f68e407d55d_814x754.gif) We are also seeing evidence that machines can infer new knowledge from lightly described datasets. Figshare works with publishers to make outputs that are not the peer-reviewed article citable, previewable and persistent. This is often associated datasets that may be checked by editorial staff at the journal. The vast majority of this content is supplemental data, with varying levels of descriptive metadata and context. Many naysayers suggest this information has no use outside the context of the individual paper however we do see evidence that this information can be re-used at scale, as exemplified by a recent paper, “[A resource for automated search and collation of geochemical datasets from journal supplements](https://www.nature.com/articles/s41597-022-01730-7?ref=openresearch.wtf)”: > *“Here, we present open-source code to allow for fast automated collection of geochemical and geochronological data from supplementary files of published journal articles using API web scraping of the Figshare repository. Using this method, we generated a dataset of \~150,000 zircon U-Pb, Lu-Hf, REE, and oxygen isotope analyses. This code gives researchers the ability to quickly compile the relevant supplementary material to build their own geochemical database and regularly update it, whether their data of interest be common or niche.”* > > https://www.nature.com/articles/s41597-022-01730-7 The authors also propose a set of guidelines for the formatting of supplementary data tables that will allow for published data to be more readily utilised by the community. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd6a7819-6ec7-4cf9-826c-a094d5cf7a4d_1600x1371.png) Whilst the above examples demonstrate that poorly-described research data can be useful in some situations, it is not something that should be encouraged and merely highlights that open data is better than no data. The research community should be focussed on the principles of open, FAIR data as the biggest potential return on investment for academic funders in the next 50 years – both in the efficiency of research and the discovery of new knowledge. For researchers, there are two immediate wins by creating and consuming FAIR data. By ppublishing well-described data, they have an easy route to greater impact of their research. There is also a wealth of information that is available today in open repositories. By becoming an early consumer of this, researchers can get the upper hand on their peers by obtaining a fuller picture of a research field. By embracing FAIR data, researchers empower themselves and the wider scientific community, enabling the acceleration of knowledge and fostering innovation. In the next two parts of this blog series, we will explore how other sectors and players can harness the power of open, FAIR data. Stay tuned for more insights on this transformative paradigm shift in data management and its implications for various stakeholders. ### Academic data publishing is the biggest ROI in research today URL: https://www.openresearch.wtf/academic-data-publishing-is-the-biggest-roi-in-research-today/ Last updated: 2026-07-31T10:00:03.000Z After a decade of running Figshare (a data publishing repository), I think the biggest cultural, and impactful change of the last 10 years to be ‘encouraging researchers to put data on the internet’. ROI = Return on Investment. With steady behavioural change, I think the next 10 years are all about making that data useful for machines. This gives funders and academia the lowest risk:reward ratio for a next paradigm in research. This is how we move further, faster. We will naturally get to better transparency and reproducibility in research by making data openly available, linked from peer reviewed publications. However, to create a real step-change in discovery of new knowledge, research needs to take advantage of artificial intelligence acting upon large swathes of reproducible research data with homogenous metadata. Alphafold from Deepmind was the first example of how monumentally powerful this combination can be in 2020. > *“AlphaFold is a once in a generation advance, predicting protein structures with incredible speed and precision. This leap forward demonstrates how computational methods are poised to transform research in biology and hold much promise for accelerating the drug discovery process.”* > > *Arthur D. Levinson, PhD, Founder and CEO Calico, Former Chairman and CEO Genentech.* Helping guide the world on areas to focus are the United Nations Sustainable Development Goals (SDGs). At its heart the 17 SDGs are an urgent call for action by all countries – developed and developing – in a global partnership. They recognize that ending poverty and other deprivations must go hand-in-hand with strategies that improve health and education, reduce inequality, and spur economic growth – all while tackling climate change and working to preserve our oceans and forests. Creative Commons is one of the more interesting recent participants in attempting to solve this. They recently received funding for a [four-year, $4 Million Open Climate Campaign](https://creativecommons.org/2022/08/30/press-release-new-four-year-4-million-open-climate-campaign/?ref=openresearch.wtf#:~:text=Creative%20Commons,-August%2030%2C%202022&text=This%20grant%2C%20which%20builds%20on,promoting%20open%20access%20to%20research.) to open knowledge to Solve Challenges in Climate and Biodiversity in collaboration with SPARC. Similarly, on 23 November 2021, UNESCO’s[ Recommendation on Open Science](https://www.unesco.org/en/natural-sciences/open-science?ref=openresearch.wtf) was adopted by 193 Member States during the 41st session of the Organization’s General Conference. > “*This Recommendation outlines a common definition, shared values, principles and standards for open science at the international level and proposes a set of actions conducive to a fair and equitable operationalization of open science for all at the individual, institutional, national, regional and international levels*.” Recognizing the urgency of addressing complex and interconnected environmental, social and economic challenges for the people and the planet, including poverty, health issues, access to education, rising inequalities and disparities of opportunity, increasing science, technology and innovation gaps, natural resource depletion, loss of biodiversity, land degradation, climate change, natural and human-made disasters, spiralling conflicts and related humanitarian crises, In North America, Federal agencies are celebrating 2023 as a Year of Open Science, a multi-agency initiative across the federal government to spark change and inspire open science engagement through events and activities that will advance adoption of open, equitable, and secure science. This has already seen commitments from[ NASA](https://science.nasa.gov/open-science/transform-to-open-science?ref=openresearch.wtf),[ NIH](https://sharing.nih.gov/?ref=openresearch.wtf),[ NEH](https://www.neh.gov/?ref=openresearch.wtf) and 5 other agencies, as documented on the new[ open.science.gov](http://open.science.gov/?ref=openresearch.wtf) page. This all sounds great – so what can we do to help move things further, faster? #### **How do we as a society find areas to focus on?** With generalist repositories, there is now a home for every dataset – no matter the research field, or funding situation. However, we also know that subject specific repositories can provide subject specific guidance around metadata schemas and ways to make the files they generate useful. Re3data[ lists 2316 disciplinary repositories globally](https://www.re3data.org/metrics/types?ref=openresearch.wtf). Whilst this is encouraging, this self-reporting database also contains a few redundant repositories. Sustainability models for generalist repositories appear to be an easier problem to solve than subject specific repositories. We also now have the tools to identify what research areas are lacking in tools to help define subject specific metadata standards, in an actionable way. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee0feb24-3dfe-4a52-9a91-91d79589b2a2_904x852.png) Using [Dimensions.ai](https://www.dimensions.ai/?ref=openresearch.wtf), we can look at SDG categorisation of all datasets catalogued by DataCite. This provides a great snapshot of which SDG has good coverage when it comes to data publishing. To normalise these counts, we can look at the patterns for traditional published research papers. The graphs below highlight a similar pattern across papers and data. Health, climate, and energy dominate. This seems in line with the most urgent of the SDGs, however, by definition, they are all urgent – and subsequently we see there is not enough research being funded for issues like poverty, clean water, and gender equality. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bc352f-4cc3-4227-ab19-33f660b82f6c_1340x588.png) Digging deeper into the Dimensions dataset, we also see how generalist repositories, whilst providing a way for all researchers to publish datasets, may also be leading researchers to a “path of least resistance”. For example, Zenodo has more datasets that can be categorised as useful for the SDG of “climate action” than the World Data Centre for Climate. I am sure that my colleagues at Zenodo would prefer that these datasets ended up in a subject specific repository that is one layer closer to being interoperable and reusable than generalist repositories, such as Figshare and Zenodo. As generalist repositories, our responsibility is to ensure that the platforms we build can support FAIR (Findable, Accessible, Interoperable and Re-usable) data. We know that the I and R of FAIR require more curation. The solution we have for this is humans. Librarians and data curators who can facilitate making research data and associated metadata as rich as possible. As a result – Institutional data repositories are often a level above generalist repositories as they can offer a layer of said curation. Our [“State of Open Data” report](https://doi.org/10.6084/m9.figshare.21276984.v5?ref=openresearch.wtf) highlighted that when researchers need to publish their data, the academic publisher is their first port of call. As such, there is a huge opportunity for publishers and societies to define metadata schemas at the subject level. When it comes to repositories, academia and all subsequent stakeholders should: - As a collective, agree on and start using consistent thematic metadata - Improve metadata quality using computers or human curation I think the next 10 years are all about making that data useful for machines. This is how we do it. In the last 10 years, I have learned to understand the power of incentives in academia. At times I feel we are trying to trick researchers into doing the right thing. When it comes to the integrity of research, we must hold researchers accountable. When it comes to understanding why researchers sometimes need perverse incentives to be a good researcher, we need to look at ways to improve the dog-eat-dog pyramid scheme that academia can become. Culture change takes time. However, researcher adoption of data publishing is happening at a fast rate. We should and will continue to champion this action and those who lean into it. Data publishing is the biggest ROI in research today. It is truly an agent of change and I for one am excited that we have made it this far. Now funders have mandated data publishing writ large, I am also excited for the road ahead. ### Halfway to happiness — what the OSTP update means in the grand scheme URL: https://www.openresearch.wtf/halfway-to-happiness-what-the-ostp-update-means-in-the-grand-scheme/ Last updated: 2026-07-31T10:00:32.000Z It is very hard for any academic to argue with the concept of equitable access and equitable ability to publish research. The Office of Science and Technology Policy recently released [their new memorandum](https://www.whitehouse.gov/wp-content/uploads/2022/08/08-2022-OSTP-Public-Access-Memo.pdf?ref=openresearch.wtf) outlining Guidance to Make Federally Funded Research Freely Available Without Delay. It is already being ‘[Hailed as a Win for Innovation and Equity](https://www.whitehouse.gov/ostp/news-updates/2022/08/29/what-they-are-saying-white-house-federally-funded-research-guidance-hailed-as-a-win-for-innovation-and-equity/?ref=openresearch.wtf)’, which is exactly what we work for at Digital Science and more specifically, Figshare, in academic publishing. What has made this task so difficult is 350 years of old habits, dying hard. The incentive structure in academia has long been broken and attempts to change this, while seeming slow moving, are up to a pace not previously seen before. That said, we are still a long way away from what academic publishing would look like if it were invented today. Figshare is 10 years old this year, and looking back at the 10 years in the industry we can see the monumental change that Open Access (OA) publishing has brought about. There has been a move from 70% of all publishing being closed access to 54% being OA in a decade. This is unstoppable momentum. What the new OSTP memo will do is accelerate the march to 100% OA (or as close as is possible). ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9746c518-1aad-4e7f-97d3-a8bb7a102e48_1063x533.png) Another side of the story is Green vs Gold OA over the same time. In 2011, 55% of OA papers were available via Green OA, compared with 70% of OA papers being Gold OA in 2021\. This means that the speed at which we are moving to OA over closed access is the same speed at which the market is deciding to publish Gold OA papers over Green OA. This move to Gold OA and a need for funding may not be in line with the goal of equitable access and equitable ability to publish research. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1c263e4-bcaa-4e31-be88-b9508c2e9382_1063x469.png) One of the most important changes is the requirement for immediate Open Access without any embargoes. This could be beneficial for all flavours of Open Access, including preprints, which can be classified as Green OA if properly managed. This improves both access and equity across research, as pointed out by Alison Mudditt, CEO of PLOS on Twitter. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F31a0edf8-c393-4d03-95dc-a8acba612317_1182x502.png) [https://twitter.com/alison\_mudditt/status/1562870887841288193](https://twitter.com/alison%5Fmudditt/status/1562870887841288193?ref=openresearch.wtf) The memo also, importantly, requires open data. When Figshare started there were no free generalist repositories to post datasets or other non-traditional research outputs (NTROs). When I think of academic data, it doesn’t mean only spreadsheets: it is the files generated that are needed to back up the conclusions in the peer reviewed papers. Figshare now houses over 6 million of these outputs. The OSTP memo has the following requirements when it kicks in in 2026: > *Immediate public access to the underlying data is required. The policy requires that scientific data underlying peer-reviewed scholarly publications be made publicly available at the time of publication. Federal agencies should develop approaches and timelines for sharing other federally funded scientific data that are not associated with peer-reviewed scholarly publications.* As Clarke and Eposito point out in [their overview of the memo](https://www.ce-strategy.com/the-brief/zero-embargo/?ref=openresearch.wtf): “This is (presently) an unfunded mandate. No additional funding to cover the cost of publication or data availability is referenced in the memo, though section 3d of the memo directs federal agencies to “allow researchers to include reasonable publication costs and costs associated with submission, curation, management of data, and special handling instructions as allowable expenses in all research budgets.” As part of the internal Figshare celebrations, I presented to the team my thoughts on how the space had moved on when it comes to academic research publishing, from the view of traditional peer reviewed papers and data with regard to equitable access and equitable ability to publish research 10 years ago, now, and 10 years into the future. Here were my thoughts: ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde15811-2f34-41c8-a0ee-7d4da74cde54_998x490.png) The above table reflects my thoughts prior to the new memo being published. It has bolstered my hope that a future with equitable access and equitable ability to publish research is clearly in reach. The roadblocks and open questions of how to get there are still prominent in a few areas. But there are a lot fewer questions now than there were 10 years ago, or back in the 80s when the first dream of Open Access publishing was born. The new memo shows the power of constantly pressuring in the right direction. We have a great opportunity in the data space with a constant pressure on funders to require FAIR data publishing. As is the tradition in academic publishing, we can continue to stand on the shoulders of giants, specifically the decades of work done by early open access advocates and SPARC. The speed at which dissemination of research is improving itself is accelerating. At Figshare, we welcome this new memo with open arms and head into the future more inspired about achieving the goal of open research. ### Why fast but good publishing matters URL: https://www.openresearch.wtf/why-fast-but-good-publishing-matters/ Last updated: 2024-02-05T15:55:18.000Z A [large number of funders](https://www.dcc.ac.uk/guidance/policy/overview-funders-data-policies?ref=openresearch.wtf) around the world now mandate publishing data needed to reproduce findings from the research they fund – at the same time as the paper is published. This is a change for academics. We have found through our [State of Open Data](https://doi.org/10.6084/m9.figshare.c.4046897.v4?ref=openresearch.wtf) reports that the majority of researchers are new to the concept of licensing content when they publish it. We also see lack of knowledge about how to describe datasets – particularly the title of the dataset. A rule of thumb is to describe a dataset in the same way that you would title a paper with a descriptive, meaningful title describing the research question, method, and/or finding. The acceleration of public awareness around preprints and open data has occurred recently in line with the huge public interest in finding a treatment for COVID-19\. This is nicely summed up by Digital Science CEO, Daniel Hook, in his recent whitepaper [‘How COVID-19 is Changing Research Culture’](https://doi.org/10.6084/m9.figshare.12383267.v1?ref=openresearch.wtf). “The response has been immediate and intensive. Indeed, the research world has moved faster than many would have suspected possible. As a result, many issues in the scholarly communication system, that so many have been working to improve in recent years, are being highlighted in this extreme situation.” Our recent work with the [National Institutes of Health (NIH) on a pilot generalist repository](https://datascience.nih.gov/data-ecosystem/exploring-a-generalist-repository-for-nih-funded-data?ref=openresearch.wtf) for NIH-funded researchers highlighted some of the early issues that researchers face when publishing their data. This leads to a lot of poorly described and unintentionally obfuscated datasets. This can be very problematic. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83290357-c73f-44e2-9035-132523346da2_606x337.png) Consider the infographic above, where several labs are working to find a new treatment. By building on top of the work that has been done by one lab, another lab can advance at a faster rate. Add in multiple labs and this productivity is compounded. However, when reusing the ideas and data produced by another lab (represented as a dotted line above), it doesn’t take much for the process to fall apart. If the data is not open and [FAIR](https://www.nature.com/articles/sdata201618?ref=openresearch.wtf), researchers will not be able to use it as a stepping stone in the hunt for a cure or treatment. This inevitably pushes back the time to find said treatment. If we’re using this analogy to represent a COVID-19 vaccine, pushing back the time potentially results in more deaths. Unfortunately, a lot of researchers are still not making their data available in an open and FAIR manner. A quick search for the text string [‘data available upon request’](https://app.dimensions.ai/analytics/publication/overview/timeline?search%5Ftext=%22data%20available%20upon%20request%22&search%5Ftype=kws&search%5Ffield=full%5Fsearch&year%5Ffrom=1971&year%5Fto=2020&ref=openresearch.wtf) gives us over 4,000 papers. [‘Data available upon reasonable request’](https://app.dimensions.ai/discover/publication?search%5Ftext=%22data%20available%20upon%20reasonable%20request%22&search%5Ftype=kws&search%5Ffield=full%5Fsearch&ref=openresearch.wtf) gives us nearly 100 results, of which nearly one third were published in 2019\. A [recent study also highlighted](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0194768&ref=openresearch.wtf#sec005) where “Other statements suggested data were openly available, but failed to give enough information for readers to actually locate the data” **Number of papers published that include the phrase** [**‘data available upon request’**](https://app.dimensions.ai/analytics/publication/overview/timeline?search%5Ftext=%22data%20available%20upon%20request%22&search%5Ftype=kws&search%5Ffield=full%5Fsearch&year%5Ffrom=1971&year%5Fto=2020&ref=openresearch.wtf) **over time.** ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa853fc02-0e64-4c85-9367-10257c668e87_599x335.png) Fortunately, a global effort by funders such as the NIH to encourage data sharing and publishers to require data availability statements means that most researchers will soon be required to make their data publicly available as much as is feasible. Data available upon request will not suffice for most cases. Making research available rapidly is important, but human checks on the content are too, as it needs to be not just openly available but also discoverable and reusable. Both preprints and data are amongst the new type of formats in which academic findings are disseminated. Whilst every preprint published on well known preprint platforms, such as ChemrXiv, TechrXiv and bioRxiv, have a basic human check – the majority of data published in generalist repositories does not. For me, there are several tiers, which can be mapped back to the FAIR data principles. This list is not intended to be exhaustive, but indicative of the complexity and effort required at each stage. Data with no checks that can be useful - Academic files and metadata are available on the internet in repositories that follow [best practice norms](https://www.coar-repositories.org/news-updates/coar-sparc-response-to-the-ostp-draft-desirable-characteristics-of-repositories-managing-data/?ref=openresearch.wtf). Top level check to make data Findable and Accessible - The metadata is sufficient to be discoverable through a google search - Policy compliant checks – the files have no PII, are under the correct license - Appropriate use of PIDs and metadata standards Interoperable and Reusable - Files are in an open, preservation-optimised format - Subject specific metadata schemas and file structures are applied in compliance with community best practice - Forensic data checks for editing, augmenting - Re-running of the results to ensure replicability A search for the term dataset on the ‘free’ and non-curated figshare.com highlights a lack of descriptive titles in the screenshot below. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e58cba-f9e2-4266-b38a-ed85ec858b4e_612x462.png) Our pilot with the NIH involved a data librarian on the Figshare team checking the metadata quality and working with the submitting researchers to make the datasets Findable and Accessible. The following checks and enhancements were carried out: - Files match the description, can be opened, and are documented. - Item type is appropriate for the NIH Figshare Instance. - No obvious personally identifiable information is found in data or metadata and submitter has affirmed no PII is included. - Metadata sufficiently describes the data or links to resources that further describe it. - Embargoes are used appropriately. - An appropriate license has been applied. - NIH funding is specified and linked (if possible). - Related publications are linked (when applicable). ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673a662d-5bde-4fb1-afb6-704ac73d89ff_618x213.png) As expected, these checks resulted in much more detailed metadata, when compared to datasets on figshare.com that were not checked or enhanced by a curation expert. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58c8323e-605e-4c2d-a9ec-5f1bba819193_1884x864.png) As raising awareness of new platforms is always tough (marketing is important) and takes time for adoption, we were aware that some NIH-funded researchers would continue to make use of the free figshare.com – even though they had the opportunity for free expert help with publishing on NIH Figshare. This actually proved to be beneficial in providing a comparison to consider the effect of metadata enhancement in datasets published in the NIH Figshare Instance. We found preliminary evidence that datasets that had been checked by a data librarian got more downloads and many more views (2.5x) than those that hadn’t been checked. We also saw that the file size of the published datasets was much larger in the NIH Figshare Instance (8x greater) suggesting that there is a gap in the repositories space for datasets 100-500GB. Taken together, these findings highlight the need for free-for-researcher generalist repositories, with human checks on the metadata to ensure FAIR-er, more impactful, and more reproducible research. If we are to enter the world of fast and good academic publishing, we need a dataset curation model that scales. ### Academic Research Data. Is it being cited? URL: https://www.openresearch.wtf/academic-research-data-is-it-being-cited/ Last updated: 2026-07-31T10:00:04.000Z Recently, we partnered with our sister company [Dimensions](https://www.dimensions.ai/?ref=openresearch.wtf) to offer daily citations updates on any object within any Figshare instance. Our [State of Open Data](http://figshare.com/collections/State%5Fof%5FOpen%5FData/4046897?ref=openresearch.wtf) report that comes out each year has consistently highlighted that increased impact and visibility is the number one driver ([78% on respondents](https://knowledge.figshare.com/articles/item/state-of-open-data-2019?ref=openresearch.wtf) value data citations as much or more than a paper citation) are the number one driver in motivating researchers to publish their datasets and the accompanying files their research generates. This also allows us to begin to investigate patterns in the citation data that may help us understand the motivations of researchers to reuse content both generally and in specific fields. Data citations still struggle to find their way into reference lists, instead being cited inline within the article itself or in data availability statements. As such, we have partnered with our sister company Dimensions to identify these citations in the full text of articles. This ensures that researchers, and our clients, are aware of all of the citations to their research. To start our investigations, we looked at the top 100 cited outputs on figshare.com, the the public data repository that is free for researchers to upload research to. In the future, it may also be interesting to compare rates of citation between institutional outputs that have been curated or checked, and those which have not. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97ab5ccb-7937-40f6-9e84-dfce39e352c3_1024x417.png) There is a long tail of outputs with 1 citation, which we can assume to be a citation of the output in the original paper describing the work by the same authors. Similarly, we can infer that outputs with more than 1 citation, cited in publication(s) by different authors, have been reused. Here are some other interesting preliminary findings. ### 4 out of the top 10 are software with a mean citation count of 50.5! The code repositories have a few things in common. - Each has a README file - Each has multiple versions - Each has lots of metadata associated with it. Sometimes this description is includeddirectly on the landing page, as is the case with [PIVlab – Time-Resolved Digital Particle Image Velocimetry Tool for MATLAB](https://doi.org/10.6084/m9.figshare.1092508?ref=openresearch.wtf) ([https://doi.org/10.6084/m9.figshare.1092508](https://doi.org/10.6084/m9.figshare.1092508?ref=openresearch.wtf)). For others, we find context to the work on other sites that link directly to the output, which are picked up by the Altmetric Explorer tool. Whether it be a [Wikipedia article](https://en.wikipedia.org/?curid=56424654&ref=openresearch.wtf) describing the software, as is the case for Common Workflow Language, v1.0 ( [https://doi.org/10.6084/m9.figshare.3115156.v2](https://doi.org/10.6084/m9.figshare.3115156.v2?ref=openresearch.wtf)) or blog posts that describe the [The graph-tool python library](https://www.r-bloggers.com/benchmark-of-popular-graph-network-packages-v2/?ref=openresearch.wtf) ([https://doi.org/10.6084/m9.figshare.1164194.v14](https://doi.org/10.6084/m9.figshare.1164194.v14?ref=openresearch.wtf)) or [PyMKS: Materials Knowledge System in Python](https://materials-informatics-class-fall2015.github.io/MIC-grain-growth/2015/12/06/final-post/?ref=openresearch.wtf) ([https://doi.org/10.6084/m9.figshare.1015761.v2](https://doi.org/10.6084/m9.figshare.1015761.v2?ref=openresearch.wtf)) Interestingly, there are 12 outputs labelled by authors as software in the top 100, despite software making up just 1.6% of outputs on figshare.com. This suggests that software can be highly reused and the impact it is having is not being captured in traditional credit mechanisms. ![](https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06a906cd-77d0-42ad-a003-21220c105ea8_1024x791.png) ### 32 out of the top 100 are datasets As of today, outputs labelled by authors as datasets make up 30% of outputs on Figshare.com. Therefore, finding that 32% of the top 100 cited outputs are datasets seems logical. Obviously with the many different types of outputs that researchers want credit for, there are some that lend themselves to reuse and traditional citation metrics better than others. Notably, posters and presentations rarely get cited. But this does not mean that it isn’t happening. There is 1 presentation and 2 posters in the 100 most cited outputs. The [author](http://figshare.com/authors/Lorena%5FA%5FBarba/97553?ref=openresearch.wtf) of said presentation has picked up multiple citations over multiple different outputs, suggesting that the credibility of the author and their work is a key factor in driving impact. Soon Figshare’s newly launched [faceted search functionality](http://figshare.com/blog/Announcing%5Fa%5FNew%5FFaceted%5FSearch%5FPage%5Ffor%5FFigshare/569?ref=openresearch.wtf) will also include the ability to filter search results by citation count and Altmetric score in any Figshare portal. We’re excited about the new ways to encourage researchers to share all of their research outputs and make sure they get credit for the work. We look forward to continuing to investigate patterns of data reuse and working with the community on methods to make all the products of research more reusable. ### Academic Data Curation: Who checks? Who Pays? How Much? URL: https://www.openresearch.wtf/academic-data-curation-who-checks-who-pays-how-much/ Last updated: 2026-07-31T10:01:22.000Z In the last decade, we’ve seen repeated reports of highlighting a lack of reproducibility and replicability in published academic research results. This has led to institutional, publisher and most significantly funder mandates for research data being made openly available at the point of publication of the paper. Over 10 years, we have seen the ground swell of these policies and mandates. This has led to large amounts of data files with associated metadata being made available, in many data repositories around the world. The birth of the concept of FAIR (Findable, Accessible, Interoperable, Reusable) data in 2014([1](https://doi.org/10.1038/sdata.2016.18?ref=openresearch.wtf)) has helped unite global initiatives with a broad, common goal. We even have been hearing from researchers themselves that this is a change they want([2](https://doi.org/10.6084/m9.figshare.9980783.v2?ref=openresearch.wtf)). Doing your research becomes so much easier when you can build on top of the raw research that has gone before and not just the summarised findings in the form of a conclusions section of a peer reviewed article. So, what are the next steps in ensuring all of this information can be turned into knowledge, to be read by fellow researchers, or to train AI models? Several recent, high profile publications have highlighted some of these problems([3](https://doi.org/10.1038/d41586-020-00287-y?ref=openresearch.wtf)). As data publishing becomes the norm, questions are surfacing about who should be responsible for checking these datasets and their associated metadata, and to what level should they be examining the research. This is in addition to basic technological needs, such as open APIs and complying with web accessibility standards. For me there are several tiers, which can be mapped back to the FAIR data principles. This list is not intended to be exhaustive, but indicative of the complexity and effort required at each stage Data with no checks that can be useful 1. Academic files and metadata are available on the internet in repositories that follow [best practice norms](https://www.coar-repositories.org/news-updates/coar-sparc-response-to-the-ostp-draft-desirable-characteristics-of-repositories-managing-data/?ref=openresearch.wtf). Top level check to make data Findable and Accessible 1. The metadata is sufficient to be discoverable through a google search 2. Policy compliant checks – the files have no PII, are under the correct license Interoperable and Reusable 1. Files are in an open, preservation-optimised format 2. Subject specific metadata schemas are applied in compliance with community best practice 3. Forensic data checks for editing, augmenting 4. Re-running of the results to ensure replicability One important discussion topic is understanding that post publication metadata curation can help improve datasets over time, either by humans or machines – with the caveat that researchers may be intentionally obfuscating the research, or providing so little descriptive metadata that the dataset will always be useless. As we move from through checks 1-7, the human curation, technical expertise and time taken increases. As such, so does the cost. Scalability of costs needs to be thought about([4](https://doi.org/10.1038/d41586-020-00505-7?ref=openresearch.wtf)). We cannot rely on volunteers to take all research to levels 2 and 3. The research data community needs to come up with a plan to move to fully FAIR data by 2030 (points 4+ above), with a full understanding of how each of the steps above is carried out and by whom. The research publishing system works. We get new drugs and new breakthrough discoveries every year. The goal of FAIR research data is to optimise this, to make use of machine learning, AI and all human knowledge to get these breakthrough discoveries, treatments for pandemics and improve our understanding of climate change faster. References 1\. Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016). [https://doi.org/10.1038/sdata.2016.18](https://doi.org/10.1038/sdata.2016.18?ref=openresearch.wtf) 2\. Science, Digital; Fane, Briony; Ayris, Paul; Hahnel, Mark; Hrynaszkiewicz, Iain; Baynes, Grace; et al. (2019): The State of Open Data Report 2019\. figshare. Report. [https://doi.org/10.6084/m9.figshare.9980783.v2](https://doi.org/10.6084/m9.figshare.9980783.v2?ref=openresearch.wtf) 3\. Nature 578, 199-200 (2020) [https://doi.org/10.1038/d41586-020-00287-y](https://doi.org/10.1038/d41586-020-00287-y?ref=openresearch.wtf) 4\. Nature 578, 491 (2020) [https://doi.org/10.1038/d41586-020-00505-7](https://doi.org/10.1038/d41586-020-00505-7?ref=openresearch.wtf)