# OpenResearch.wtf > FAIR data, metadata and research infrastructure for an AI-driven science. Essays and working prototypes on what machines need before they can discover. Public Ghost content for AI and LLM tooling. Use `/llms-full.txt` for consolidated page and post context. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages - [About this site](https://www.openresearch.wtf/about.md) - Mark Hahnel on the metadata, identifiers and infrastructure that decide what machines can find and reuse — plus four working prototypes. ## Posts - [Whoever writes the paper, the tools have to be trustworthy](https://www.openresearch.wtf/whoever-writes-the-paper-the-tools-have-to-be-trustworthy.md) - If we are going to hand more of the writing, the citing and eventually the analysis to machines, then the tools doing that work have to verify before they assert, keep provenance attached, and leave ownership with the human whose name is on the paper. - [Do we actually need to write research papers at all?](https://www.openresearch.wtf/do-we-actually-need-to-write-research-papers-at-all.md) - If we are going to lean on the creativity and the cross-domain thinking of human researchers, then surely that should be where the human time goes. - [BioSingularity. The substrate determines the ceiling](https://www.openresearch.wtf/biosingularity-the-substrate-determines-the-ceiling.md) - Biology becoming computationally tractable depends less on model size than on the data substrate underneath. The ceiling is infrastructure. - [Ground the model, or it invents the evidence](https://www.openresearch.wtf/ground-the-model-or-it-invents-the-evidence.md) - arXiv now bans authors for unchecked LLM output. The real fix is trusted, curated, subject-specific models rather than more policy. - [The new era of going “Fast and Far” in research](https://www.openresearch.wtf/the-new-era-of-going-fast-and-far-in-research.md) - A Nobel Prize in Chemistry went, in effect, to software. The old trade-off between small fast teams and large slow ones has stopped holding. - [What parts of the academic knowledge creation & dissemination pipeline can we automate with AI?](https://www.openresearch.wtf/what-parts-of-the-academic-knowledge-creation-dissemination-pipeline-can-we-automate-with-ai.md) - Four working prototypes testing which parts of the research pipeline, from writing to review to publishing, can actually be automated. - [OpenAccess.ai: How Cheap Can We Make Academic Publishing?](https://www.openresearch.wtf/openaccess-ai-how-cheap-can-we-make-academic-publishing.md) - Article processing charges run to thousands per paper. OpenAccess.ai tests how much of the editorial stack can genuinely be automated. - [Preprints.ai: How Much of Peer Review Can We Automate?](https://www.openresearch.wtf/preprints-ai-how-much-of-peer-review-can-we-automate.md) - A large share of peer review is mechanical checking. Preprints.ai tests how much of it can be scored automatically, at preprint scale. - [OpenScience.ai: Using Open Data to Generate Research That Doesn't Yet Exist](https://www.openresearch.wtf/openscience-ai-using-open-data-to-generate-research-that-doesnt-yet-exist.md) - What claims are already latent in ClinVar, GTEx, STRING and Open Targets that nobody has written up yet? OpenScience.ai explores that gap. - [FAIRdata.ai: Making Open Data FAIR-er](https://www.openresearch.wtf/fairdata-ai-making-open-data-fair-er.md) - Most repository data is technically FAIR and practically unusable. FAIRdata.ai is an experiment in closing that gap automatically. - [Moving from "share data because you should" to "share data because you'll get something back."](https://www.openresearch.wtf/moving-from-share-data-because-you-should-to-share-data-because-youll-get-something-back.md) - Allen AI's AutoDiscovery generated its own hypotheses from an open dataset. That return on sharing is exactly why I founded Figshare. - [Machine-First FAIR: Realigning Academic Data for the AI Research Revolution](https://www.openresearch.wtf/machine-first-fair-realigning-academic-data-for-the-ai-research-revolution.md) - FAIR treats humans and machines as equal priorities. It should not. Prioritising machines is the faster route to human benefit. - [Academia has a new preprints problem](https://www.openresearch.wtf/academia-has-a-new-preprints-problem.md) - Researchers are pulling the LLM slot-machine lever for novel physics, then wrapping the output in LaTeX. What that looks like from inside a preprint platform. - [Have we already hit the peer review breaking point?](https://www.openresearch.wtf/have-we-already-hit-the-peer-review-breaking-point.md) - Notes from the Royal Society's Future of Scientific Publishing: what does review look like when every paper is AI-generated or co-authored? - [AI Needs New Facts – The Value of Novel Scientific Research](https://www.openresearch.wtf/ai-needs-new-facts-the-value-of-novel-scientific-research.md) - Hearing Demis Hassabis at SXSW London: if models train on what already exists, genuinely novel experimental results become the scarce input. - [From Gold to Diamond: Is Equitable Open Access Still a Mirage?](https://www.openresearch.wtf/from-gold-to-diamond-is-equitable-open-access-still-a-mirage.md) - Gold Open Access growth is slowing, and Hybrid is filling the gap rather than Diamond. What the numbers say about equitable publishing. - [The Perpetual Research Cycle: AI's Journey Through Data, Papers, and Knowledge](https://www.openresearch.wtf/the-perpetual-research-cycle-ais-journey-through-data-papers-and-knowledge.md) - Testing ChatGPT, Gemini and Claude at turning datasets into papers, and what a genuinely self-feeding research cycle would require. - [How is every country/funder/ institution doing at data sharing?](https://www.openresearch.wtf/how-is-every-country-funder-institution-doing-at-data-sharing.md) - An interactive app joining the Data Citation Corpus to Dimensions, ranking countries, funders and institutions on how well they link data. - [The Role of Foundations and Technology Companies in Fuelling Optimism in Academic Research](https://www.openresearch.wtf/the-role-of-foundations-and-technology-companies-in-fuelling-optimism-in-academic-research.md) - As North American grant funding wobbles, Evo-2 and Google's AI co-scientist show where research optimism is now coming from instead. - [DataCite DOI prevalence in the published literature by country](https://www.openresearch.wtf/datacite-doi-prevalence-in-the-published-literature-by-country.md) - Combining the Make Data Count Citation Corpus with Dimensions to see which countries actually link to datasets from their published papers. - [Types of AI Agents and Their Applications in Academic Research](https://www.openresearch.wtf/types-of-ai-agents-and-their-applications-in-acaresearch.md) - Model-based, goal-based, learning and utility agents: a practical taxonomy of AI agents and where each one fits in academic research. - [Have we all forgotten about SciHub?](https://www.openresearch.wtf/have-we-all-forgotten-about-scihub.md) - Sci-Hub's founder proposes a colour-coded taxonomy that absorbs Black OA into legitimate Open Access. A more interesting idea than it sounds. - [Are we ready to start rewarding researchers for Open data?](https://www.openresearch.wtf/are-we-ready-to-start-rewarding-researchers-for-open-data.md) - The State of Open Data 2024 pairs what researchers say with what they do, using Dimensions, data availability statements and the Citation Corpus. - [Innovating Academia From The Outside - The Web3 Solution to Academia’s Peer Review Problem](https://www.openresearch.wtf/innovating-academia-from-the-outside-the-web3-solution-to-academias-peer-review-problem.md) - ResearchHub pays for peer review in crypto. That terrifies plenty of researchers, but the rate of innovation on the platform is hard to ignore. - [The Data Citation Corpus - tracking NIH funded open academic data](https://www.openresearch.wtf/the-data-citation-corpus-tracking-nih-funded-open-academic-data.md) - Joining the Wellcome-funded Data Citation Corpus to Dimensions to track how well the NIH open data policy is actually working in practice. - [Some Gold Open Access (OA) Article Processing Charges (APC) data](https://www.openresearch.wtf/some-apc-data.md) - A new Gold Open Access article processing charge dataset, and what it says about whether we can publish open access faster and cheaper. - [OpenResearch WTF April 2024](https://www.openresearch.wtf/openresearch-wtf-april-2024.md) - April 2024: the Barcelona Declaration on Open Research Information, DataCite's full public metadata release, and the month's open research news. - [Open Access: Mo money, mo problems](https://www.openresearch.wtf/open-access-mo-money-mo-problems.md) - From discovering PLOS ONE in a stem cell lab to today's article processing charges: how Open Access got expensive, and what that cost us. - [OpenResearch WTF - March 2024](https://www.openresearch.wtf/openresearch-wtf-march-2024.md) - March 2024: the Gates Foundation drops APCs and mandates preprints, plus the first State of Open Data supplementary report. - [OpenResearch WTF - February 2024](https://www.openresearch.wtf/openresearch.md) - February 2024: eLife's new model one year on with 6,200 submissions and 1,300 reviewed preprints, plus the launch of TL;DR Shorts. - [Making poorly described data FAIR-er using GenAI](https://www.openresearch.wtf/making-poorly-described-data-fair-er-using-genai.md) - A small experiment: can ChatGPT-4 improve the metadata, and therefore the FAIR-ness, of badly described datasets already published openly? - [Who benefits when, from FAIR data? Part 3 - The Public](https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-3.md) - Part 3: the public fund research through taxes but cannot read most of it. What open academic data owes the people paying for it. - [Who benefits when, from FAIR data? Part 2 – Machines](https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-2-machines.md) - AI and machine learning can be used to create more detailed, FAIR-er datasets to be consumed by the machines. - [All of the knowledge for all of the machines](https://www.openresearch.wtf/all-of-the-knowledge-for-all-of-the-machines.md) - Two slow revolutions in publishing, blurred peer review and universal Open Access, decide what machines can legally read and actually trust. - [Who benefits when, from FAIR data? Part 1 – Researchers](https://www.openresearch.wtf/who-benefits-when-from-fair-data-part-1-researchers.md) - Part 1: seven years on from the FAIR principles, what do researchers themselves actually get back from making their data reusable? - [Academic data publishing is the biggest ROI in research today](https://www.openresearch.wtf/academic-data-publishing-is-the-biggest-roi-in-research-today.md) - A decade of running Figshare says the big win was simply getting data onto the internet. The next ten years are about making it usable. - [Halfway to happiness — what the OSTP update means in the grand scheme](https://www.openresearch.wtf/halfway-to-happiness-what-the-ostp-update-means-in-the-grand-scheme.md) - The OSTP memo on federally funded research is real progress built on decades of SPARC advocacy. It is also only half the journey. - [Why fast but good publishing matters](https://www.openresearch.wtf/why-fast-but-good-publishing-matters.md) - If we are to enter the world of fast and good academic publishing, we need a dataset curation model that scales. - [Academic Research Data. Is it being cited?](https://www.openresearch.wtf/academic-research-data-is-it-being-cited.md) - Daily citation updates across Figshare show which research outputs actually attract citations, and why datasets still trail papers badly. - [Academic Data Curation: Who checks? Who Pays? How Much?](https://www.openresearch.wtf/academic-data-curation-who-checks-who-pays-how-much.md) - The research publishing system works. We get new drugs and new breakthrough discoveries every year. The goal of FAIR research data is to optimise this. ## Optional - [RSS Feed](https://www.openresearch.wtf/rss/) - [Sitemap](https://www.openresearch.wtf/sitemap.xml) - [Full content of pages and posts](https://www.openresearch.wtf/llms-full.txt)