Opinion

Your Lab Notebook Is a Data Product

When a licensing partner’s diligence team asks for the data behind a preclinical programme, many research organisations find the evidence spread across notebooks, instrument PCs, and shared drives. Four typical diligence requests show what research data management has to deliver, and where researchers will meet the change halfway.
Your Lab Notebook Is a Data Product — hero image

When your corporate development team has a term sheet for out-licensing a preclinical programme, the partner’s diligence list is likely to run to several hundred lines. While most will be routine, a handful will land on the research organisation, and they tend to look like these four, paraphrased from the kind that appear in life sciences licensing and M&A diligence:

  • Provide the primary data behind the in vivo efficacy figures in the data room summary, including raw instrument output.

  • Confirm the analysis can be reproduced: methods, software and versions used, and any exclusions applied.

  • For each key experiment, identify who generated the data, when, and under which employment, CRO, or collaboration agreement, with the supporting records.

  • List any third-party data in the package (CROs, academic collaborators, material transfer agreements, licensed datasets, and any government funding) and the terms under which it can be used and transferred.

The science is rarely what slows these requests down. The effort goes into answering them in a way the partner will rely on when it comes to price, because anything it can’t verify is likely to come back as a warranty or a price adjustment. The answers usually exist, spread across electronic lab notebooks (ELNs), instrument PCs, a scientific data management system (SDMS), departmental shares, CRO portals, and the memories of scientists, some of whom now work on other programmes.

Research data in most pharma and biotech organisations is kept in several ways at once: by individual scientists in the way that suits their work, in shared drives and an ELN under lab-level rules, and, less often, as versioned, described, and owned datasets built for people who weren’t there when they were generated. Diligence tests all of them, because it asks the questions day-to-day research rarely does.

The raw data request

For a typical in vivo efficacy study, the raw data is mostly where the team left it. Caliper and body-weight readings sit in a spreadsheet, imaging files on the instrument PC or an external drive, and the processed numbers in a workbook named in a way that made sense at the time. The ELN entry records the protocol and attaches the final figure as an image. An SDMS may have captured the imaging output, but nothing records which files fed which figure, so the link between the summary in the data room and the data underneath it depends on folder conventions and on the people who remember them.

What answers this request is an identifier for the study, with the raw output, the processed data, and each figure linked to one another and to the ELN entry that describes the experiment. With that in place, the response is an export of a defined set of records to the data room, and nothing has to be assembled for the occasion.

The reproducibility request

This is often the request that takes longest. The FAIR principles, a widely adopted framework for research data stewardship, ask for data to be findable, accessible, interoperable, and reusable, by machines as well as people, and for data and metadata to be associated with detailed provenance (Wilkinson et al., 2016, R1.2). For an analysis, provenance means more than the files: which script, which version, which parameters, and which animals were excluded and why.

The method is usually well written up in the ELN. The code that ran it is more often in an analyst’s personal repository, and exclusion decisions made in a lab meeting may be recorded only as rows deleted from a spreadsheet. Reproducing the result then means finding the analyst, recovering the script version, and reconstructing the reasoning for the exclusions from meeting notes.

A research data platform built for this stores the script and its version with the output it produced, records parameters as data, and logs each exclusion with its reason. This is the same engineering any production data platform uses, applied to lab data.

The chain-of-title request

This request is about chain of title. A partner wants to know that the organisation owns what it’s licensing, which means knowing who generated each key dataset and under which agreement: an employee’s contract, a CRO’s statement of work, or a collaboration agreement with its own terms on ownership.

The lab notebook’s role in research IP has changed. Since the US moved to a first-inventor-to-file system for applications filed on or after 16 March 2013, a dated notebook entry no longer wins a priority contest against someone who filed first (USPTO, n.d.). Notebooks still matter as evidence of who conceived an invention, in disputes with collaborators, in derivation proceedings, which must be supported by substantial evidence (USPTO, n.d.), and for patents filed under the earlier rules that remain in force. Inventorship is a separate question from who generated the data, and a diligence team will usually ask about it separately.

The NIH’s data management and sharing policy, which applies to research NIH funds or conducts, draws a line that’s useful well beyond its scope. It defines scientific data as data of sufficient quality to validate and replicate research findings, and excludes laboratory notebooks, preliminary analyses, and communications with colleagues from that definition (NIH, n.d.). Many organisations haven’t drawn that line internally, so data gets pasted into notebook entries and notebook pages get photographed into data folders. Keeping them linked but distinct, with the ELN entry recording who did what and why and the dataset recording what was produced and by which instrument or provider, lets this request be answered with a query across the two records. In GLP studies, original observations recorded in a notebook are themselves raw data under FDA rules (FDA, n.d., §58.3(k)), so the split applies mainly to discovery and other non-GLP work, which is where most efficacy data sits.

The third-party data request

The last request can change the deal terms. A preclinical package often includes data generated by CROs, material received under transfer agreements, academic collaborations, and licensed third-party datasets, each under its own agreement. FAIR asks for data to be released with a clear and accessible data usage licence (Wilkinson et al., 2016, R1.1), and that’s a sensible standard inside a company as well as outside it. Data from an NIH-funded academic collaborator may also come with sharing commitments made in that collaborator’s data management plan (NIH, n.d.).

Notebooks and shared drives rarely carry those terms. The CRO’s study report is in the archive, and the contract that says who owns the underlying data is with legal. When source and usage terms are carried as metadata, the package can be filtered to what’s transferable before the data room opens. FAIR’s principle that metadata should stay accessible even when the data itself is no longer available (Wilkinson et al., 2016, A2) matters here as well, because an agreement may require data to be returned or deleted while the record that it existed has to remain.

This is the argument Telemetry Is Becoming the Business made about operational data, applied to the bench: a by-product of the work that needs an owner, a contract, and a structure once someone outside the lab relies on it. Sakura’s Data & AI practice builds research data platforms on top of the ELN and instrument estate a lab already runs, which is usually where the pharma data architecture work starts.

What works at the bench

None of this works unless scientists use it, and research culture has good reasons for its habits. The line tends to fall in predictable places.

Capture that happens without the scientist is usually welcome. Instruments that write raw output straight to a managed store, with the run’s metadata attached, remove work. So does an ELN template that pulls sample and instrument identifiers from the systems that already know them. Anything that removes a manual copy step tends to be adopted quickly.

Structure is accepted when it pays the scientist back. Someone who can find an older assay dataset in seconds, or rerun a colleague’s analysis without a meeting, sees the point of identifiers and versioning. A form asking for fifteen metadata fields for someone else’s benefit gets skipped.

What doesn’t work is a governance process that slows exploratory work. Most experiments are exploratory, and many are superseded, so treating every early run as a curated asset costs effort the science can’t spare. The workable approach is for automated raw capture to run for every experiment, and for curation, review, and ownership to start at a defined promotion point: when a result is going to be presented, included in an invention disclosure, or built on. The notebook stays the scientist’s space, and the data product is what the organisation builds from the experiments that matter.

That promotion step, with its integrations and access rules, has to be run as instruments, partners, and programmes change. Our Managed Services team keeps it running for research organisations. Done well, the scientist notices fewer copy steps, and the diligence team notices nothing at all.

References

National Institutes of Health (NIH), n.d. Data Management and Sharing Policy Overview. NIH Grants and Funding. Available at: https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/policy-overview [Accessed 5 October 2026].

United States Patent and Trademark Office (USPTO), n.d. America Invents Act (AIA) Frequently Asked Questions. USPTO. Available at: https://www.uspto.gov/patents/laws/america-invents-act-aia/america-invents-act-aia-frequently-asked [Accessed 5 October 2026].

US Food and Drug Administration (FDA), n.d. 21 CFR 58.3: Definitions (Good Laboratory Practice for Nonclinical Laboratory Studies). Electronic Code of Federal Regulations. Available at: https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-58/subpart-A/section-58.3 [Accessed 5 October 2026].

Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J. et al., 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. Available at: https://doi.org/10.1038/sdata.2016.18 [Accessed 5 October 2026].