Healthcare and biotech leaders face an increasingly expensive temptation: when an AI project underperforms, assume the answer is a larger or more sophisticated model. Yet the central argument highlighted by TechCrunch’s report on a cancer-AI startup is more fundamental: AI is not close to curing cancer, and progress depends heavily on the data available to researchers and systems—not on model ambition alone.
That distinction matters operationally. A model cannot recover information that was never captured, reconcile clinical definitions that vary invisibly between hospitals, or create valid consent for a new use of patient records. Nor can computational scale make a weak proxy into a meaningful medical endpoint. Before commissioning another model, healthcare organizations need a data-readiness plan that establishes whether the proposed clinical question can be answered responsibly.
Start with the clinical decision, not the model specification
A healthcare AI initiative should begin with a precise statement of intended use. “Improve oncology” is not a usable objective. Leaders need to identify the population, setting, user, decision and outcome involved. Is the system intended to prioritize cases for review, identify candidates for a study, predict treatment response, or support diagnostic interpretation? Each purpose requires different records, labels, validation methods and safeguards.
The intended use also determines the cost and timeline. A retrospective research tool assembled from an existing dataset has different requirements from a system expected to influence care across multiple clinical sites. If the team cannot state who will act on an output and what action could follow, it is too early to select an architecture or promise a deployment date.
Define success in clinical and operational terms before considering conventional model metrics. Sensitivity or accuracy may be relevant, but they do not independently show that a tool improves decisions, fits workflow, performs across patient groups or avoids harmful delays. The endpoint must match the claim. An association found in historical data cannot automatically support a claim about treatment benefit, survival or a cure.
Inventory what can actually be used
A data inventory should document more than the existence of electronic health records, imaging archives, genomic files or trial datasets. For each source, record the period covered, patient population, provenance, format, missingness, update frequency, custodial owner and permitted uses. Include the practical route to access: ethics review, contractual approval, consent restrictions, de-identification requirements and technical extraction work.
This exercise often exposes the difference between nominal and usable volume. Millions of records may contain few patients with the required diagnosis, intervention, follow-up period and outcome. Relevant information may be embedded in free text, stored under incompatible coding systems, split across institutions or absent because patients received later care elsewhere.
Teams should then test fitness for the defined question:
- Completeness: Are the predictors, interventions and outcomes available at the required points in time?
- Consistency: Do sites use the same definitions, units, codes and collection procedures?
- Representativeness: Does the dataset reflect the populations and care environments where the system would operate?
- Label integrity: Were outcomes measured reliably, or inferred through proxies that could introduce error?
- Temporal validity: Have clinical practice, equipment or coding conventions changed during the collection period?
A small, relevant and well-characterized cohort can be more informative than a larger collection assembled without a clear lineage. That is an operator judgment, not a claim about any particular startup’s method.
Resolve rights, privacy and collaboration before engineering
Access is part of readiness, not an administrative task to postpone until a prototype works. Healthcare organizations must establish whether consent, law, institutional policy and contracts permit the proposed use, linkage and transfer of data. De-identification does not eliminate every privacy risk, particularly when datasets contain rare conditions, detailed timelines or genomic information.
Multi-institution projects also need an explicit collaboration design. Leaders should decide whether information will be centralized, analyzed within each institution, or made available through another controlled arrangement. The appropriate pattern depends on permitted use, security capabilities, reproducibility needs and the sensitivity of the records. Whatever the design, access should be limited, logged and reviewable.
Agreements should address who can use the data, for which purpose, for how long, and whether derived datasets or models can be retained. They should also cover incident response, publication, intellectual property, withdrawal where applicable and responsibility for downstream use. If these questions remain unresolved, apparent engineering speed may simply move legal and ethical risk closer to deployment.
Standardization is a product dependency
Clinical data rarely arrive as a model-ready table. Standardization may require mapping terminology, reconciling units, normalizing dates, resolving duplicate identities and documenting how missing values are handled. Imaging and molecular data add their own acquisition and processing differences. Free-text extraction introduces another layer of uncertainty that must be measured rather than hidden.
Preserve provenance throughout this work. Every feature and label should be traceable to its source and transformation. Version datasets, code, inclusion criteria and annotation guidance so that an analysis can be reproduced. A model result without a recoverable data lineage is difficult to audit and harder to improve safely.
Budget accordingly. Data engineering, specialist annotation, privacy review, security controls and clinical validation may dominate both expenditure and schedule. Leaders should therefore create milestones around data feasibility: an access-approved sample, a completed quality profile, a harmonized pilot dataset and evidence that the intended endpoint can be measured. Only then is a larger training investment defensible.
Validate for clinical meaning, bias and failure modes
Internal test performance is only one layer of evidence. Validation should use data separated appropriately from development and, where the intended claim requires it, evaluate performance across sites, time periods, equipment, workflows and relevant patient groups. The team must look for hidden shortcuts—for example, a model learning institutional practices or documentation patterns rather than the biological or clinical signal it is supposed to detect.
Subgroup analysis is essential because aggregate performance can conceal unequal error rates. Representation alone does not guarantee fairness, but major gaps should trigger additional collection, a narrower intended population or explicit limits on use. Uncertainty and abstention behavior also matter: a system should not present confidence where its input falls outside the conditions represented in development.
Clinical experts should review false positives and false negatives in context. Ask what harm each error could cause, whether users can detect it and how the output changes an actual decision. Prospective evaluation may be necessary before making claims about workflow or patient outcomes. Communications must match the evidence: promising retrospective results do not demonstrate improved survival, successful treatment or a cure for cancer.
Make readiness an ongoing governance function
Readiness does not end when a model ships. Assign named owners for dataset stewardship, clinical safety, privacy, security, model performance and approval of new uses. Establish change controls for new sites, data fields, clinical guidelines and model versions. Monitor input drift, missingness, subgroup performance and the relationship between outputs and meaningful endpoints.
Before increasing model size, leadership should require evidence that eight conditions are met: the clinical question is specific; the necessary data exist; quality and representativeness are characterized; consent and access rights are valid; records are standardized with documented provenance; collaboration is secure; validation matches the proposed claim; and accountable owners can monitor the system over time.
If one of these conditions fails, the right next step may be narrower scope, better collection, longer follow-up or a non-AI intervention—not more compute. The broader lesson from the source’s data argument is not that model research lacks value. It is that in healthcare, credible model ambition starts with accessible, clinically meaningful and governable evidence.
