BioinvestGPT ApS · Research use only
How to evaluate BVCT predictions
A BVCT prediction is evaluated by fixing the question before the readout (drug, population, comparator, endpoint and predicted effect size), timestamping it, and scoring it against the reported outcome under a published rule. Accuracy, PPV and NPV summarize the binary GO/NO-GO calls; the error of a predicted effect size needs its own measure.
Last updated · Maintained by BioinvestGPT ApS
What is publicly described about the approach?
BVCT simulates the mechanistic interaction between a drug and a disease-matched virtual human body and outputs a quantified effect size against a named comparator, with the mechanistic rationale. It is patient-data-free: it works from the drug's design and mechanism of action, the disease context and the protocol, not from patient records.
This page explains how to assess the reported predictions. It does not disclose the proprietary model specification or assert that every evaluation item below has already been independently verified.
What should a prospective evaluation specify?
Before the relevant readout, fix the information cutoff, the eligible evaluation cohort, the intervention and comparator, the endpoint, the time horizon, the prediction and the scoring rule. Keep a versioned record of amendments. Distinguish the prediction time, the timestamp-verification time, the first public disclosure and the adjudication time.
How should binary predictions be reported?
| Measure | Definition | Interpretation |
|---|---|---|
| Accuracy | (TP + TN) / (TP + TN + FP + FN) | Correct calls among adjudicated calls |
| PPV | TP / (TP + FP) | Positive calls meeting the positive-outcome rule |
| NPV | TN / (TN + FN) | Negative calls meeting the negative-outcome rule |
| Sensitivity | TP / (TP + FN) | Positive outcomes captured |
| Specificity | TN / (TN + FP) | Negative outcomes captured |
TP means true positive; TN, true negative; FP, false positive; and FN, false negative under the published adjudication rule. Report a zero-denominator measure as unavailable. State exclusions and pending outcomes outside these denominators.
Use confidence intervals appropriate to the data structure. Multiple endpoints or predictions from one drug may be correlated; counting them as independent can overstate precision. Report the number of unique predictions and trials as well as the number of scored readouts: BVCT's published figures count 705 scored readouts from 692 predictions (see the evidence page).
How should quantitative predictions be evaluated?
Define the error measure in advance, such as the mean absolute error of a specified ratio or the error on a log-ratio scale. Show the distribution of errors, calibration where applicable and interval coverage. Reporting a ratio at 0.1 precision (one decimal place) is different from demonstrating an error within 0.1.
Keep HR, OR and RR distinct, and compare like endpoints and populations. An odds ratio is not a Relative Ratio (a ratio of proportions), and a hazard ratio describes a time-to-event comparison.
How should generalization be assessed?
Evaluate performance by therapeutic area, phase, modality, comparator and prediction horizon, with adequate denominators. Show misses and outcomes that could not be adjudicated. Use external or blinded evaluation to assess transportability. Historical agreement is evidence to examine, not a guarantee for a new program.
What should an independent review disclose?
The reviewer, date, cohort, accessible source materials, adjudication procedure, conflicts of interest and exact scope of work. Timestamp review, arithmetic checking and scientific validation answer different questions.