Turn a completed BVCT into a structured 11-module protocol package: a Statistical Analysis Plan decision engine plus ten downstream protocol modules, each scored for readiness — an AI draft for qualified human review, never a submission-ready document.
BVCT Packages takes one completed BioinvestGPT Virtual Clinical Trial (BVCT) result and expands it into an eleven-module protocol package — the scaffolding a sponsor would assemble around a single trial design. Module 1 is a constrained Statistical Analysis Plan (SAP) decision engine; modules 2–11 are ten downstream protocol modules (power, safety, consent, data management, CMC, regulatory, and more). Each module is scored, gaps are logged, and the open judgment calls are routed to the human who must own them.
The panel lives at the BVCT Packages route and is reached from the validator sidebar. Its header carries a permanent, non-dismissible banner that states the operating contract of the whole feature:
AI-generated draft — not for regulatory submission · requires qualified human sign-off · Green ≠ approved. Every claim in this chapter is bounded by that banner. A "green" readiness light means no internal inconsistency was detected, not that the package is approved, correct, or fileable.
| Principle | What it means in the package |
|---|---|
| Never invent | No numeric safety, efficacy, tox, CMC, stability, or prior-clinical value is ever stated unless it is present in the inputs. When such evidence is needed but absent, the module emits a "MISSING — source required" gap instead of fabricating a number. |
| Modules never self-score | Readiness is computed deterministically, outside the AI. A module cannot grade its own homework. |
| Disqualifier-dominated | A single critical violation (a red flag, a missing source, a missing sign-off) forces Red regardless of how well everything else scores. You cannot launder a fatal gap with strong scores elsewhere. |
| Failure-isolated | Modules generate independently; if one fails it becomes an error stub with a Red light, and the other ten still complete. |
Where this fits: Packages is a post-result step. You first run a BVCT to a GO / NO-GO decision (see Understanding Your Results); the package then takes that completed result as its starting point. See Open BVCT Packages below.
Scope & track record: BVCT outputs and the package built on top of them are model-based decision-support analyses, not investment advice and not regulatory submissions. Prospective track record: data.bioinvestgpt.com.
You open the package generator from the validator, pick a completed BVCT as the source, and let the panel pre-fill the inputs from that result. The flow has three states before generation: select a result, review inferred inputs, then generate.
From inside the Validator, open the BVCT Packages panel (the violet header with the stacked-files icon, titled "11-Module Protocol-Package Generator"). You can deep-link to it in a new tab at bvct.bioinvestgpt.com/#/validator/packages.
If you have not finished any BVCT yet, the panel shows an empty state: "No completed BVCT results yet — complete a BVCT workflow to generate a protocol package here." Run a BVCT to completion first.
Under "Source result", the panel lists your completed BVCTs as chips, each labelled with its BVCT ID. Click one to select it.
The panel then does one of three things automatically:
The package reads the completed BVCT result and infers the trial design parameters it needs, wrapping each one with its provenance so you can see where every value came from. Each input field carries a small tag:
| Tag | Meaning |
|---|---|
| auto-filled | Inferred from the BVCT result (the result PDF / rationale / metadata). |
| edited | You typed over the inferred value. |
| default | A safe default was used because nothing could be inferred. |
The inferred / editable inputs include the trial drug, indication, phase, comparator arm, primary endpoint and its type, the design type (superiority / non-inferiority / equivalence), the effect-measure type (HR / RR / OR / rate ratio / mean difference), the endpoint polarity and numerator orientation (the direction trio), the predicted effect magnitude and precision interval, the predicted estimand type, the NI margin (if non-inferiority), the target population, allocation ratio, target power, one-sided alpha, expected control event rate, and the BVCT's GO / NO-GO verdict.
Tip: review the auto-filled fields before you generate — especially design type, effect measure, and the direction trio (endpoint polarity + numerator orientation). The SAP decision engine branches hard on these, and the favorable direction is derived from them, never assumed. A wrong polarity can flip the entire interpretation.
There is also an Evidence tab on the selected result. Evidence-bearing modules (Safety, Risk-Benefit, Consent, CMC) will attempt an evidence-grounded, cite-or-gap drafting pass against any source documents you have made available; with no embedded evidence they fall back to the never-invent module and emit the corresponding gaps.
Module 1 is not a writer — it is a constrained decision engine for the Statistical Analysis Plan, operating under ICH E9 and ICH E9(R1). It evaluates eight statistical decision nodes, each returning a structured verdict, and then a separate deterministic finalization gate turns those verdicts into eight 0–100 scores and a single Red / Yellow / Green classification.
| Decision node | Gate score | What it checks |
|---|---|---|
| Estimand | Estimand alignment | Is the E9(R1) estimand fully specified (all five attributes), and are intercurrent events given a strategy rather than left UNRESOLVED? |
| Effect-size alignment | BVCT effect-size compatibility | Does the planning effect size align with the estimand, with any null-ward adjustment justified? |
| HR suitability | HR suitability | Is a hazard ratio an appropriate summary measure (proportional-hazards risk, prespecified diagnostics, NI-margin consistency)? |
| Treatment switching | Treatment-switching robustness | Is crossover-induced dilution handled, and is the chosen adjustment method (RPSFT / IPCW / two-stage) feasible given the data actually collected? |
| Post-discontinuation | Post-discontinuation data adequacy | Is follow-up after treatment discontinuation collected adequately for the estimand? |
| Missing data | Missing-data robustness | Is there a credible missing-data / sensitivity-analysis strategy? |
| Regulatory defensibility | Regulatory defensibility | Would the plan withstand regulatory scrutiny — capped so it can never exceed its weakest load-bearing pillar (estimand, effect size, post-discontinuation)? |
| Sponsor claim | Sponsor-claim alignment | Does the analysis actually support the claim the sponsor intends to make? |
The gate is disqualifier-dominated — the classification is decided by violations first, scores second:
| Classification | How it is reached |
|---|---|
| Red | Any red flag is raised, or any of the eight scores falls below 70. |
| Yellow | No red flags and every score ≥ 70, but at least one score is below 85 — open judgment items remain. |
| Green | No red flags and every score ≥ 85. Means "no internal inconsistency detected — ready for human statistical/regulatory review," not approved. |
Red flags come from the engine verdicts (e.g. an incomplete estimand, an unresolved estimand↔effect-size mismatch, a non-inferiority design with null-ward attenuation applied, a summary measure swapped while the NI margin is HR-defined, or a switching method recommended whose required data is not collected) and from independent validators. The gate also reports a human-decisions-required count and a one-line rationale.
The eight scores are uncalibrated internal-consistency indicators, not measurements. The panel deliberately leads with the R/Y/G light and the verdict checklist, not the numbers. A score's "min revision to green" tells you how many points short of 85 it is.
The sample size is computed deterministically (never narrated into existence by the AI) and stamped with a re-derivation provenance hash for reproducibility. If it had to be suppressed, the SAP tab says so and flags it as a gap rather than guessing. Modules that need it (the SAP itself and the Power module) cannot be "ready" without a non-suppressed, provenance-stamped N. The SAP narrative draft is then written around these locked facts as a thirteen-section document for human review.
Overriding a Red gate: a Red SAP gate can be overridden only by an authenticated user who records a written justification. The override does not mutate the gate — the original Red classification is retained and shown in the exported PDF and the audit trail. An override is a documented human decision, not an erasure.
Module 1 is the SAP (above). Modules 2–11 are ten downstream protocol modules. Each is generated by a single constrained AI call returning structured sections plus a list of gaps and sponsor flags; the safety-bearing modules (Safety, Risk-Benefit, Consent, CMC) are additionally lint-checked, and any fabricated safety/CMC number is down-classified to a gap. The ten modules run in two waves of five, failure-isolated, while you watch a per-module progress checklist.
| # | Module | What it drafts | Sign-off owner |
|---|---|---|---|
| 1 | Statistical Analysis Plan (SAP) | The decision engine + finalization gate + deterministic sample size + 13-section narrative draft. | Statistician |
| 2 | Power & Sample-Size | Power / sample-size rationale narrated around the deterministically-computed N, with an assumptions table and a sensitivity discussion. Never invents or alters N. | Statistician |
| 3 | Safety Monitoring Plan | AE/SAE/AESI definitions, causality & severity framework, reporting cadence, DSMB/DMC charter outline. AE rates and numeric stopping boundaries are emitted as gaps, never invented. | Medical |
| 4 | Risk-Benefit Assessment | The benefit-risk framework (disease burden, hazards, mitigation, alternatives). Cannot assert a "favorable" balance without real safety data — classifies as "uncertain — requires sponsor safety data." | Medical |
| 5 | Informed Consent (template) | A consent template with placeholders. No specific risks, frequencies, or benefit claims; no predicted effect size enters the document; every risk/benefit line is marked "requires source + IRB/IEC review." | Regulatory |
| 6 | Schedule of Assessments | The screening → follow-up assessment grid; verifies every endpoint has a collecting assessment. | ClinOps |
| 7 | Data Management Plan | Data flow, eCRF structure, edit checks, query management, database-lock criteria; captures intercurrent-event and switching predictors. | ClinOps |
| 8 | Monitoring & Site Operations | Risk-based monitoring / site-operations plan structure (ICH E6(R3)). RBM thresholds and SDV % are emitted as gaps tied to the risk assessment, never invented. | ClinOps |
| 9 | Investigational Product Handling | Dispensing, accountability, blinding/labeling, returns/destruction. Storage / temperature / stability specifics are emitted as gaps without a source. | ClinOps |
| 10 | CMC / Nonclinical / Prior-Clinical | A required-evidence CHECKLIST only (drug substance/product, manufacturing, stability, impurities, tox, starting-dose rationale, prior human safety). The system has none of this evidence, so every item is "MISSING — source required" and zero specific numbers are generated. | Regulatory |
| 11 | Regulatory & Ethics Documentation | The submission / IRB-IEC checklists, trial-registry field map, essential documents, and country considerations. Asserts no approval and predicts no outcome. | Regulatory |
Reading a module tab: each module tab shows a Red/Yellow/Green pill with its score, the drafted sections, and a yellow "Gaps / missing evidence" box listing every open item. The module score starts at 100 and decays 10 points per gap and 6 per sponsor flag; a module with any open gap can never be Green.
CMC is intentionally near-empty. Module 10 is a checklist of missing evidence by design — the platform holds none of the manufacturing, stability, or tox data, so it lists what a human must supply rather than inventing it. A red CMC light is the expected, honest output, not a bug.
After all eleven modules settle, a deterministic (no-AI) synthesis pass folds their outputs into three cross-cutting views:
| View | What it contains |
|---|---|
| Issue log | One row per problem, with a severity (critical / high / medium / low) and a recommendation. SAP gate red flags become critical issues; a module's gaps become high if that module is Red, otherwise medium. |
| Sponsor decision table | One row per sponsor flag, each marked blocking when its module is Red — the explicit decisions a sponsor must make before the package can move. |
| Missing-info list | Every "MISSING — source required" gap, with which module needs it and why — the shopping list of source evidence to gather. |
There are two distinct readiness layers, and it helps to keep them separate. The SAP finalization gate (covered above) scores the statistical plan's internal consistency. The submission readiness rubric on the Readiness tab scores how close each module is to being structurally fileable into an eCTD dossier for a chosen regulator.
Submission readiness is deterministic and AI-free. Each module's readiness is the product of four independent factors, each in 0–1:
| Factor | Goes to 0 when… |
|---|---|
| Evidence | The module's generation failed, fabrication was blocked, or (for source-bearing modules) the required source document is missing. Decays gently with the number of open gaps otherwise. |
| Compute | The module needs a validated sample size (SAP and Power do) but the N is suppressed or not provenance-stamped. Modules with no compute requirement score 1. |
| Sign-off | The module's required human role (statistician / medical / regulatory / clinops) has not signed it off. |
| eCTD conformance | The module is not yet mapped to its eCTD leaf in the dossier. |
Multiplicative by design. Because the four factors are multiplied, any factor of 0 zeroes the whole module. You cannot offset a missing source, a missing sign-off, or an un-mapped leaf by scoring well on the other three. The breakdown — which factor is zero — matters far more than the headline percentage.
The package percentage is the mean of the per-module readiness × 100. The Red/Yellow/Green classification is again disqualifier-dominated:
| Classification | Condition |
|---|---|
| Red | Any single module is zeroed, or the package percentage is below 50%. |
| Yellow | Between the two — open work remains, but nothing is fully disqualifying. |
| Green | Package ≥ 80% and every module is at least 0.5 ready. |
The percentage is not a probability of acceptance. The rubric ships with calibrated: false — the number is an internal-consistency / structural-completeness index, not a validated probability that a regulator will accept the dossier. It will stay uncalibrated until an external benchmark study maps these factor products to observed sponsor-review pass rates.
The Readiness tab leads with the breakdown table (evidence / compute / sign-off / eCTD per module) and a "Human-owned remainder" list — the irreducible work a qualified human must still own, phrased as the disqualifying factor first, e.g. "M3 Safety Monitoring Plan: requires medical sign-off" or "M10 CMC: evidence missing, blocked, or SAP red-flagged."
Two tabs feed the Readiness rubric live:
m5.3.5.3, Risk-Benefit → m2.5, CMC → m3.2.s) with a "mapped" checkbox, a traceability matrix, and CT.gov / EU-CTIS registry fields (missing fields render as TO-DO, never invented). The target regulator (FDA / EMA / PMDA; NMPA and global fall back to FDA-shaped requirements) is taken from the inputs.Tip: the honest path to a higher percentage is to supply real source evidence, get the right human to sign each module, and map every module to its eCTD leaf — in that order. There is no shortcut that bypasses the zeroing factors.
Once a package is generated, an Export PDF button appears in the panel header. It compiles the whole package — the SAP draft and gate, all ten downstream modules with their sections and gaps, the issue log, the sponsor decision table, the missing-info list, the readiness breakdown, and any recorded gate overrides — into a single document for circulation and human review.
With a package showing, click Export PDF in the top-right of the header. The button shows a spinner while the document is assembled, then downloads.
If a Red SAP gate was overridden, the export preserves the original Red classification alongside the override record (who overrode it and the justification) — the document never hides that a human chose to proceed past a red flag.
| The export IS | The export is NOT |
|---|---|
| A structured AI draft of an 11-module protocol package, with every gap and open decision made explicit. | A regulatory submission, an approved protocol, or a fileable eCTD dossier. |
| A review artifact you route to statisticians, medical, regulatory, and clinical-operations owners. | A substitute for those qualified humans' sign-off — the banner and the readiness rubric both enforce that. |
| An audit-honest record that retains original Red classifications and override justifications. | A document that launders a missing source, a suppressed sample size, or a missing sign-off into a "green." |
Why is my package Red even though most scores look high?
Both readiness models are disqualifier-dominated. A single SAP red flag, a single score below 70, a missing source on an evidence module, a missing sign-off, or an un-mapped eCTD leaf is enough to force Red on its own — by design, so a fatal gap can never be averaged away.
Can I trust the readiness percentage as a chance of approval?
No. It ships uncalibrated (calibrated: false) and is an internal structural-completeness index, not a probability of regulatory acceptance. Read the per-module factor breakdown and the human-owned remainder, not the headline number.
A module came back nearly empty (especially CMC). Did it fail?
Usually not. The never-invent rule means modules that need source evidence the platform does not hold will list that evidence as "MISSING — source required" rather than fabricate it. Module 10 (CMC) is a missing-evidence checklist by design.
What if I disagree with a Red SAP gate?
An authenticated user can record an override with a written justification. The override is logged and the original Red is retained in the export — it documents a human decision, it does not erase the finding.
Continue to Understanding Your Results for the BVCT result that feeds this package, or to the Glossary for terms like estimand, hazard ratio, eCTD, and ICH E9(R1). BVCT outputs and packages are model-based decision-support analyses, not investment advice. Prospective track record: data.bioinvestgpt.com.
BioinvestGPT BVCT Platform User Guide — BVCT Packages — 11-Module Protocol. BVCT outputs are model-based decision-support analyses, not investment advice. Prospective track record: data.bioinvestgpt.com.