
The FDA has spent nearly three decades telling drug developers what proof looks like. Earlier this year, it started revising that definition. A draft guidance updating how sponsors must demonstrate substantial evidence of effectiveness for human drug and biological products – the first such revision since 1998 – does more than reflect changes in drug development; it exposes a question that has long been uncomfortable to ask plainly: when the familiar machinery of large, comparative trials cannot run on schedule, what actually counts as defensible evidence?
That question is not unique to pharmaceuticals. It recurs across specialist clinical practice and aviation safety with the same underlying structure: the obligation to act arrives before conventional evidence-generation tools can complete their work. In each domain, the most responsible response has not been to wait indefinitely, but to build explicit, rigorous, and transmissible frameworks for acting under uncertainty. Credibility, it turns out, comes not from the completeness of the evidence but from the rigour of the framework that structures action in its absence.
Structural Challenges, Not Exceptions
Evidence-based practice generally assumes that the conditions for generating robust evidence will be present before decisive action is required: enough comparable cases to study, clearly defined interventions and controls, reproducible protocols, and time to observe outcomes. In much of medicine, engineering, and public safety, that assumption holds well enough to support the standard hierarchy of trials, audits, and guidelines. The trouble starts in those corners where the caseload is sparse, the systems are intricate, and the consequences of delay are severe.
In those corners, the gap is structural. A review in Orphanet Journal of Rare Diseases describes rare-disorder evaluation as constrained by small patient populations, substantial genotypic and phenotypic diversity, and incomplete understanding of disease progression. A NASA report on safety cases for complex software notes that assurance can never amount to a complete proof, because it is impossible in practice to test every path through a system. Geotechnical standards such as Eurocode 7 explicitly anticipate sites where prediction of behaviour is difficult and direct engineers to review and adapt design during construction. Different domains, different problems – the same structural answer: act, observe, and adapt under defined conditions rather than wait for certainty that the problem’s own structure will not deliver. Institutions still tend to treat this as an edge case. It isn’t.
Under genuine structural constraint, experience either accumulates into something transmissible or evaporates case by case, depending almost entirely on what practitioners do with it. The difference between an explicit, repeatable pathway and careful improvisation is not a matter of individual skill; it is whether others can inspect and extend what was learned. Waiting for trials that the problem’s structure cannot deliver is sometimes labelled rigour – but the more honest label is avoidance. What actually separates the two is whether the framework for acting exists only in one person’s head or is built to survive beyond the encounter.

Institutions Define Frameworks
One influential response is emerging from the bodies that define professional standards themselves. The FDA’s revised draft guidance on demonstrating substantial evidence of effectiveness directly revisits the core definition of proof. The agency explains that sponsors may, in some circumstances, rely on one adequate and well-controlled clinical investigation together with other confirmatory evidence to satisfy the statutory standard, reflecting changes in drug development and advances in data quality. Because the document is open for public comment and intended to replace 1998 guidance, it makes explicit what was already implicit: the benchmark for proof is a live, revisable construct rather than a fixed rule.
A companion draft guidance on individualised therapies for ultra-rare diseases goes further by naming an alternative evidentiary architecture. Where randomised trials are not feasible, the FDA’s Plausible Mechanism Framework organises approval around four components: define the disease-causing abnormality and mechanism, situate expectations in well-characterised natural history, and confirm that the target has been successfully engaged. For traditional approval, the agency still expects evidence of improved clinical outcomes, disease course, or validated biomarkers – but it acknowledges that how that evidence is assembled must differ, and it invites public scrutiny of that structure.
Aviation regulators face a version of this problem that is more structurally acute: for technologies with no compliance history at all, the evidence-building phase cannot sit alongside certification – it must precede it, because there is no existing framework from which to deviate. The Federal Aviation Administration responded by launching an Advanced Air Mobility and electric vertical take-off and landing (eVTOL) Integration Pilot Program, selecting eight proposals from more than 30 submissions spanning passenger transport, cargo and logistics, emergency medical response, autonomous flight, and offshore and energy-sector operations across 26 states, with operations expected to begin in 2026.
What makes the FAA’s framing significant is its deliberateness about what the pilot phase is and is not – regulators who name the evidence-building process explicitly are doing something important: they’re ensuring it can’t later be mistaken for a workaround. Donnell Evans, an FAA spokesperson quoted in a WIRED report on the integration pilot projects, drew the line clearly: “The pilot program is focused on informing standards and future policy development and is not a mechanism to bypass certification requirements.” The operative principle is that legitimacy under uncertainty comes from bounded, reviewable trials that build the evidentiary base for certification – in addition to, not instead of, its requirements.
Bedside Logic
At the bedside, the same structural gap appears in conditions that are both uncommon and technically demanding to treat. Atlantoaxial osteoarthritis – affecting the C1–C2 joint at the top of the cervical spine – carries serious functional consequences yet occurs too infrequently for conventional multicentre randomised trials to be a realistic prospect. Posterior C1–C2 fixation is anatomically constrained and technically complex, making consistent standardisation across sites difficult, while prolonged watchful waiting is not a neutral option for patients with severe pain or instability. Here, as in rare-disease pharmacology and eVTOL regulation, the obligation to act arrives well before traditional evidence machinery can deliver its strongest products.
The risk of acting without structured tracking is not theoretical. A New England Journal of Medicine Perspective on metal-on-metal hip implants describes what can happen when adoption outpaces documentation: more than 500,000 patients in the United States received these prostheses, most between 2003 and 2010, and subsequent data showed they failed at higher rates than alternative designs. The authors argue that weaknesses and delays in surveillance and evaluation meant the field was slow to recognise the scale of the problem. The lesson for high-complexity interventions is pointed: when randomised trials cannot realistically precede widespread use, disciplined, peer-legible outcome tracking becomes the main safeguard against harms that surface only in retrospect. That lesson shapes what a structured pathway is actually for.
Dr Timothy Steel, a neurosurgeon and minimally invasive spine surgeon at St Vincent’s Private and Public Hospitals, has developed a complex cervical reconstruction pathway for atlantoaxial osteoarthritis – a structured, documented clinical approach refined over a defined patient cohort. The approach runs from preoperative CT and MRI planning through C1–C2 posterior fixation using transarticular screws and Harms constructs, intraoperative navigation, and defined postoperative imaging to confirm fusion – each step named explicitly so the approach can be evaluated or replicated by another team rather than remaining in a single practitioner’s recall. An external study of 23 patients treated between 2005 and 2015 reported Visual Analogue Scale neck pain scores falling from 9.4 to 2.9, Neck Disability Index scores from 72.2 to 18.9, radiographic fusion in 95.5 per cent of patients, and 91 per cent willingness to undergo the procedure again. The value of those numbers is not their magnitude – it’s that they exist at all as a shared, inspectable record rather than a practitioner’s private recall. That distinction is precisely what turns local expertise into something other clinicians can evaluate, adapt, or reject.
The structure of this pathway mirrors, at a smaller scale, the logic in the FDA’s Plausible Mechanism Framework: a clearly defined condition and indication, a standardised intervention mechanism, a followed cohort, and documented radiographic and patient-reported outcomes.
Credible Frameworks Versus Preferences
In innovation-heavy, evidence-constrained settings, the central question is how to tell a credible expert-derived pathway from a set of rationalised preferences. Four practical features provide a working answer: a defined cohort with documented outcomes, explicit reasoning linking indications and techniques to expected results, some form of peer scrutiny even without a randomised trial, and design for transmission rather than proprietary use. The BMJ’s IDEAL framework for evaluating surgical innovation – which describes the evidence base in surgery as vastly weaker than in drug development and makes transparency and staged evaluation a key principle – gives these features a formal methodological footing as safeguards when trial-grade proof is hard to obtain early on. They are not substitutes for that proof; they are what makes expert-derived practice legible enough to be questioned. The line between a tested framework and a confident practitioner’s accumulated anecdote is thinner than it looks – and these four features exist precisely to make that line visible.
When those features are absent, the problem is not merely weaker evidence – it’s that knowledge stays locked in individual memory, and the difference between a genuine pathway and a trail of anecdotes becomes invisible from the outside. A BMJ Open systematic review of studies citing the IDEAL framework warns that potentially harmful procedures can become widespread before their risks are recognised and finds that many modifications to techniques are simply not documented, reported, or shared. Structured approaches such as the FDA’s individualised-therapy guidance and the FAA’s eVTOL integration pilot program aim to counter that failure mode by tying innovation to clear rationales, defined cohorts, and reviewable data. Without explicit documentation and peer exposure, expert enthusiasm and rigorous evidence look identical until it’s too late to tell them apart.
The risk of flexibility, however, is real. A STAT First Opinion commentary has argued that the FDA’s move toward more flexible oversight of cell and gene therapies may accelerate beneficial development but also creates new risks for patients, the field, and potentially the agency itself. By building room for judgement into both the evidence standard and the approval mechanism, the FDA is necessarily creating space for interpretation about what counts as adequate confirmation and acceptable mechanistic plausibility – and the commentary’s warning underlines how that room can be misused if safeguards are weak. The same is true for any expert-derived clinical pathway: the more weight it carries in practice, the more tempting it becomes to present personal conviction as if it were a tested framework. Without the four distinguishing features, that’s a temptation with no structural check at all.
Honest Instruments for Permanent Conditions
The FDA’s decision to revisit its definition of substantial evidence of effectiveness can be read less as a loosening of standards than as an attempt to make them honest about what different problems can realistically yield. In the same landscape sit the eVTOL integration pilots commissioned by the FAA, the framework for individualised therapies in ultra-rare diseases, and clinician-built pathways such as the structured approach to atlantoaxial cervical reconstruction developed by Dr Timothy Steel – each turning unavoidable uncertainty into something that can be argued about in public rather than left as private comfort.
For rare, high-complexity problems with small numbers, heterogeneous presentations, and high consequences for delay, the evidence gap is not a temporary inconvenience that patience will eventually close. It is a recurring feature of how these systems work. Professions that handle this condition best equip themselves with honest instruments – explicit frameworks, rigorous documentation, and designs that invite scrutiny and transmission. The alternative is not safety; it’s opacity. And opacity, as the metal-on-metal hip implant episode demonstrated, has its own clinical consequences.
Gearfuse Technology, Science, Culture & More
