FDA Proposes Competency-Based Testing for Generative AI Medical Devices

FDA Proposes Competency-Based Testing for Generative AI Medical Devices

The FDA is preparing to establish new regulatory pathways for generative AI-enabled medical devices, recognizing that existing frameworks designed for traditional software and conventional artificial intelligence may not adequately address the unique risks these systems present. The agency has published a discussion paper seeking stakeholder feedback on how to ensure such devices remain safe and effective while reaching patients in a timely manner.

GenAI-enabled medical devices differ fundamentally from traditional devices and conventional AI systems in ways that complicate evaluation and monitoring. These systems learn continuously, producing variable outputs that change over time. They accept open-ended inputs, making it impractical to test every possible scenario using standard premarket methodologies. The underlying models often come from third-party developers with limited transparency into training data, architecture, and evaluation methods, making it difficult to trace specific behaviors or errors to their source.

The risks are concrete. GenAI systems may misinterpret data or fill knowledge gaps with plausible but false information, a phenomenon called confabulation. They can generate hallucinations: false facts presented as reliable output. These outputs often sound authentic to clinicians and patients, creating a dangerous gap between apparent credibility and actual accuracy. An algorithm might incorrectly associate a symptom with the wrong diagnosis, with potential consequences for patient safety.

scientist examining AI diagnostic software interface
FDA regulatory review

How The FDA Plans to Evaluate GenAI Safety

The FDA’s Center for Devices and Radiological Health has proposed a discussion paper on considerations for regulating GenAI-enabled medical devices that explores competency-based testing as a potential premarket evaluation model. This approach would parallel medical training and licensure frameworks: devices would be assessed on clinical knowledge, analytic capabilities, safety behavior, communication, and generalizability to their intended use.

The agency recognizes that competency-based evaluation alone may not fully capture how a device performs in actual clinical settings. Clinical confirmation, real-world validation with patients and clinicians, may be required to supplement laboratory assessment. A risk-based approach is anticipated, with higher scrutiny applied to functions that actively direct clinical decisions compared to systems that merely provide informational support. Agentic AI systems capable of autonomous actions would face additional oversight.

The FDA has not yet decided whether formal regulation is necessary or where guidance alone would suffice. The feedback period runs through October 19, 2026, and responses will inform whether new methodologies for premarket evaluation and postmarket monitoring are warranted. The agency aims to develop approaches that are efficient, scientifically grounded, and impose the least regulatory burden consistent with patient protection.

medical device executives and regulators in discussion
regulatory stakeholder engagement

Why GenAI Devices Cannot Use Conventional Testing

Traditional medical device regulation relies on fixed, reproducible testing protocols. Manufacturers demonstrate safety and effectiveness through premarket trials that examine device performance under defined conditions. This model assumes the device behaves consistently, the same input produces the same output, and all relevant inputs can be tested.

GenAI systems violate these assumptions. A continuously learning model may produce different outputs for identical inputs after it has incorporated new training data. The range of possible inputs is theoretically infinite, making exhaustive testing impossible. Third-party foundation models create an additional challenge: manufacturers may not have full visibility into how the underlying architecture works or what training data shaped its responses.

This opacity creates regulatory uncertainty. If a GenAI medical device produces an incorrect recommendation, determining whether the error stems from the device developer’s implementation, the foundation model itself, or unforeseen interactions between them becomes difficult. Traditional premarket pathways cannot resolve these questions before launch.

Postmarket Monitoring and Device Benchmarking

The FDA discussion paper addresses postmarket oversight alongside premarket approval. As GenAI devices operate in the field, manufacturers and regulators will need new methods to confirm the devices remain safe and effective. Device benchmarking, comparing a device’s performance against established clinical standards and competing systems, may become a routine monitoring tool.

Periodic reevaluation and public reporting of device performance could create ongoing transparency about how GenAI systems behave in clinical practice. This continuous feedback loop differs sharply from traditional device regulation, where postmarket monitoring typically focuses on adverse event reporting rather than active performance assessment.

The stakes are substantial. AI has shifted from hidden infrastructure to active clinical decision support in healthcare systems, expanding its influence on diagnosis, treatment planning, and patient communication. GenAI medical devices amplify both the potential benefits and the risks of that transition.

What Stakeholders Must Decide

The FDA is asking manufacturers, clinicians, researchers, and the public to address a fundamental question: should GenAI-enabled medical devices follow a separate regulatory track, or should existing frameworks be adapted? The answer will shape how quickly beneficial devices reach patients and how thoroughly they are scrutinized before and after market entry.

Manufacturers face pressure to innovate rapidly while demonstrating safety to a regulator developing rules in real time. Clinicians must understand how to evaluate and trust outputs from systems that may behave unpredictably. Patients depend on both accuracy and transparency, knowing not just what a device recommends, but how confident the device is in that recommendation and when human judgment remains essential.

The FDA’s willingness to solicit broad input before finalizing guidance reflects the genuine novelty of the challenge. No agency has fully solved the problem of regulating systems that improve themselves over time, accept boundless inputs, and operate as black boxes even to their developers. The feedback submitted by October 19, 2026, will determine whether the FDA’s next moves protect innovation or patients more, and whether both are possible simultaneously.

Facebook
Pinterest
LinkedIn
WhatsApp

Michael Peres (Mikey Peres) is a software engineer, journalist, tech investor and founder of Her Forward News.   Peres has developed an interest in exploring the unique mindsets of life’s outliers: extraordinary people who have weaponized their perceived limitations and found a way to succeed. His passion is to share their stories, giving strength and inspiration to those who are trying to find their way in life.

Related Articles