A medical device AI company recently shared its validation summary with a hospital compliance team. The accuracy number was strong — 94.3% across the full test dataset. The clinical team was impressed. The legal team was not.
The legal team's question was not about the 94.3%. It was about what that number hid.
Had the company measured performance separately across age groups? Across racial and ethnic subgroups? Across geographic regions that have systematically different disease prevalence or imaging characteristics? Had a clinician outside the development team reviewed whether the model's outputs were clinically appropriate across those populations — not just statistically similar?
The answer to every question was no. Not because the company was careless. Because internal QA processes are designed to optimize aggregate performance, not to audit the demographic distribution of failure modes.
That distinction — aggregate performance vs demographic performance parity — is precisely what FDA's bias documentation requirements are designed to surface. And it is the gap that stops more hospital and health system AI procurement decisions than any other single compliance issue.
What FDA's 2025 AI Guidance Actually Requires
FDA's guidance on transparency and bias for AI-enabled devices is specific about what manufacturers must document. It is not a general requirement to "consider fairness." It requires documented evidence of performance across specific demographic dimensions relevant to the intended use population.
The dimensions FDA specifically identifies include:
- Age — performance across pediatric, adult, and geriatric populations where clinically relevant
- Sex and gender — performance parity across male and female populations, with attention to conditions that present differently across sex
- Race and ethnicity — performance across racial and ethnic subgroups with particular attention to conditions with known prevalence or presentation differences across groups
- Socioeconomic and geographic factors — performance across populations that may have different access to care, different imaging equipment quality, or different disease burden at presentation
The requirement is not to achieve identical performance across all groups — which is often statistically impossible given differences in prevalence and presentation. The requirement is to measure, document, and disclose performance across these dimensions so that clinicians, procurement teams, and regulators can make informed deployment decisions.
FDA requires manufacturers to characterize the performance of their AI model across the demographic subgroups present in the intended use population — not just report aggregate accuracy. Where performance differences exist across subgroups, those differences must be disclosed and clinically contextualized. The absence of this documentation is itself a finding.
The Aggregate Accuracy Problem — What 94% Actually Hides
The most common bias documentation gap I see in MedTech AI evaluations is not that companies have poor demographic performance. It is that they have never measured it at the subgroup level — only at the aggregate.
A model that is 94% accurate overall can perform very differently across demographic subgroups. Here is a realistic illustration of what aggregate accuracy can conceal:
The 94% aggregate number is mathematically accurate. A large well-represented subgroup with 96% performance pulls the aggregate up significantly, masking a 67% performance result for a smaller but clinically significant population. A hospital deploying this AI for that population would be deploying a tool that is wrong one in three times — while reporting 94% overall accuracy to their governance committee.
This is not a hypothetical risk. It is the documented pattern in multiple published analyses of clinical AI performance across demographic subgroups. And it is the gap FDA's bias documentation requirements exist specifically to prevent.
What Internal QA Is Not Built to Catch
Internal QA processes are designed to answer the question the development team is optimizing for: does this model perform at or above the target threshold on the validation dataset?
That is a legitimate and necessary question. It is not the same question FDA is asking.
Three structural reasons why internal QA almost always misses demographic bias:
The validation dataset is the training population. Models are typically validated on datasets that look similar to the data they were trained on. If the training data underrepresents a demographic subgroup — which is the norm, not the exception, for most clinical datasets — the validation set underrepresents that subgroup as well. Poor performance on an underrepresented group produces a small statistical signal in aggregate accuracy. The QA process passes the model. The clinical failure is invisible until deployment.
The metric is chosen by the people who chose the model. Internal QA teams select the performance metric they report. A team optimizing for sensitivity may report sensitivity. A team optimizing for AUC may report AUC. FDA's demographic bias documentation requires performance reporting on the metrics most relevant to clinical harm across demographic subgroups — which may be different from the metrics the development team chose to optimize. The same people who built the model are selecting how to measure it.
Clinical appropriateness is not a statistical test. A model can produce outputs that are statistically similar across demographic groups but clinically inappropriate for specific populations — because disease presentation, normal reference ranges, imaging characteristics, or clinical context differs across those groups in ways that aggregate statistics do not capture. Catching this requires a credentialed clinician reviewing the model's outputs across populations with clinical domain knowledge. That review is not part of any standard QA process.
Internal QA is produced by people with an interest in the outcome. The team that built the model, selected the training data, chose the performance metrics, and designed the validation protocol is the same team reviewing whether the model passed. That is not independence — it is consistency. FDA's bias documentation requirements exist precisely because self-reported bias assessment cannot carry the weight of third-party reliance.
What the CMS Non-Discrimination Requirement Adds
For AI tools touching Medicare or Medicaid patients — which includes most hospital-deployed clinical AI — FDA's bias documentation requirements are compounded by CMS non-discrimination requirements that are specific and enforceable.
CMS requires that AI tools used in prior authorization, clinical decision support, and utilization management decisions do not produce discriminatory outcomes across race, color, national origin, sex, age, or disability. The requirement is not to demonstrate identical accuracy across groups. It is to demonstrate that the AI tool does not produce systematically different — and potentially harmful — outcomes for protected populations.
The documentation standard CMS applies is similar to what FDA requires: not an assertion of non-discrimination, but documented evidence of performance measurement across the relevant demographic dimensions, reviewed by a qualified independent party.
A hospital that deploys an AI tool without this documentation and faces a CMS audit is in a difficult position. The vendor's 94% aggregate accuracy number does not answer the CMS question. Independent demographic performance documentation does.
What Adequate Bias Documentation Actually Looks Like
Based on FDA's guidance and the documentation standard that satisfies hospital legal teams and CMS auditors, adequate bias documentation has four components:
| Component | What it contains | What internal QA typically produces |
|---|---|---|
| Subgroup performance data | Performance metrics broken down by each FDA-specified demographic dimension, with sample sizes for each subgroup and statistical confidence intervals | Aggregate accuracy on the validation dataset. Subgroup breakdown absent or limited to the groups most represented in training data. |
| Clinical contextualization | Explanation of whether observed performance differences across subgroups are clinically significant, and what the clinical implications are for each flagged subgroup | Statistical comparison of subgroup means. Clinical interpretation absent or provided by the development team without independent clinical review. |
| Independent expert review | Review by a licensed clinician with domain expertise in the relevant clinical area, conducted without prior access to the development team's conclusions | Internal clinical advisor review if any. No formal blind review by an independent credentialed clinician. |
| Formal documented record | A dated, formally structured document that can be produced in a regulatory review, attached to a vendor due diligence file, or cited in a governance committee report | Internal validation summary. Not structured for external reliance. May not be dated or formally signed. |
How Independent Evaluation Closes the Gap
The demographic bias assessment in a ClearanceAI evaluation addresses all four components through a structured process that internal QA cannot replicate by design.
The automated evaluation layer runs the AI model through a structured battery of test prompts specifically designed to surface performance differences across FDA-specified demographic subgroups. The results are scored against the applicable regulatory framework — FDA 2025 AI Guidance for medical devices, CMS non-discrimination requirements for Medicare and Medicaid contexts — and flagged for expert review where performance differences exceed clinically significant thresholds.
The Layer 02 credentialed expert review then takes the flagged outputs and assesses them for clinical appropriateness. This is the component that automated testing cannot substitute for. A 13% performance gap between male and female populations may be clinically significant in one context and explainable by known physiological differences in another. That determination requires clinical judgment from a licensed domain expert — not a statistical algorithm.
The findings are compiled into the Bias and Fairness Assessment section of the formal 9-section Compliance Assessment Report — structured, dated, SHA-256 anchored, and signed by an independent third party. That is the document that answers the hospital legal team's question, satisfies the CMS auditor's request, and provides the clinical governance committee with the evidence it needs to authorize deployment with confidence.
The 94% accuracy number is not wrong. It is incomplete. The question FDA, CMS, and hospital procurement teams are increasingly asking is not whether the aggregate number is strong — it is whether the performance distribution across the populations the model will actually be used on has been independently measured, documented, and reviewed by someone outside the organization that built it.
That question has one answer. And internal QA, however rigorous, is not structured to produce it.