
What You Should Know
- Cognita Imaging, Inc. (a wholly owned subsidiary of Mosaic Clinical Technologies, Inc., operating within the broader Radiology Partners ecosystem) was awarded a $1.29 million research contract by the U.S. Food and Drug Administration (FDA) to pioneer new methodologies for evaluating autonomous, generative AI radiology reports.
- The 18-month contract, effective June 22, 2026, was funded under the FDA’s Broad Agency Announcement (BAA) program for advanced research and development in regulatory science.
- The initiative, titled “Virtual Subject Matter Expert Agents for Radiology Report Evaluation,” addresses the clinical limits of traditional human reader studies, which evaluate only a few hundred cases and fail to capture rare edge cases, subtle equipment discrepancies, or regional workflow variations.
- Cognita will build and validate an “LLMs-as-a-jury” algorithmic framework that runs multiple large language models concurrently to grade and cross-examine both human-written and AI-generated imaging reports, mitigating single-model bias.
- The evaluation engine will be applied across a massive dataset of approximately 1 million patient exams from a diverse U.S. patient cohort, analyzing performance variance across demographics, care settings, imaging hardware manufacturers, and low-prevalence pathologies.
The “LLMs-as-a-Jury” Evaluation Architecture
The project is led by Principal Investigator Dr. Akshay Chaudhari (Cognita co-founder and Associate Professor of Radiology and Biomedical Data Science at Stanford University) and co-investigator Dr. Louis Blankemeier (Cognita co-founder and CEO):
- Multi-Model Consensus Scoring: Rather than relying on a single large language model (which inherits that model’s idiosyncratic biases and failure modes) to score diagnostic drafts, the architecture aggregates evaluation outputs across an ensemble of distinct LLMs acting as a virtual panel of expert reviewers.
- Building on the Open-Source GREEN Benchmark: The system expands upon Cognita’s prior research developing GREEN (Generative Radiologist Evaluation and Error Notation), an open-source evaluation framework designed to capture clinically meaningful discrepancies between ground-truth radiologist reports and AI-generated text.
- Targeted Radiologist Escalation: The platform isolates high-discrepancy outliers and clinically significant disagreements for expert human review. Radiologists assess whether errors originated from the generative drafting model, discordance within the LLM jury, or ambiguity in the initial human radiologist’s ground truth.
Validation Across 1 Million Real-World Patient Exams
Following core architectural development, the framework will be stress-tested across an enterprise dataset of approximately 1 million patient imaging exams drawn from a diverse U.S. clinical cohort:
- Edge-Case & Subgroup Discovery: Analyzes model behavior across variable scanner manufacturers, geographic site types, patient demographics, and low-incidence pathologies that are routinely omitted from standard hundreds-of-cases clearance studies.
- Sample-Size Sensitivity Analysis: Researchers will extract smaller validation sub-cohorts from the 1-million-exam dataset to empirically quantify how many edge-case diagnostic failures are missed when regulatory or institutional evaluations rely on small datasets.
- Enterprise Radiology Integration: The initiative leverages clinical workflows and data pipelines from MosaicOS and the broader Radiology Partners network, which covers over 4,000 radiologists interpreting 55+ million imaging studies annually.
Regulatory Science Deliverables
Cognita will provide the FDA with open software code, benchmarking guidelines for constructing multi-agent LLM juries, and comparative performance analyses examining premarket clearance and post-market safety monitoring. The contract serves to advance regulatory science standards and does not constitute formal FDA product approval or commercial clearance.

