Data plane
Checks cohort integrity, missingness, documentation inequity, leakage, and spurious signals.
Distributed Open Justice Oversight
DOJO is an open, community-driven platform for continuous adversarial evaluation of health AI—from the data that shapes a model to the workflow in which people rely on it.
One system. Continuous evidence.
01 / What is DOJO
MIT Critical DataDistributed Open Justice Oversight is an open, community-driven platform for adversarial evaluation and remediation of health AI systems.
It helps researchers, clinicians, and evaluation teams examine the full system—not only a model score—before and during clinical use.
Shared tools, challenge sets, and evidence that can be reviewed and improved.
Designed to work near the data and within institutional security requirements.
Clinical, technical, and lived expertise shape what gets tested.
02 / Architecture
Three connected planes make failures visible at the point where they enter the system.
Checks cohort integrity, missingness, documentation inequity, leakage, and spurious signals.
Challenges robustness, calibration, rare cases, shifted inputs, unsafe confidence, and abstention.
Tests interoperability, infrastructure, human-AI interaction, and readiness in context.
Every consequential claim stays linked to the model, data, evaluator, environment, and re-test that produced it.
Selected work from the MIT Critical Data research program and DOJO team, spanning data quality, model evaluation, bias, safety, and clinical deployment.
The primary DOJO paper describes the community-driven evaluation layer, its three planes, and its system-level approach to health AI.
Read the DOJO paperExamines how shortcuts spread through clinical multi-agent systems and why independent oversight is needed to detect benchmark gaming and socially plausible errors.
Open paperIntroduces paired image-swap audits to test whether report-conditioned medical vision-language models actually respond to images.
Open paperTests a program-based solver for clinical calculators, separating tool use, formula access, and execution reliability from model arithmetic.
Open paperMaps how bias enters AI systems and perpetuates healthcare disparities across data, design, deployment, and evaluation.
Open paperFrames clinical AI as a living system that needs continual monitoring, updating, and quality improvement after deployment.
Open paperShows why performance in one setting may not transfer to another and why external validation is central to clinical readiness.
Open paperAudits whether a general-purpose medical model reproduces racial and gender biases in healthcare responses.
Open paperExamines the ethical risks of large language models in medicine, including accountability, safety, and evaluation.
Open paperIdentifies common pitfalls across the electronic health record data lifecycle that can undermine reliable clinical AI.
Open paperCalls for operational fairness metrics and practices that make equity measurable in machine learning for healthcare.
Open paperConnects reproducibility, transparent reporting, and trustworthy evidence in digital medicine.
Open paperExplains why technical accuracy alone is not enough to make AI useful, safe, or equitable in healthcare.
Open paperAdvocates embedding ethical reflection throughout AI development rather than treating governance as a final checkpoint.
Open paper03 / Team
DOJO is developed by an interdisciplinary team working across clinical research, machine learning, engineering, and public accountability.
Contact
For collaboration, contribution, or institutional questions, start a conversation through the DOJO repository.