Distributed Open Justice Oversight

Clinical AI should be tested like a system, not celebrated like a score.

DOJO is an open, community-driven platform for continuous adversarial evaluation of health AI—from the data that shapes a model to the workflow in which people rely on it.

DATA + MODEL + WORKFLOW

One system. Continuous evidence.

01 / What is DOJO

MIT Critical Data

A shared evaluation layer for clinical AI.

Distributed Open Justice Oversight is an open, community-driven platform for adversarial evaluation and remediation of health AI systems.

It helps researchers, clinicians, and evaluation teams examine the full system—not only a model score—before and during clinical use.

01Open

Shared tools, challenge sets, and evidence that can be reviewed and improved.

02Local-first

Designed to work near the data and within institutional security requirements.

03Community-driven

Clinical, technical, and lived expertise shape what gets tested.

02 / Architecture

From data to care, evaluated as one system.

Three connected planes make failures visible at the point where they enter the system.

01

Data plane

Checks cohort integrity, missingness, documentation inequity, leakage, and spurious signals.

03

Clinical workflow

Tests interoperability, infrastructure, human-AI interaction, and readiness in context.

01Define riskFreeze the intended use and threat model.
02ChallengeRun controlled stress tests across the planes.
03ObserveTrace findings to evidence and provenance.
04AssureReport readiness, limits, and next action.
Evidence objectIdentity · Challenge · Observation · Assurance · Lineage

Every consequential claim stays linked to the model, data, evaluator, environment, and re-test that produced it.

Selected publications from MIT Critical Data / DOJO

Research that informs the DOJO evaluation layer.

Selected work from the MIT Critical Data research program and DOJO team, spanning data quality, model evaluation, bias, safety, and clinical deployment.

02Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Cajas Ordóñez et al. · arXiv preprint · 2026

Examines how shortcuts spread through clinical multi-agent systems and why independent oversight is needed to detect benchmark gaming and socially plausible errors.

Open paper
03ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Cajas Ordóñez et al. · arXiv preprint · 2026

Introduces paired image-swap audits to test whether report-conditioned medical vision-language models actually respond to images.

Open paper
04Towards a Deterministic Math Solver for Clinical Language Models

Ocampo Osorio et al. · arXiv preprint · 2026

Tests a program-based solver for clinical calculators, separating tool use, formula access, and execution reliability from model arithmetic.

Open paper
05Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review

Celi et al. · PLOS Digital Health · 2022

Maps how bias enters AI systems and perpetuates healthcare disparities across data, design, deployment, and evaluation.

Open paper
06Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare

Feng et al. · npj Digital Medicine · 2022

Frames clinical AI as a living system that needs continual monitoring, updating, and quality improvement after deployment.

Open paper
07The myth of generalisability in clinical research and machine learning in health care

Futoma et al. · The Lancet Digital Health · 2020

Shows why performance in one setting may not transfer to another and why external validation is central to clinical readiness.

Open paper
08Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study

Zack et al. · The Lancet Digital Health · 2023

Audits whether a general-purpose medical model reproduces racial and gender biases in healthcare responses.

Open paper
09Ethics of large language models in medicine and medical research

Li et al. · The Lancet Digital Health · 2023

Examines the ethical risks of large language models in medicine, including accountability, safety, and evaluation.

Open paper
10Leveraging electronic health records for data science: common pitfalls and how to avoid them

Sauer et al. · The Lancet Digital Health · 2022

Identifies common pitfalls across the electronic health record data lifecycle that can undermine reliable clinical AI.

Open paper
11Equity in essence: a call for operationalising fairness in machine learning for healthcare

Gichoya et al. · BMJ Health & Care Informatics · 2021

Calls for operational fairness metrics and practices that make equity measurable in machine learning for healthcare.

Open paper
12The reproducibility crisis in the age of digital medicine

Stupple et al. · npj Digital Medicine · 2019

Connects reproducibility, transparent reporting, and trustworthy evidence in digital medicine.

Open paper
13The “inconvenient truth” about AI in healthcare

Panch et al. · npj Digital Medicine · 2019

Explains why technical accuracy alone is not enough to make AI useful, safe, or equitable in healthcare.

Open paper
14An embedded ethics approach for AI development

McLennan et al. · Nature Machine Intelligence · 2020

Advocates embedding ethical reflection throughout AI development rather than treating governance as a final checkpoint.

Open paper

03 / Team

Research Team

DOJO is developed by an interdisciplinary team working across clinical research, machine learning, engineering, and public accountability.

Principal Investigator

Multidisciplinary Team Researchers

Contact

Help build the evaluation layer clinical AI needs.

For collaboration, contribution, or institutional questions, start a conversation through the DOJO repository.

Leo Celi lceli@mit.edu
Sebastian Cajas asebasmos@mit.edu