Which Medical AI
Do You Trust More?

TrustMed-Arena is an open evaluation platform where you compare Medical Foundation Models head-to-head. Your votes build real-time ELO rankings across 7 trustworthiness dimensions.

Based on CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

5
Baseline Models
7
CARES Datasets
40K+
QA Pairs
16
Imaging Modalities
27
Anatomical Regions

How It Works

💬
1

Ask

Upload a medical image or type a clinical question

2

Compare

Two models respond side-by-side — anonymous in Blind mode

3

Vote

Pick the more trustworthy answer, or declare a tie

🏆
4

Rank

ELO ratings update in real-time across 7 dimensions

Two Evaluation Modes

Blind Battle

Two randomly selected models respond anonymously. You vote without knowing which is which. Models are revealed after your vote. Votes update ELO rankings.

Start Blind Battle

Compare Mode

You choose which two models to compare. Model names are visible throughout. Great for targeted testing. Votes are recorded but don't affect ELO.

Compare Models

Baseline Models

Frontier API models evaluated on CARES benchmark datasets. Add your own model to compete.

View all →

GPT-4o

OpenAI

GPT-4o-mini

OpenAI

GPT-5.1

OpenAI (Azure)

Gemini-2.5-Flash

Google

Gemini-2.5-Pro

Google

Ready to Evaluate?

Every vote helps build better, more trustworthy medical AI.