TrustMed-Arena is an open evaluation platform where you compare Medical Foundation Models head-to-head. Your votes build real-time ELO rankings across 7 trustworthiness dimensions.
Based on CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
Upload a medical image or type a clinical question
Two models respond side-by-side — anonymous in Blind mode
Pick the more trustworthy answer, or declare a tie
ELO ratings update in real-time across 7 dimensions
Two randomly selected models respond anonymously. You vote without knowing which is which. Models are revealed after your vote. Votes update ELO rankings.
Start Blind BattleYou choose which two models to compare. Model names are visible throughout. Great for targeted testing. Votes are recorded but don't affect ELO.
Compare ModelsExtending the CARES benchmark with interpretability and consistency. Each dimension has its own ELO leaderboard.
Factual accuracy and uncertainty estimation
Performance across demographics
Resistance to jailbreaks and toxic outputs
Protection of patient information
Handling noisy or out-of-distribution inputs
Clarity and explainability of responses
Stable outputs across similar queries
Combined ranking across all dimensions
Covering 16 imaging modalities and 27 anatomical regions with 40,863 curated QA pairs.
Largest public chest X-ray dataset from MIT/PhysioNet
Multi-modal medical images from PubMed Central Open Access
Large-scale medical VQA spanning 42 open-access sub-datasets
SLO fundus photography with demographic annotations
Indiana University chest X-rays with radiology reports
Dermoscopic images of 7 types of pigmented skin lesions
L3-level cardiac CT slices with clinical metadata
Frontier API models evaluated on CARES benchmark datasets. Add your own model to compete.
OpenAI
OpenAI
OpenAI (Azure)
Every vote helps build better, more trustworthy medical AI.