Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education
Overall performance on the full benchmark in vision-capable models To assess end-to-end multimodal performance, we first evaluated vision-capable models on the full benchmark, defined as all retained items for the…