Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education
Overall performance on the full benchmark in vision-capable models
To assess end-to-end multimodal performance, we first evaluated vision-capable models on the full benchmark, defined as all retained items for the respective examination set (Fig. 1A,B). On M1 (n = 3747), Gemini 3.1 Pro achieved the highest overall accuracy at 99.31% (95% CI 98.99–99.53), followed by GPT-5.4 at 98.48% (95% CI 98.03–98.82) and Claude Opus 4.6 at 96.93% (95% CI 96.33–97.44). Among open-weight vision-capable models, Kimi K2.5 reached 96.72% (95% CI 96.10–97.24), followed by Qwen3-VL at 94.45% (95% CI 93.67–95.14) and Mistral Large 3 at 87.75% (95% CI 86.66–88.76).

Ranked accuracy for vision-capable models on the full benchmark, defined as all retained items, on (A) M1 and (B) M2, aggregated across examinations and weighted by item counts. The gray bar in each panel indicates mean student accuracy across the corresponding item set. Full numerical results (k/n, accuracy, and 95% Wilson confidence intervals) are provided in Suppl. Table S2.
On M2 (n = 3738), Gemini 3.1 Pro again achieved the highest overall accuracy at 98.37% (95% CI 97.91–98.73). Claude Opus 4.6 and GPT-5.4 followed at 97.30% each (95% CI 96.73–97.77). Among open-weight vision-capable models, Kimi K2.5 achieved 94.52% (95% CI 93.74–95.20), followed by Qwen3-VL at 93.18% (95% CI 92.32–93.94) and Mistral Large 3 at 91.33% (95% CI 90.39–92.19). Smaller vision-capable open-weight models, including Gemma 3 and Ministral 3, showed lower overall performance and are included in Fig. 1; full numerical results for all evaluated models are provided in Suppl. Table S2.
Performance on the shared text-only subset across all models
Because not all evaluated models supported image input, we next compared all models on the shared text-only subset (M1 n = 3227; M2 n = 3244), which provides the fairest common basis for cross-model comparison (Fig. 2A, B).

Ranked text-only accuracy for all evaluated models on (A) M1 and (B) M2, aggregated across examinations and weighted by item counts. Models are grouped by access type and size tier. The gray bar in each panel indicates mean student accuracy across the corresponding text-only item set. Full numerical results (k/n, accuracy, and 95% Wilson confidence intervals) are provided in Suppl. Table S2.
On M1 text-only items, Gemini 3.1 Pro achieved the highest accuracy at 99.63%, followed by GPT-5.4 at 99.50%, Claude Opus 4.6 at 99.26%, GLM-5 at 99.10%, Kimi K2.5 at 98.76%, DeepSeek V3.2-Thinking at 98.73%, and Qwen3-VL at 98.42%. Additional strong performance was observed for MiniMax M2.5 at 98.17% and gpt:oss 120B at 97.27%, whereas Mistral Large 3 and gpt:oss 20B reached 94.45% and 94.24%, respectively.
On M2 text-only items, Gemini 3.1 Pro again ranked first at 98.86%, followed by Claude Opus 4.6 at 98.15% and GPT-5.4 at 98.03%. The strongest open-weight models on this subset were GLM-5 at 96.42%, DeepSeek V3.2-Thinking at 95.84%, Kimi K2.5 at 95.75%, MiniMax M2.5 at 95.01%, and Qwen3-VL at 94.73%. gpt:oss 120B and Mistral Large 3 reached 93.74% and 93.16%, respectively, whereas gpt:oss 20B achieved 89.61% (Fig. 2A, B; Suppl. Table S2).
Several large open-weight models also achieved very high accuracy on the shared text-only subset, particularly in M1.
Modality-stratified performance and comparison with the human reference
We next quantified performance separately on the shared text-only subset and on image-present items for vision-capable models (Fig. 3A, B). Across the evaluated model inventory, accuracy on image-present items was consistently lower than on text-only items.

A, B Accuracy on the shared text-only subset versus the image-present subset for vision-capable models on (A) M1 and (B) M2, with the corresponding mean student accuracy shown as a reference. Each model is shown with paired text-only and image-present accuracies connected by a line; numerical labels indicate the absolute text–image gap in percentage points. Models are ordered consistently across panels and grouped by access type and size tier. C, D Error-rate ratio between image-present and text-only items, defined as (1 − accimage)/(1 − acctext), for the same models on (C) M1 and (D) M2. Values are shown on a logarithmic scale; the dashed reference line at 1 × corresponds to equal error rates. Full numerical results for absolute accuracies are reported in Suppl. Tables S3 and S4.
On M1, Gemini 3.1 Pro achieved 99.63% on text-only items versus 97.31% on image-present items, corresponding to a text–image gap of 2.32 percentage points. GPT-5.4 achieved 99.50% versus 92.12% (gap 7.39 points), and Claude Opus 4.6 achieved 99.26% versus 82.50% (gap 16.76 points). Among open-weight models, Kimi K2.5 reached 98.76% on text-only items but 84.04% on image-present items (gap 14.72 points), whereas Qwen3-VL showed a larger gap, with 98.42% on text-only items versus 69.81% on image-present items (gap 28.61 points). Mistral Large 3 reached 94.45% on text-only items but only 46.15% on image-present items (gap 48.30 points).
On M2, Gemini 3.1 Pro achieved 98.86% on text-only items and 94.44% on image-present items (gap 4.41 points). GPT-5.4 achieved 98.03% versus 92.59% (gap 5.43 points), and Claude Opus 4.6 achieved 98.15% versus 90.74% (gap 7.41 points). Among open-weight models, Kimi K2.5 reached 95.75% on text-only items and 85.65% on image-present items (gap 10.10 points), Qwen3-VL reached 94.73% versus 81.71% (gap 13.02 points), and Mistral Large 3 reached 93.16% versus 77.78% (gap 15.38 points) (Fig. 3A, B; Suppl. Tables S3 and S4).
To contextualize these model-level gaps against the human reference, we computed the corresponding mean student accuracies on the same item sets from the official item-level student correctness proportions. Mean student accuracy was 71.21% on M1 text-only items versus 64.27% on M1 image-present items (gap 6.94 percentage points), and 74.66% on M2 text-only items versus 70.74% on M2 image-present items (gap 3.92 points). Image-present items were therefore also harder for human candidates than text-only items, although direct comparison of absolute percentage-point gaps is limited by floor and ceiling effects.
We additionally quantified the modality-specific performance drop using the error-rate ratio, i.e. the ratio of the error rate on image-present items to the error rate on text-only items, which is invariant to the baseline accuracy level. For human candidates, the error-rate ratio was 1.24 × on M1 and 1.15 × on M2, indicating only a modest increase in error rate from text-only to image-present items. Across all evaluated vision-capable models, the error-rate ratio was substantially larger: on M1, it ranged from 7.2 × (Gemini 3.1 Pro) to 23.5 × (Claude Opus 4.6) for proprietary frontier models, from 9.7 × to 19.1 × for open-weight models ≥ 100B parameters, and from 2.9 × to 3.3 × for open-weight models < 100B parameters. On M2, error-rate ratios ranged from 3.8 × to 5.0 × for proprietary frontier models, from 3.2 × to 3.5 × for open-weight models ≥ 100B parameters, and from 1.8 × to 1.9 × for open-weight models < 100B parameters (Fig. 3C, D).
Taken together, these results show persistent text–image performance gaps across models. Image-present items were also more difficult for human candidates, but the modality penalty was disproportionately larger for models: in absolute terms, all evaluated models except Gemini 3.1 Pro on M1 showed a larger drop than students, and in relative terms the multiplicative increase in error rate was far larger for models (≈ 1.2 × for students versus several-fold for most models). Because this ratio normalizes each solver against its own text-only performance, this contrast indicates that image-present items pose an additional, model-specific difficulty beyond the general difficulty they also present to human candidates. The error-rate ratio is nonetheless sensitive to baseline accuracy and should not be used to rank individual models against one another.
Item-level correlation between human and model difficulty
We next assessed whether item-level model difficulty aligned with official student difficulty on eligible items from the shared text-only subset. For M1, item-level human difficulty and model difficulty showed a positive correlation (ρ = 0.318, 95% CI 0.286–0.346; n = 3227 items). For M2, the association was slightly stronger (ρ = 0.333, 95% CI 0.302–0.364; n = 3244 items). These findings indicate moderate, but incomplete, alignment between human and model difficulty at the item level (Suppl. Fig. S1).
Difficulty-enriched text-only subsets reveal residual model differences and human–model discordance
Given only moderate alignment between human and model difficulty, we next evaluated difficulty-enriched subsets of the shared text-only benchmark separately for M1 and M2 (Fig. 4). Human-hard-text was defined by an official student p-value of ≤ 0.30, whereas model-hard-text comprised items missed by at least two models from the top-tier comparison pool.

Performance of evaluated models on difficulty-enriched text-only subsets in M1 and M2. (A) M1 human-hard-text, defined as items with student p ≤ 0.30; (B) M1 model-hard-text, defined as items on which at least two models from the top-tier comparison pool answered incorrectly; (C) M2 human-hard-text; and (D) M2 model-hard-text. Bars indicate model accuracy, error bars denote 95% Wilson confidence intervals, and the gray bar in each panel shows mean student accuracy for the corresponding subset. The model-hard-text subsets showed stronger descriptive separation among top-performing systems than the human-hard-text subsets. Full numerical results are reported in Suppl. Tables S5 and S6.
In M1, the human-hard-text subset comprised 128 items, with a mean student p-value of 0.221. On this subset, GPT-5.4 achieved 99.22%, Gemini 3.1 Pro 98.44%, and Claude Opus 4.6 and GLM-5 each 96.88%. In M2, the human-hard-text subset comprised 197 items, with a mean student p-value of 0.194. On this subset, Gemini 3.1 Pro achieved 97.97%, GPT-5.4 and Claude Opus 4.6 each 95.94%, and GLM-5 92.89%.
By contrast, performance was more dispersed across models on the model-hard-text subset. In M1, this subset comprised 93 items, with a mean student p-value of 0.558. Gemini 3.1 Pro achieved 89.25%, GPT-5.4 84.95%, Claude Opus 4.6 80.65%, and GLM-5 75.27%. In M2, the model-hard-text subset comprised 323 items, with a mean student p-value of 0.633. Gemini 3.1 Pro reached 89.16%, followed by Claude Opus 4.6 at 82.97% and GPT-5.4 at 79.88%; among open-weight models, Kimi K2.5 achieved 60.06%, followed by Mistral Large 3 and Qwen3-VL at 53.87%, and GLM-5 at 52.01% (Fig. 4; Suppl. Tables S5 and S6).
Overlap between the two subset definitions was limited, with 12 shared items in M1 and 40 in M2 (Suppl. Table S7). Cross-classification further showed that most model-hard-text items were not human-hard-text (81/93 in M1 and 283/323 in M2), whereas most human-hard-text items were not model-hard-text (116/128 in M1 and 157/197 in M2).
To characterize what distinguishes these two off-diagonal cells, we asked five of the evaluated language models to describe, in qualitative terms and without knowledge of the selection criteria, the recurring properties and the competence probed by a sample of items from each cell (Set A, items that were model-hard-text but not human-hard-text; Set B, items that were human-hard-text but not model-hard-text; Suppl. Note 1). Across these five models, the two cells were distinguished primarily by the cognitive task type they probe. In M1, the models converged on contrasting the classification of described scenarios into concepts and terminology, with a marked psychosocial component (Set A), against the explanation of mechanisms and the application of scientific principles in the basic natural sciences (Set B); here, the task-type contrast was additionally aligned with a behavioral- versus natural-science disciplinary tilt. In M2, an equally consistent task-type axis emerged that cut across clinical specialties rather than tracking a single subject area: the models broadly converged on characterizing Set A as probing situation-appropriate, procedural, and norm-governed clinical decision-making, namely selecting the appropriate next action within realistic, management-oriented scenarios, with several models additionally noting a recurring legal, administrative, or ethical component. By contrast, Set B was characterized as probing recognition of specific clinical or laboratory patterns and recall of narrowly defined entities, tests, or standard therapies. Because item content is protected, these characterizations are reported as a convergent qualitative impression across models rather than as a quantified agreement metric (Suppl. Note 1).
Exam-level robustness and empirical probes for direct contamination
To assess robustness across individual examination administrations and to empirically probe for direct benchmark contamination, we examined exam-wise accuracy across the 12 M1 and 12 M2 examinations on the shared text-only subset. Reported pre-training cutoffs of the evaluated models span June 2024 (gpt:oss 20B and gpt:oss 120B) to August 2025 (GPT-5.4), allowing the same exam-wise data to support both a robustness assessment and a temporally structured contamination probe.
Across administrations, the strongest models showed consistently high performance with relatively narrow exam-wise ranges, whereas some lower-performing systems exhibited broader dispersion. On M1, exam-wise text-only accuracy ranged from 98.92% to 100.00% for Gemini 3.1 Pro, from 98.90% to 100.00% for GPT-5.4, and from 98.06% to 100.00% for Claude Opus 4.6. Strong open-weight models also remained comparatively stable, including Kimi K2.5 (97.70% to 99.61%), Qwen3-VL (96.93% to 99.29%), GLM-5 (98.47% to 100.00%), and DeepSeek V3.2-Thinking (97.32% to 99.64%). On M2, Gemini 3.1 Pro ranged from 98.10% to 100.00%, GPT-5.4 from 96.54% to 100.00%, and Claude Opus 4.6 from 96.93% to 100.00%. Among open-weight models, Kimi K2.5 ranged from 93.10% to 98.55%, Qwen3-VL from 93.54% to 96.74%, GLM-5 from 94.64% to 98.85%, and DeepSeek V3.2-Thinking from 93.92% to 99.64%. Full exam-wise heatmaps for all models are provided in Suppl. Fig. S2 (M1) and Suppl. Fig. S3 (M2).
To probe for direct contamination effects, we first examined whether models showed an accuracy drop across the pre-training cutoff boundary. Three evaluated open-weight models have reported cutoffs that fall between the Spring 2024 (administered March–April 2024) and Fall 2024 (administered October 2024) examinations: gpt:oss 20B and gpt:oss 120B (both June 2024) and Gemma 3 (August 2024). If direct memorization of pre-training-resident items contributed substantially to performance, these models should perform measurably worse on Fall 2024 than on Spring 2024 items. The observed Spring-to-Fall 2024 differences on the shared text-only subset were − 0.22 percentage points (pp) (95% CI − 4.32, + 3.79) on M1 and − 1.03 pp (95% CI − 6.14, + 4.04) on M2 for gpt:oss 20B; − 0.14 pp (95% CI − 3.57, + 3.21) on M1 and − 3.35 pp (95% CI − 7.24, + 0.14) on M2 for gpt:oss 120B; and − 0.86 pp (95% CI − 7.84, + 6.08) on M1 and − 2.80 pp (95% CI − 9.66, + 4.01) on M2 for Gemma 3. All six confidence intervals comfortably include zero. Differences of comparable magnitude were observed for proprietary frontier models whose cutoffs fall well after both 2024 administrations (e.g., GPT-5.4 M2 − 3.76 pp; Claude Opus 4.6 M2 − 2.62 pp; Gemini 3.1 Pro M2 − 2.26 pp), indicating that the modest Spring-to-Fall drops observed on M2 reflect an examination-specific difficulty difference affecting all models uniformly rather than a contamination-driven effect on cutoff-critical models. On M1, Spring-to-Fall differences were centered near zero across the inventory. This boundary-crossing test is necessarily restricted to the three open-weight models whose cutoff lies within our examination range; for the remaining models, direct boundary testing is not feasible.
Second, we examined whether examinations from earlier years showed systematically higher accuracy than recent examinations, as would be expected if older items had propagated through web-indexed materials over time. Per-model Spearman correlations between chronological examination index (1–12) and exam-wise accuracy on the shared text-only subset were small in magnitude and not consistently signed across the model inventory (Suppl. Table S8). For M1, correlations ranged from − 0.586 (GPT-5.4) to + 0.357 (gpt:oss 20B) with no systematic pattern by model class or reported cutoff. For M2, correlations ranged from − 0.336 (Mistral Large 3) to + 0.322 (gpt:oss 120B), again with mixed signs across models with both early and late cutoffs. This second probe is necessarily a coarser signal than the boundary-crossing test, but applies to the full model inventory and tests an alternative contamination mechanism (gradual web propagation of older material) rather than direct ingestion of specific item text.
Taken together, exam-wise performance was robust across the 2019–2024 administrations, and the two contamination probes are consistent with low direct benchmark contamination across the evaluated model inventory. Neither approach rules out indirect exposure through paraphrased or restructured material in secondary educational platforms, and reliable corpus-level contamination detection would require access to provider pre-training data, which is not publicly disclosed for most evaluated models.
Secondary analyses
Among the three evaluated base models without native thinking support, two-stage chain-of-thought prompting improved performance relative to direct constrained answering in both M1 and M2, with the largest gains observed in Ministral 3 (Suppl. Fig. S4; Suppl. Tables S9 and S10). In a separate within-family sub-analysis on a fixed recent text-only subset, performance across successive OpenAI model generations increased from 66.16% for GPT-3.5 Turbo to 98.30% for GPT-5.4, illustrating the rapid compression of benchmark performance toward ceiling levels (Suppl. Fig. S5).