TY - JOUR T1 - Overview of ImageCLEFmedical 2025 – Medical Concept Detection and Interpretable Caption Generation A1 - Damm, Hendrik A1 - Pakull, Tabea M. G. A1 - Becker, Helmut A1 - Bracke, Benjamin A1 - Eryilmaz, Bahadir A1 - Bloch, Louise A1 - Brüngel, Raphael A1 - Schmidt, Cynthia S. A1 - Rückert, Johannes A1 - Pelka, Obioma A1 - Schäfer, Henning A1 - Idrissi-Yaghir, Ahmad A1 - Abacha, Asma Ben A1 - de Herrera, Alba G. Seco A1 - Müller, Henning A1 - Friedrich, Christoph M. JA - CEUR-WS.org/Vol-4038 - CLEF 2025 Working Notes Y1 - 2025 VL - 4038 IS - Paper 170 UR - https://ceur-ws.org/Vol-4038/paper_170.pdf KW - computer vision KW - Explainable AI KW - Image Captioning KW - Image Understanding KW - ImageCLEF KW - Multi-label classification KW - Radiology N2 - The ImageCLEFmedical 2025 Caption task follows challenges held from 2017–2024 and comprises three subtasks: concept detection, caption prediction, and a newly introduced explainability task. The goal is to extract Unified Medical Language System (UMLS) concepts, generate fluent captions from medical images, and provide humaninterpretable justifications for the outputs. This year’s edition used an enlarged version of the Radiology Objects in COntext version 2 (ROCOv2) dataset, which was expanded with new articles and the inclusion of the optical coherence tomography (OCT) imaging modality. For concept detection, the F1-score was used to evaluate predictions against UMLS terms. For caption prediction, evaluation was updated to a composite score averaging six metrics to assess both relevance and factuality. The new explainability submissions were manually judged by a radiologist. The 2025 task attracted 80 registered research groups, with 11 teams submitting a total of 149 graded runs across the three subtasks. Top-performing systems for concept detection were predominantly based on ensembles of Convolutional Neural Networks (CNNs). For caption prediction, a general shift towards fine-tuning Vision-Language Models (VLMs) was observed, with adapted architectures like BLIP leading to strong results across the new composite metrics. Finally, the inaugural explainability task saw initial submissions of post-hoc visualizations, establishing a baseline and clarifying the need for model-intrinsic explanations in future editions. ER -