Author
Listed:
- Amine Moujdi
(School of Computer Science, Nanjing University of Information Science and Technology, Nanjing, China)
Abstract
Large vision-language models have been used to create cross-modal retrieval systems, including CLIP, that have achieved large performance improvements but often act as black boxes, which makes it more difficult to use these models and apply them in more critical areas. This non-disclosure is a great impediment to the responsible implementation of such systems in high stakes applications. As a measure to counter this shortcoming, we suggest an explainability-based design with an embedded post-hoc interpretation modules as part of the CLIP retrieval pipeline. The framework provides sensible, dual- mode accounts of bidirectional retrieval tasks; to begin with, it produces visual heatmaps that both outline the regions in the image that have the strongest impact on a retrieval decision and to a second end, it deactivates word-level attribution to quantify the relative significance of textual tokens in the query or caption. As we will discuss later, with our implementation and subsequent evaluation of our system on Flickr8k one can see that we provide these interpretable insights whilst maintaining, and in fact slightly increasing the baseline retrieval accuracy of the vanilla CLIP model. The empirical evidence confirms that the incorporation of interpretability layers is not accompanied by the trade-off in terms of performance. Together, this work confirms that principled explainability mechanisms should be augmented to multimodal retrieval systems in order to foster trustful, responsible AI solutions. Based on the increased transparency, the approach prepares the foundation of more solid and trustworthy human-AI cooperation.
Suggested Citation
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bjf:ijltem:v:15:y:2026:i:3:a:2227. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Dr. Pawan Verma (email available below). General contact details of provider: https://www.ijltemas.in/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.