Skip to content

Author

Josef Kittler

We have 3 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Probabilistic Embeddings With Evidence Learning and Refinement for Text–Video Retrieval

This paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the information granularity and abstraction levels between the two modalities hinder a reliable sample-level alignment. Moreover, redundant visual content, sparse textual descriptions, and temporal variability in videos introduce additional uncertainty, resulting in ambiguous matching and suboptimal performance. In this paper, we propose a novel method named Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR), which models video-text pairs as probability distributions and captures uncertainty through the evidence theory. Specifically, we perform distribution-level representation learning to resolve the semantic ambiguity of video-text pairs. To improve the alignment further, we introduce a distribution-based embedding refinement module to ameliorate the semantic consistency across modalities. The proposed PE2LR is able to pull positive sample pairs closer in the embedding space, while pushing the negative pairs apart. Comprehensive experiments on several benchmark datasets (including MSRVTT, DiDeMo, and ActivityNet Captions) demonstrate that our PE2LR achieves state-of-the-art search performance.

Donglin Zhang, Zhenghao Rao, Xintao Xu et al. · 0 citations
Aug 2026

A Knowledge-Imparting Generative Modelling Framework for Heterogeneous Federated Learning

Federated learning aims to provide security for client data privacy in practical machine learning applications. In principle, a global server aggregates the models produced by local clients to obtain a global model. However, the server is challenged when collaborating with local clients handling non-identically distributed data without authorisation to access it. Therefore, advanced solutions advocate the use of generative modules to deliver surrogate data to local clients during a server-agent interaction, without revealing private particulars. We argue that such a unidirectional transfer of surrogate patterns cannot fully represent and harmonise knowledge during the server-client interactions. To this end, we propose a knowledge-imparting generative modelling framework (FedKIG) based on adversarial feature learning and bidirectional knowledge distillation, to explore the potential of interactive generative modelling. In particular, Fed-KIG trains a feature discriminator for each local client to identify the surrogate patterns extracted by the global model. Under the supervision of the local feature discriminators, the server learns a global generator to generate pseudo samples that convey its global perspective. In this manner, local models are enabled to absorb global knowledge, thereby mitigating the training data divergence caused by data heterogeneity. In addition, we develop a bidirectional knowledge distillation strategy to support the entire learning process. This strategy breaks the rigidity of federated distillation by updating knowledge transfer between the server and the clients iteratively, thus overcoming the learning-forgetting issue. The proposed privacy-protected server-client interaction solution supports explicit knowledge generation for exploitation in federated learning. Extensive experimental results indicate that FedKIG significantly improves the generalisation performance and the stability of the model in heterogeneous federated learning scenarios.

Hong-Yao Chen, Tianyang Xu, Xiao-Jun Wu et al. · 0 citations
Jul 2026

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

It is demonstrated that prediction-level feedback substantially improves the reliability of training-free RVOS, with ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels.

Yuanjia Li, Tianyang Xu, Tao Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.