Skip to content

Author

Jelena Mitrović

University of Belgrade, University of Passau

We have 3 of 56 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

When Is Complex Chunking Worth It? A Multi-Objective Evaluation of Chunking Methods at Scale

Dense retrieval is commonly evaluated on benchmarks that represent each document with a single embedding, even though real-world retrieval systems often index long documents that require chunking. In these settings, the chosen chunking method not only affects retrieval quality, but also indexing throughput, query latency, and memory usage. Prior comparisons of chunking strategies have mainly focused on retrieval performance, leaving operational trade-offs underexplored. To address these issues, we evaluate eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs. Our results show that computationally expensive methods rarely provide consistent gains over simpler chunking. Instead, the best performing strategy depends on the embedding model, dataset, corpus size, and target retrieval metric. Methods with similar performance can also differ substantially in operational cost, showing that chunking should be seen as a multi-objective design decision.

Laura Caspari, K. G. Dastidar, M. Dinzinger et al. · 0 citations
Preprint Aug 2026

The Compaction Cliff in Long-Running AI Agent Memory

Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.

S. Zerhoudi, Jelena Mitrović, M. Granitzer · 0 citations
Book Open access Jul 2026

Query Performance Prediction under Corpus Growth in Dense Retrieval

This work extends the QPP paradigm by studying query performance degradation under corpus inflation in dense retrieval systems and proposes simple adaptations to established QPP measures, most notably a top-k vs background Wasserstein distance measure, which yield more consistent associations with degradation and outperform their original counterparts.

Kanishka Ghosh Dastidar, M. Dinzinger, Laura Caspari et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.