The Small Sample Challenge: Unsupervised Authorship Clustering in Croatian
Many forensic methods have been developed for uncovering authors of threatening letters to get rid of their anonymous nature. With the development of new tools for the digital creation of threatening letters, such methods became obsolete, yet new options for their digital analysis also became available. This paper covers a method for digital analysis of threatening letters written in the Croatian language using machine learning algorithms. First the techniques for extraction of style elements of a text are covered, focusing on the most influential elements, and then transforming the extracted data into a vectorized form which can be used in machine learning models. The models are built upon the autoencoder architecture in combination with two text style vectorization methods. Multiple such models are implemented and their results are compared and analyzed, focusing on their effectiveness and the quality of their results. Finally, additional possibilities for further development of machine learning models for the analysis of threatening letters written in the Croatian language are covered and their possible effectiveness explained.