Integrating generative AI into forensic workflows: A case study from the Netherlands Forensic Institute.
Abstract
Large Language Models (LLMs) such as ChatGPT promise efficiency gains in routine text-based work, but their use in forensic settings raises challenges for confidentiality, governance, and forensic validation. Prior studies have evaluated LLMs on single tasks within single forensic disciplines, leaving open the question under which conditions such models can be embedded in the day-to-day workflows of a forensic institute. This paper reports a three-month pilot at the Netherlands Forensic Institute, in which 50 forensic and support staff used Robin, a ChatGPT-4o instance hosted in a Trusted Cloud environment of the Dutch Ministry of Justice and Security, governed by a Data Protection Impact Assessment and a GDPR-compliant processing agreement. We report user-experienced applications gathered from semi-structured interviews, walk-in sessions, and follow-up mini-projects, and group them into three recurring categories: generating new text, reducing larger inputs to relevant information, and transforming text between styles, languages, or representations. Participants found Robin most useful for low-risk writing support, while expectations quickly extended towards system integration, document-grounded retrieval, coding assistance, and case-oriented support. We also discuss the technical, organisational, ethical, and legal conditions for sustainable deployment, and argue that workflow support and forensic reasoning require different levels of human oversight and discipline-specific validation. The pilot is presented as an exploratory institutional case study rather than as a performance evaluation of an LLM for forensic analysis.