Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems
An unsupervised fine-tuning pipeline that harvests reasoning trajectories via in-context learning inference via in-context learning inference is proposed, enabling Large Language Models (LLMs) to access external knowledge and produce factual responses.
Mark A. Ferro, Oier López de Lacalle
· 0 citations