Skip to content
Open access

High-Dimensional Programmable Personas: A Comprehensive Framework for Synthesizing Customizable Role-Playing Agents

2026 · IEEE Access · Vol 14, pp. 143398-143418 · 0 citations · 36 references

Abstract

Role-playing language models are commonly conditioned on fixed historical or fictional identities, whereas many applications require personas assembled from user-specified attributes. We study this setting through a structured data-construction pipeline rather than a new model architecture. The pipeline maps four controls, namely career, aspiration, traits, and skills, to personal and social profiles, situates the resulting personas in generated scenes, and conditions multi-turn interactions on emotion and topic labels. Applying the pipeline with GPT-4-1106 produced SimsConv, a synthetic corpus of 68 personas, 1,360 scenes, and 13,971 dialogue turns. We then use standard supervised fine-tuning to obtain SimsChat variants based on LLaMA-3-8B-Instruct, Qwen2-7B-Instruct, and Qwen3-8B. On the in-domain SimsConv interview evaluation and the external WikiRoleEval and CharacterBench benchmarks, the fine-tuned variants improve most consistently in persona fidelity, character-specific recall, and rejection of out-of-scope questions. Component ablations associate these gains with scene grounding, detailed profiles, explicit emotion/topic controls, and the predefined attribute scaffold. Human ratings complement the automatic evaluation, although the synthetic origin of the training corpus, partial corpus audit, and use of model-based judges limit the strength of the conclusions. Accordingly, we present SimsConv as a reproducible test bed for controlled persona synthesis, not as evidence that synthetic data can replace human dialogue. The released artifacts are available at https://github.com/Bernard-Yang/SimsChat.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.