Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans
Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li, Cekai Weng, Yaodong Yang, Yue Luo, Li Liu, Jing Tang, Zhijing Shao, Zeyu Wang

TL;DR
This paper introduces a high-fidelity, real-time conversational digital human system that combines realistic visuals, expressive speech, and knowledge-grounded dialogue, enabling natural and immersive interactions.
Contribution
It presents an integrated system with novel retrieval-augmented methods and an asynchronous pipeline for seamless, real-time multi-modal digital human interactions.
Findings
Supports emotionally expressive prosody and wake word detection
Achieves minimal latency in multi-modal coordination
Enables context-aware, knowledge-grounded responses
Abstract
High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human system that seamlessly combines a visually realistic 3D avatar, persona-driven expressive speech synthesis, and knowledge-grounded dialogue generation. To support natural and timely interaction, we introduce an asynchronous execution pipeline that coordinates multi-modal components with minimal latency. The system supports advanced features such as wake word detection, emotionally expressive prosody, and highly accurate, context-aware response generation. It leverages novel retrieval-augmented methods, including history augmentation to maintain conversational flow and intent-based routing for efficient knowledge access. Together, these components form an…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsMultimodal Machine Learning Applications · Social Robot Interaction and HRI · Human Motion and Animation
