A syllable-character collaborative model for enhanced Pinyin and Chinese recognition

Zeyuan Chen; Cheng Zhong; Danyang Chen

PMC · DOI:10.1371/journal.pone.0325045·July 7, 2025

A syllable-character collaborative model for enhanced Pinyin and Chinese recognition

Zeyuan Chen, Cheng Zhong, Danyang Chen

PDF

Open Access

TL;DR

This paper introduces a new model for Chinese speech recognition that improves accuracy by combining syllables and characters during training.

Contribution

The novel SCCM model uses phonetic elements and an ensemble approach to reduce recognition errors in Chinese speech.

Findings

01

The SCCM model reduces pinyin and character error rates compared to prior methods.

02

It achieves a 45.7% relative reduction in Character Error Rate on the AISHELL-1 dataset.

Abstract

In Chinese speech recognition, end-to-end speech recognition models usually use Chinese characters as direct output and perform poorly compared with other language models. The main reason for this phenomenon is that the relationship between Chinese text and pronunciation is more complex. Inspired by the learning process of Chinese beginners, who first master initials, finals, and pinyin before learning characters, we propose the Syllable-Character Collaborative Model (SCCM), which incorporates these phonetic elements into the training process. Additionally, we design a Pinyin-Ensemble module that employs an ensemble learning approach to reduce pinyin recognition errors, which in turn leads to a reduction in text recognition errors. Experiments on AISHELL-1 show that our approach not only reduces pinyin and character error rates compared to a prior end-to-end method using pinyin as…

Linked entities

Genes, proteins, chemicals, diseases, species, mutations and cell lines named across the full text — each resolved to its canonical identifier and authoritative record.

Species1

Homo sapiens(human · species)

Cell lines1

AISHELL-1— Mus musculus (Mouse) · Hybridoma

Chemicals1

CTC

Diseases1

SCCM

Figures50

Click any figure to enlarge with its caption.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques · Topic Modeling