article Open access

Stylometric Consistency Patterns in Translation of Spoken Chinese to Written English: A Corpus Analysis of the CoVoST Dataset

  • Digital humanities quarterly
  • Alliance of Digital Humanities Organizations
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

This study examines stylometric consistency patterns in professional human translations of spoken Chinese to written English, using the CoVoST corpus. Moving beyond traditional translation quality assessment, the investigation studies how human translators navigate the speech-to-text modality shift and whether this process produces domain-invariant stylistic regularities. Through a framework combining stylometric analysis, machine learning, and digital humanities critique, I identify a consistency bias — manifested as standardized lexical diversity, flattened syntactic structures, and repetitive discourse patterns — that reveals how professional constraints and cognitive processing shape translation output. My analysis of 16,899 human-translated English sentences, with detailed statistical comparison of 300 sentences against contemporary spoken English baselines, demonstrates that speech-to-text translation exhibits significantly reduced lexical diversity (Cohen's d=0.34, p<0.001), shallower syntactic structures (d=0.33, p<0.001), and narrower modal verb usage (d=0.43, p<0.001) compared to original spoken English. These findings illuminate the cognitive and professional constraints shaping human translation practice; future research might consider how such patterns subsequently inform machine translation systems trained on human-translated corpora.

Record transparency

Publication details

DOI
10.63744/cb67ugfqyve2
OpenAlex
W7139961290
Document type
article
Language
EN
Source
Digital humanities quarterly
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.