日本经济产业省在1978年建立的JIS X 0208编码标准中混入了一批来源不明的字符,研究人员后来将其称为"幽灵文字"[1]。这些神秘字符在随后的几十年里一直无人问津,直到1997年,有关研究人员发起了一项调查以追踪这些字符的真实来源[1]。
经过调查发现,这些幽灵文字大多源于编目工作中的人为错误[1]。例如,字符"妛"是在剪切粘贴"山"和"女"两字时因笔画误读而产生的[1];而字符"彁"唯一既无明确来源也无历史记录可循,最有可能是对"彊"字的误读[1]。这些字符最终被纳入了Unicode标准[1],成为全球计算机系统的一部分。
Japan's Ministry of Economy, Trade and Industry established the JIS X 0208 character encoding standard in 1978, which inadvertently included a mysterious set of characters whose origins remained unexplained [1]. These so-called "ghost characters" were discovered to be largely the result of cataloging errors accumulated during the encoding process [1]. When researchers launched an investigation in 1997 to trace the source of these anomalies, they found that many originated from misreadings of brush strokes that occurred during copy-and-paste operations [1]. The character '妛', for instance, was traced back to such a mistake when two separate characters were combined incorrectly [1].
Among the ghost characters, '彁' stands out as particularly enigmatic, lacking both a clear origin and any historical documentation; researchers concluded it most likely resulted from a misreading of the character '彊' [1]. Despite their dubious provenance, these ghost characters were eventually incorporated into the Unicode standard [1], becoming a permanent fixture of modern computing systems worldwide.