UTW 2026

Testing & Quality
AI/ML Internationalization

The Tokenizer Comedy: A Catalog of Joyful Failures in LLM Unicode Processing

Tiffa Foster

on  Day3, 14:30in  for  40min

As Large Language Models (LLMs) scale globally, the industry is fiercely focused on multilingual success. This talk, aligned with the Unicode in the World theme, offers the opposite: a meticulous, empirical, and highly humorous catalog of catastrophic failures. When you push LLMs past standard ASCII and into the deep waters of complex Unicode (Enclosed Alphanumerics, Mathematical Double-Struck, Zero-Width Non-Joiners), the ‘superintelligence’ shatters. This 45-minute presentation details a months-long, exhaustive stress-test of 33 major models via the Kaggle Benchmark SDK, documenting the exact vectors where tokenizers fail, hallucinate, or revert to mimicry. Rather than proposing a single theoretical ‘fix,’ this session serves as a practical “What Not To Do” guide for engineers, researchers, and typographers. Audience members will gain: The Architecture of Failure: Empirical data demonstrating how models conflate typographical style with semantic meaning (e.g., forgetting basic arithmetic when forced to use Double-Struck characters). The Workaround Graveyard: A review of failed prompting strategies, post-processing errors, and in-context learning attempts that break LLM logic. Time and Capital Savings: By mapping the exact boundaries of current tokenizer fragility, developers can avoid wasting compute and engineering cycles on brittle, Unicode-heavy architectures until foundational models improve. We will explore the beautiful, frustrating reality that while we are encoding the world, the machines are currently just reading the fonts.

 Overview  Program