Early bird registration is open now through August 1st. Use code UTW2026Early when registering.
Check back soon for new sessions still being added!
The Tokenizer Comedy: A Catalog of Joyful Failures in LLM Unicode Processing
As Large Language Models (LLMs) scale globally, the industry is fiercely focused on multilingual success. This talk, aligned with the Unicode in the World theme, offers the opposite: a meticulous, empirical, and highly humorous catalog of catastrophic failures. When you push LLMs past standard ASCII and into the deep waters of complex Unicode (Enclosed Alphanumerics, Mathematical Double-Struck, Zero-Width Non-Joiners), the ‘superintelligence’ shatters. This 45-minute presentation details a months-long, exhaustive stress-test of 33 major models via the Kaggle Benchmark SDK, documenting the exact vectors where tokenizers fail, hallucinate, or revert to mimicry. Rather than proposing a single theoretical ‘fix,’ this session serves as a practical “What Not To Do” guide for engineers, researchers, and typographers. Audience members will gain: The Architecture of Failure: Empirical data demonstrating how models conflate typographical style with semantic meaning (e.g., forgetting basic arithmetic when forced to use Double-Struck characters). The Workaround Graveyard: A review of failed prompting strategies, post-processing errors, and in-context learning attempts that break LLM logic. Time and Capital Savings: By mapping the exact boundaries of current tokenizer fragility, developers can avoid wasting compute and engineering cycles on brittle, Unicode-heavy architectures until foundational models improve. We will explore the beautiful, frustrating reality that while we are encoding the world, the machines are currently just reading the fonts.