🔤 The Ge'ez Tokenizer Playground
Experiment with a specialized Ge'ez-script tokenizer (Amharic, Tigrinya, Tigre, Ge'ez) and compare it with standard tokenizers from the GPT family — everything below runs locally in your browser, nothing is sent to a server.
Token count comparison — same input, every tokenizer
Ge'ez Tokenizer ours
TOKENS–
CHARACTERS–
🤖 Standard Tokenizer cl100k_base
TOKENS–
CHARACTERS–
Loading tokenizers… (the Ge'ez tokenizer is ~8MB, first load may take a few seconds)
If you find this tool helpful in your research or work, please consider citing our paper.
@inproceedings{teklehaymanot-etal-2025-movoc,
title={MoVoC: Morphology-Aware Subword Construction for Ge'ez Script Languages},
author={Teklehaymanot, Hailay Kidu and Fazlija, Dren and Nejdl, Wolfgang},
booktitle={Findings of the Association for Computational Linguistics: EMNLP 2025},
year={2025}
}