unIIcode
Inventing new characters to fill the gaps in Unicode.
[2026]
The [unicode](https://en.wikipedia.org/wiki/Unicode) is a text encoding standard used by most modern computers to support 161 scripts used by the world's major writing systems.
Unicode encoding contains gaps between scripts and ensures that a new script starts from a '00'. This allows for new characters to be added to existing scripts without having to disrupt other scripts.
What would a character look like if it belonged to a script, but nobody had ever written it? Since 2021 I have been training machine learning models to fill these gaps with characters of their own.
An Alphabet of the Gaps
The latest iteration learns from every character Unicode has assigned, over a hundred thousand of them, each drawn as a 128 by 128 pixel image. A diffusion model learns what each script looks like and is then asked to invent a character for every empty code point, shaped by the characters that sit around it in the chart.
Printed in full, the result is an alphabet of the gaps. Around 120 pages where the existing characters sit in grey and the invented ones in red, each labelled with its place in the standard.

Seen on its own, an invented character is just a shape. Placed back between its neighbours, it starts to read like a letter that was always meant to be there.






A second model maps every character into a single latent space of shapes, conditioned on the script it belongs to. Asked to write the same letter in another script, it produces an A as it might have looked in Greek, Devanagari, Tamil or Thai.
Each script leaves a trail through this space. Following the characters of one block in the order Unicode assigns them, the path spreads across the whole space, weaving through the shapes of every other script.
The Exhibition
This work was invited by the Shenzhen Graphic Design Association to [PROMPT](https://www.sgda.cc/news-event/news/news/94.html), an international exhibition on design and AI, shown in its AI for Design section at K11 HACC in Shenzhen from July to October 2026. Four pages of the alphabet were printed as scrolls, each 594 by 1782 mm, and mounted on aluminium panels.



Installation photos courtesy of the Shenzhen Graphic Design Association.
The First Iteration (2021)
This character does not exist in any language. It was created by the ML model.

To train the model, I created 128*128 pixel PNGs of all the characters.



The Tamil script was used to generate more Tamil characters.
These were then 3D modeled and printed to create physical characters.
