Remix.run Logo
▲ andai 2 hours ago

For me the really interesting part was:

> Haven’t we seen LLMs do this already?

> Yes, BabelTele (arXiv, June 2026) demonstrated that LLMs can encode text in compact, non-standard forms — omnilingual word fragments, symbols, emoji — that other models recover with high fidelity (99.5% semantic fidelity at 27.9% of original length, by their metrics), including cross-model transfer, agent memory, and multi-agent communication. It proves the general phenomenon: human readability is not a requirement for model-to-model text.

I remember people testing early GPT-4 (2023?) in similar ways, to compress text, it would emit a string of strange text, Unicode, emojis, but was able to decode the compressed version very reliably.

This seems to cut usage by another ~50%, at the cost of being incomprehensible to humans.

▲HPsquared 2 minutes ago | parent | next [-]

Good for chain of thought, perhaps?

▲apefulsin 2 hours ago | parent | prev | next [-]

Once, a GAN model that was trained to convert between satellite images and drawn maps was caught encoding the original satellite image in imperceptible dots

The field of ML is Goodhart's law reified. We might have temporarily forgotten some of the basics of the field amidst this LLM craze.

▲alwa 33 minutes ago | parent [-]

https://arxiv.org/abs/1712.02950 (2017)

…if you’re curious and missed that one like I did. Snack-sized paper with lots of satisfying visual examples.

▲Theory42 an hour ago | parent | prev | next [-]

Yeah! LLM compression is a spectrum between 'normal stuff we can read' and vectors. Depending on trust in the models and the desire for compression, there's a choice to be made on which formats you want to allow. Pretty interesting stuff.

▲andai 2 hours ago | parent | prev [-]

[dead]