A+

Start or pause the reading of the main content. Pressing Escape stops the reading.

Tip: Press the Escape key to reset all display settings.

I had an AI image generator create a photorealistic image of a mall in a shopping center. It was supposed to include stores, such as a pharmacy and an ATM, as well as signs saying “Special Sale,” “Information,” etc. But why did the image generator make so many mistakes with the text in the generated image? Even after several attempts, it still couldn’t spell the term “ATM” correctly. Time and again, despite all my correction instructions, it came out as “Gelderautomat.” The pharmacy had two “e”s in its store sign, and so on. When I tried to change individual words again using a prompt, the same error—or a new one—often reappeared in the image. It seemed to me as if the AI—the smartest of all intelligences—were illiterate. 

What the background is

I could hardly believe it. After all, most discussions these days tend to focus on when artificial intelligence will surpass human intelligence. And yet my AI assistant was illiterate? To solve this mystery, one must consider that humans and artificial intelligence function differently.

When I analyze the caption using my “human intelligence” and spot errors in the text, I do so based on my understanding of language. In other words, I distinguish between letters and combine them to form a word. But this is precisely what an AI image generator cannot do. It does not analyze the text based on its sequences of letters, but rather examines the pixel pattern and the resulting visual image. Images are thus treated as visual patterns composed of pixels. It does not recognize individual letters and the linguistic logic behind them, but rather the image of a word. On the one hand, this form of image analysis makes it possible, for example, to still recognize a logo even when the individual letters are difficult to distinguish. However, it is of little help when it comes to correcting errors within a text. For AI systems, letters are merely visual shapes within an image. The AI has learned what text looks like but cannot decipher it linguistically. Letters and words are thus treated as images that look like letters; however, the linguistic logic is not accessible to such a system.

Where People Are (Even) Better 

It is reasonable to assume, however, that AI will learn in the future to better recognize the linguistic logic of a text and incorporate it into the image-generation process. Nevertheless, it seems comforting to be able to say today that the players behind this highly acclaimed artificial intelligence are, in fact, illiterate. But, to put it less polemically: AI can only give us back what has already been built into the algorithms as analytical capability. True, AI can continue to fantasize, as expressed in its hallucinations. But that, in turn, only leads to new images that were once again created against a backdrop of visual pixel patterns—and not based on linguistic knowledge. That is precisely the point where I nearly despaired, only to end up with a usable image after all. When I asked the AI for advice, it recommended commissioning an image without text—and then inserting the text manually myself using a specialized graphics program.

Write a comment

Your email address will not be published. Required fields are marked with *