“Diffusion models are notoriously weak at generating thin, continuous, terminating structures.”
We don’t usually know why anyone would use AI slop to sell something intended to look appetizing, particularly when the food is presumably right there to photograph. But we do know a bit about why AI is so adept at producing such exquisitely nauseating images.
There are many reasons why this AI food is so wrong, from the technical nuances of how AI systems generate images and the materials used to train them to the psychological baggage humans perceive the images through. Frequently the problem starts from the very outset. Many of the leading image generators produce images using diffusion. Put simply, diffusion models start with an image of pure noise — something like a screenful of static — and gradually remove noise little by little in order to create the requested visual. “This means that initially coarse structures are recovered first with fine texture details” coming at the end, explained Chris Russell, a professor of AI, government, and policy at the University of Oxford and an expert in computer vision.
Russell said that in many of these unsettling food images things have already gone wrong by the time those finer details are added. The model might get the basic structure of an object wrong at an earlier stage, then lump vivid texture details on top of that structure. “This is the same kind of failure as you see when a person is generated with six fingers instead of five,” Russell said. (This could explain the donut shrimp too.)
The model might get the basic structure of an object wrong at an earlier stage, then lump vivid texture details on top of that structure.
Even when the underlying structure is solid, finer details can go awry in their own ways, said Giovanbattista Califano, a behavioral scientist who studies responses to AI-generated imagery at the University of Naples Federico II in Italy. “Diffusion models are notoriously weak at generating thin, continuous, terminating structures,” he said. “Noodles, strands, and tendrils are exactly the kind of geometry that trips this up, so you get spaghetti-like artifacts bleeding into places with no anatomical or culinary logic.” In other words, once a model starts generating something like this, it can struggle to figure out where it should stop or what it should be attached to. Other repeating textures like bubbles and seeds are similarly hard for diffusion models to contain within sensible boundaries, he added, meaning they often spill into areas they should not be in. That helps explain why so many AI food images are so relentlessly noodly, unsettlingly patterned, and riddled with the kind of clustered holes that can trigger trypophobia.
It doesn’t help that AI has no idea what a sandwich actually is. Or a noodle. Or a burrito. It has no understanding of the objects it’s creating or the physical world they inhabit. It has learned, broadly, what these things tend to look like on a statistical level, but not why they look that way or how they’re supposed to behave. “AI image generation reproduces looks without proper knowledge about the world,” explained Roland Meyer, a professor for digital cultures and arts at the University of Zurich in Switzerland. The result is an approximation of food divorced from any understanding of the thing itself.
That lack of understanding can lead to some stomach-churning aesthetic interpretations by humans who view the images, said Michael Cook, a senior lecturer in computer science at King’s College London. Hence the ice cream that resembles cracked concrete, or burgers seemingly fashioned from rocks. AI models have no understanding of why food should not look like other non-food images in that way. “These textures might look totally normal if used in an architectural context,” Cook explained. “But they become wrong when we imagine it as edible food.”
The images used to train these models can compound the problem. “Because we know so little about the training processes of these systems, we don’t really know what mix of content they’re receiving, or what associations they’re making,” Cook said.
Food photography is often highly stylized, full of sharp contrasts, intense colors, glossy lighting, and exaggerated shapes. Sometimes the “food” being photographed isn’t even food. Meyer said AI models can pick up on these surface qualities and visual conventions, but reproduces them without understanding the context behind them. “In other words, AI image generation perfectly imitates the look of photography, but not its professional aesthetic strategies,” Meyer said. “That is ultimately what makes them so unsettling.”
... continue reading