Are AI labs pelicanmaxxing?
A recent investigation explores whether major artificial intelligence labs are deliberately training their image models to generate specific, unusual concepts like pelicans riding bicycles. This quirky test serves as an informal benchmark for spotting whether developers are overfitting their models to handle bizarre user prompts.
The study used a rigorous testing methodology across seven different models, including versions from OpenAI, Anthropic, Google, xAI, Qwen, GLM, and DeepSeek. The researcher ran forty-eight different prompts combining eight animals and six vehicles through each model multiple times, and subsequently used other AI tools to help evaluate the visual outputs.
The findings show no evidence that AI labs are engaging in pelicanmaxxing. Models do not draw pelicans any better than other animals, nor do they render bicycles better than other vehicles. Furthermore, the combined images of pelicans on bicycles show no special improvement beyond what the models' baseline drawing skills would already predict, and the generated scenes do not appear to be memorized.
While one model from GLM showed a slight positive boost for the exact pelican-bicycle combination, the effect was too small to be statistically significant. Ultimately, the research confirms that AI labs are not artificially optimizing their generative models for this specific whimsical test case.