Rendered at 18:32:17 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
suprjami 2 days ago [-]
> The Astra pelicans are much better.
It's absolutely in the training data now.
Time to retire this (previously mildly amusing) benchmark as frontier models are now pelicanmaxxed.
conception 2 days ago [-]
What training data? Is there a big library of excellent SVG pelicans to train of?
suprjami 1 days ago [-]
Yes, all the people talking about this benchmark since OP created it.
orbital-decay 1 days ago [-]
Talking, sure. But have they created any good pelicans to train on? Pelicanmaxxing is clearly not a thing, which can be demonstrated by asking it for something else, or migrating to another domain (voxels, 3D modeling, and other things an LLM is supposed to be bad at, plus an unlikely combination of concepts). Looking at how relatively good for an LLM Astra is at Blender, it's obvious it was either purposefully trained for better spatial performance, or just generalizes better.
I actually think the pelican and similar tests are much less useful than it seems, but not because of pelicanmaxxing. They are supposed to serve as a vibe check of model's out-of-distribution performance, but do nothing to disentangle the generalization and memorization, which is the hard part. The combination of both is still useful though, and if you look at pelicans over time you'll see their quality is more or less correlated with that, with the exception of models specifically trained to produce vector graphics and 2D layouts.
graemep 20 hours ago [-]
Lots of SVG pelicans and lots of comments on what is good and bad about each.
williamDafoe 2 days ago [-]
1. Head tube should aligned with the fork. GPT-5.6 Luna is particularly terrible here.
2. Bikes are almost always depicted going rightwards so you can see the drivetrain. 4 failures in this.
3. Only ~7 of the 23 models achieved a diamond bicycle frame shape
4. Almost all the models have the bird directly over the crankset; as little as 4 or 5 have the birds sitting on a set-back bicycle seat.
amluto 2 days ago [-]
Is it just me or is it kind of remarkable that all the remotely competent pelicans on bicycles kind of look the same?
I assume that these have been in the training set for quite a while, but that they’re still a bit hard for the models because it’s not covered by the RL pipelines.
It's absolutely in the training data now.
Time to retire this (previously mildly amusing) benchmark as frontier models are now pelicanmaxxed.
I actually think the pelican and similar tests are much less useful than it seems, but not because of pelicanmaxxing. They are supposed to serve as a vibe check of model's out-of-distribution performance, but do nothing to disentangle the generalization and memorization, which is the hard part. The combination of both is still useful though, and if you look at pelicans over time you'll see their quality is more or less correlated with that, with the exception of models specifically trained to produce vector graphics and 2D layouts.
2. Bikes are almost always depicted going rightwards so you can see the drivetrain. 4 failures in this.
3. Only ~7 of the 23 models achieved a diamond bicycle frame shape
4. Almost all the models have the bird directly over the crankset; as little as 4 or 5 have the birds sitting on a set-back bicycle seat.
I assume that these have been in the training set for quite a while, but that they’re still a bit hard for the models because it’s not covered by the RL pipelines.