The sheer waste that accompanies image and video genAI is actually terrifying.
I challenge OpenAI to put a "your carbon footprint" field next to each generation. If you have nothing to hide, more information is surely better, right?
Every SDXL image uses something like 0.5 Watt-hour. That means that 2000 sdxl images uses about 30 cents of electricity. If we multiply that by 30x for video (a fair assumption for models like minimax H3, which take that much longer to make video), that means you get almost 700 5-second videos for a dollar of electricity, which is about a kilo of carbon dioxide. More or less depending on where your electricity comes from.
Someone needs to name this fallacy. The existence of existing harmful norms does not, ex ante, justify introducing additional harmful norms. I propose naming it the "GenAI Defender Fallacy".
Separately--yes, I would love carbon receipts for everything. Such a system would be fascinating and do real good in this world.
What's the goal of complaining about it though? What's more likely, you persuade people to stop all of these habits that use electricity that burns carbon, or you convert the electricity supply to renewables?
It's a combination of argument from analogy and reductio ad absurdum. It's not a fallacy. You're free to dispute the validity of the analogy, or you can dispute the absurdity, which you did, by confirming that you would indeed love carbon receipts for everything.
Considering that especially leftists have spent literal decades hating on vegans for telling them about their individual footprint, there is no real fallacy here.
AI haters have somehow discovered personal responsibility. Awesome. They've never ever applied that to anything else they PERSONALLY do, though. It's only things others do. For any personal choices, suddenly it's all the governments fault or the corporations fault. No personal responsibility to be seen.
My 5-ish years of veganism have easily made up for a life time of prompting. Yet, I still would never generate video or images in part because of their impact.
How many of you non-vegan AI haters are going to be vegan now?
This is such a strange reply I don't even know what to do with it. Good on you for being vegan, and not prompting?
I have gone lacto-ovo vegetarian in the past year, so I'll take your 5 years of veganism as inspiration. Otherwise, I welcome people waking up to the harm they cause, regardless of their past hypocrisy.
Being flawed isn't the issue. Everyone is flawed. The issue is blaming people for problems like environmental damage, while at the same time needlessly causing more damage yourself. The problem is being hypocritical. Some people are certainly more hypocritical than others.
I mean, I actively advocate for better fuel efficiency standards but also haven't entirely eliminated red meat from my diet. If I'm discussing fuel efficiency standards and you call me hypocritical, even if I granted that, it doesn't seem like a useful turn of the conversation for anyone unless it's a rhetorical move to change the topic of conversation?
I've gone chicken only. Mostly because pork and beef are too expensive. It is very hard to get 150+ grams of protein a day on a vegan diet. And butter is in almost every delicious baked good. Avoiding it is like having the worst allergy.
Bad examples. Use meat consumption. It will blow any other kind of "non-essential" resource use out of the water, on ALL fronts.
It's really hard to believe that people care about the personal impact of their decisions, when it's only ever other peoples decisions that get talked about.
No one was talking about the ethical part (which you may or may not agree with). Carbon footprint is an actually measurable quantity that we can compare - what's your problem?
that would be actually pretty cool and would cause more awareness for a lot of people. but something like this would go in both directions i am afraid, as with everything there surely would be people trying to do carbonmaxxing just for the sake of it, kind of like breaking a highscore. i am sure people like this already exist anyway, though.
You must have your units really messed up because there's no way that's right. A GPU drawing 700W for 10s would be 1.9444 Wh or 0.0019444 kWh and that's if the GPU was only serving your image gen for the whole 10s which is unlikely, most of the time is probably queuing since even a much lower powered GPU in a desktop PC can generate the same image in well under 10s.
Do you have a source for 1-2.5kWh for 10s of content? It takes about a minute or two to generate, so you'd be talking about a GPU consuming 30_000W-150_000W, it's just not possible. Even if it was distributed (which I don't think it is), that'd be 30-100 datacenter GPUs running at 100% to generate one clip? There's no way that would make financial sense.
reddit used to shared how much gold purchases went towards their server costs (pretty sure it wasn't 100% accurate but at least good to an order of magnitude). It was approximately $1/hour and they served ~100M monthly active users. So I don't think those metrics would demonstrate how wasteful traditional data centers are like you think it would.
this would be amazing, absolutely. Especially if they were able to see specifically where my power comes from, the embedded carbon in the chipsets used to serve my requests, all the costs associated with the datacenter(s), etc. A weekly breakdown would be excellent.
I propose they use my body for compost to power their electricity turbines, in exchange for a advance in tokens while I am alive...I am overweight and that is good, more burning mass. We could settle at 2 million USD.
The waste overall for data centers is beyond terrifying. Enough water for 1.2 billion people. It’s like all the GenX and Boomers looked at the challenges younger people will face and decided make it as bad as possible start kicking them in the ribs for good measure.
A round trip cross country plane ride for one passenger is equivalent to about 500k - 1 million image generations. So all one has to do to offset their lifetime usage of image generators is forego one vacation. And this of course is only for current day carbon usage, that could go down (or up I guess but compute-wise usually things get cheaper over time).
The 'environmental' consideration is much less than the argument over 'paper or plastic'. It's going to be an issue of regulation and cost - not some arbitrary decision over 'whether I should make an image or just use text'.
Second - there are really no 'social or moral implications' for the most part. This is your dystopian delusion.
People now have 'automated Photoshop' and a slightly more power to express themselves. It's not going to change that much, and it's not a bad thing, it's a good thing.
I hate it because I hate having to look at these AI generated images. In fact I want an internet where I have to explicitly opt-in to see AI generated content in general.
You are joking, right? "It's fun. But people are having fun." You literally fake situations and you call that "fun"? Before calling someone else's comment ridiculous, imagine being someone being blackmailed by photos created by ChatGPT Images. You can't. Because it's all "fun", right? Jeez, people like you are totally crazy.
Do you want to know how many identical questions ChatGPT gets asked every week, recomputing its answer each time because people stopped sharing results? I don't, but I'm sure it far exceeds 3 billion. "How to center a div", "how to install this and that package", "what is this pop culture reference", "who is this famous person",...
Groan. Your comment is the most depressing thing I've read today. What a sad way to look at the world. People are doing stuff and having fun. Meanwhile a whole sub-cult is sucking from the doom straw and mumbling carbon, water, greed, blah. The sport of Golf consumes far more resources than any AI data center.
I’m not commenting on whether generative AI is especially resource-wasteful or not, but a lot of people online seem to have newly discovered the existence of data centers.
I was curious and tried to come up with a crude estimate - all of the golf courses in the US use approximately the same amount of energy (mowing, fertilizer, etc) as one 1GW DC does.
If you're not familiar of the pitfalls of this tech, and the fact that there are more ethical tools with which to do those things, I'm not sure what to tell you.
>Generating 1,000 images with a powerful AI model, such as Stable Diffusion XL, is responsible for roughly as much carbon dioxide as driving the equivalent of 4.1 miles in an average gasoline-powered car. In contrast, the least carbon-intensive text generation model they examined was responsible for as much CO2 as driving 0.0006 miles in a similar vehicle
But 3 billion pictures per week is equivalent to 12.4 million miles or 639.6 million miles extra per year.*
It's not about you. There are no 3 billion people each creating one picture per week. Its probably more like 10.000 assholes creating 2.5 billion pictures per week to satisfy some stupid online feed and the other users creating a picture per month on average.
That study is from 2023 when the cost to produce 1000 images was 2.91 Wh per image. More recent numbers for Stable Diffusion's 2025 models that puts it at 1.3 Wh per image and that number continues to decrease.
For context that means a days worth of image generation emits about the same amount of carbon dioxide as a single transatlantic flight (New York to London). There are approximately 1500 transatlantic flights per day.
So every day the people of Vermont emit more CO2 just commuting than all image generation through OpenAI per week. Nice. That does put it in perspective. It’s really efficient.
Ours cars are atrocious. We can have the most powerful artificial minds imagine 1000 images from simple prompts, and that takes as much energy as moving a human being 4 miles.
And that will get more energy efficient. Gas cars have barely budged.
This is an embarrassingly idiotic article. They're comparing against the least intensive text model, ie some million parameter model nobody uses. Stable Diffusion is a 3.5 billion parameter model, while the GPT models a billion people are using for text generation are over 10 trillion parameters. To say nothing of the differences in average context size usage. The actual ratio of SDXL to text generation pollution is probably literally reversed from what the article claims by lying with statistics.
(Note, however, that OpenAI's image model is much much larger than SDXL; however, we don't have precise numbers for it. Nonetheless, misinformation is misinformation.)
Some people feel the need to attach an image to almost everything the post - even in chat rooms (including slack at work). These images serve no purpose, but the poster feels it is useful - though it was often reaction gifs in the past I'm seeing a lot more AI content than I ever saw reaction gifs, maybe because of novelty or maybe because it's ultra personalised.
Sometimes images help to visualise something or to get a point across, but I see so many people who think it's necessary to reply to a discord message with a cat with human limbs doing a dance, or a photo of "themselves" climbing a mountain with the Rust logo to show them mastering Rust... Ok?
Image models are useful and I'm thankful for much better visual reasoning but it really really frustrates me the constant need to burn money for all of this slop.
Not OP but my guess is that they're already using bad AI images for their kebab business, and OP wants them to use better AI for their images (so they may actually lose fewer customers because the AI imagery isn't so awful)
Why do you care what your kebab shop uses for their menus? These people are trying to run a business. If this can allow them to clearly communicate prices in a well designed pleasant way, it's great. Why must they slave away in design and technology or pay someone to do it for them? I think it's great
If I eat at a restaurant and that restaurant has images of their food, I want those to be “real images” of what they will bring me when I pay. Otherwise it’s misleading.
I can't disagree more. A lot of shops in our city have done this and I actually miss the shitty photoshopped images that did not even look like the real life dish anyway. AI signs are way worse.
I thought that was parents point, that currently they're using so shitty AI generated logos, that even if we despise that as a concept, at least better models output slightly less sloppy shit. But re-reading it, I'm not sure that was the right initial reading.
I am wondering if in 2-3 years AI will be heads and shoulders above us in both images and text. Then will we continue to see the slop accusations? Maybe its slop is better than most of our work.
Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
The sad part is my mother would love "remixing" my old child photos of me.
Seriously, are these really the best examples they could come up with?
What is the point of having a fake picture of your dog in a costume? What’s the point of having a fake picture about being at a party?
The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
> Omg, I love how the first examples just show how easy you can fake things. Fake being at a party with your friends. Didn't make your bed, no problem, just fake it.
Well, that's like the entire point of social media anyway, isn't it? People there will _love_ this. Ugh ...
I’m less concerned about faking being at a party with friends and more with the 2006 timestamp on said photo!
Can you imagine finding a stack of photos in the basement with timestamps of the Before AI times and wondering whether they are real or just got swapped with generated and printed fakes? Scary!
It’s cool, folks… nothing bad is going to happen. Right? Right?!
Because she will pounder in old memories that never existed in the first place. Instead of cherrying the current times. Past moments of long gone time will be elongated beyond their actual existence.
Imagine a time where a small moment is actually a smaller amount experienced then the remixed one. At one point you will have more memories of fake events, and more emotional beats for said fake events than real ones.
AI videos showing you a second life if you just had chosen a different path.
There’s a wonderful Ted Chiang short story called “Anxiety is The Dizziness Of Freedom” that explores how people become enmeshed and addicted to alternate versions of how their life could have gone. It’s reminiscent of what you’re describing.
Because elderly people who never had a connection to tech or sci-fi will have no feel for, and therefore no resistance to the kind of trickery that seems fun but is ultimately corrosive. Cheerfully meddling with the artifacts of memory at that life stage is corrosive.
For thousands of years people have told embellished stories, and humanity has thrived upon it. Tall tales, fish stories. Heck, that's still the average person's experience unless they've really worked their critical thinking muscles.
Today's working adults are used to the short thirty-year "safe space" of smartphones and internet. We grew up in a temporary meta stability where "truth" was "recorded" and could be "relied upon", and now that the fundamentals are shifting, we're complaining that the physical world is amenable to storytelling and imagination once again.
Cry me a river. This is awesome and is a direct consequence of everything I ever wanted the future to be: magical creative superpowers. I wanted to graduate into the world that is emerging now rather than spend the first third of my career in incrementalism and slow progress.
2008 - 2020 sucked. What a total lull. AI is healing these things and putting us back on track for the jet pack future we grew up dreaming about. It's taking us back to a creative world without shitty platforms controlling what we say and do, and without a ceiling on what we can accomplish.
It feels like every day we're unwrapping a new present or several. Not small things, but reality-shattering things that fill me with inspiration to build and explore. It's so much fun.
To tell a story, to even exaggerate one is something completely different from providing "prove" it did really happen exactly the way it was told.
There are many many people that will take an image as an undeniable fact of the real world. While before you could fake(photoshop) things it took more time then 30 seconds.
I wouldnt regulate these aspects, because I believe pandoras box is already wide open. Regulating anything wouldn't change much.
I use gpt image 2 very heavily for my current project (ai UI design tool). The biggest improvement I'm seeing with this is in speed. I've generated around 50k images with gpt-image-2 via api, the the average latency has held at around 104s.
It's wild how much of a difference this is - images are coming in at around 35-40s. Very noticable, and makes a difference when you're iterating quickly: https://jjcm.org/gpt-image-2.5-speed.mp4
I've really enjoyed using AI to generate images. For example, Long time ago, I read an amazing five-book saga called Riverworld, and I used AI to recreate many scenes, places, and ideas from the story. Seeing the books come to life through hundreds of images was a great experience. Definitely one of the coolest things to enjoy in 2026. Try it with your favorite books.
The "composite party photo", while impressive, shows that still the miniscule details are being lost, like the teeth structure of the guy in the middle or the fact that the guy on the left is holding the cup with three fingers. Wondering why they chose this edit for the showcase.
I've also found the OpenAI image models to lose fine detail on image edits compared to Nano Banana or Flux models which faithfully retain input source image geometry and details. I was hoping this might be different but it sounds similar to previous OpenAI image models where something is lost in translation during image editing.
Even with LM Arena being flawed, this is significant. I was planning to do a writeup on the original gpt-image-2 as it crushed every complex image comprehension benchmark I had...I'm glad I procrastinated since ChatGPT Images 2.5 seems like an even better starting point to test out what these models can actually do nowadays.
Many people still think AI images output the wrong number of fingers on a regular basis. (EDIT: this was an ironic comment to make in hindsight and I own it)
So this is a valid point (and I admit I eat crow on my earlier statement), but not for that reason. In that photo, there are three fingers in front of the cup, but you would not expect 5 fingers because the way humans hold cups, the thumb will be occluded by the cup itself. That said, I don't think there is a way to hold a cup with both thumb and index occluded, so the correct number of fingers would be 4 in that case.
The Facebook Marketplace experience for used items has depreciated considerably with the widespread adoption of LLMs. Between placing attractive models in photos to help sell items to "improving" the visual condition of items that completely misrepresent the real condition of the item, it's becoming a trickier landscape to navigate to find what you're looking for.
You can get semi-decent sprite work out of GenAI models, but you still have to put in some manual work (scale normalization, palette reduction, grid alignments, etc). It's definitely not "out-of-the-box" yet.
The key is making them one frame at a time, rather than asking for the whole sheet. Adherence frame-to-frame with a reference image is really good, so just prompt with the previous + direction. I finally got Nano Banana to make fairly decent fluid animations that way.
Interesting. I'm surprised they haven't prioritised/deliberately trained for this more because of how useful it is to ask codex to generate some images/assets/sprites when building sill games. It would dramatically enhance how polished it's games were.
That said for single images the old model was already okayish for prototypes
Multi-turn consistency is the real feature here: the same face, the same hands, the same cup across a whole evening, which is more than most actual party photos manage.
Surprised to see no acknowledgement of how "AI Menu Slop" has become deeply associated with ChatGPT.
Everywhere I go now I see places that blatantly used ChatGPT for their menus or posters and they all look the same.
It feels almost existential to the service, I know its a bit of survivor bias but so many images you can tell immediately are ChatGPT vs other image providers, and I feel many people are sick of them.
It's not like a lot of these small-time outfits used their own photography to begin with.
They'd pick a kebab from a menu of professionally made kebab pictures the printer has in stock. Or worse, they'd take pics of their own plated food with a dead-centre point flash in a dark cupboard or something and you end up with awful-looking food, no matter how good.
I'm not sure if I'm working with ChatGPT Images 2.5 or 2.0 here...
The original photo I took was at https://www.reddit.com/r/Tovala/comments/1pfuwrb/meatloaf_pa... - it's a cheeseburger meatloaf with potato wedges taken with a phone camera. And I was going for a consistent documentation approach for the photograph, not trying for menu proper.
I suspect that someone doing a menu could take a properly plated meal from the kitchen (rather than me photographing on top of my oven) and have it get redone for a good image for a menu without fundamentally changing what is being served.
It's the great desemantification machine: want to look, read, sound, feel just like the median, without personal traits or individual expression? This is for you! (And people used to talk about communism and everything being same. Well, tech-broism actually achieved this.)
I think quite a lot of people can distinguish between AI and real (maybe 30%?) but almost no one is worried about distinguishing between ChatGPT and other image generators
What if I told you.. the average person LIKES ai slop images and prefers them to "regular" ones.
The same way the average person prefers McDonalds food to healthy food
Earlier today I couldn't get ChatGPT to draw a horse on the moon. Idea taken from Stable Diffusion's wikipedia page. I tried three times with that prompt, "a photograph of an astronaut riding a horse" and again with a horse on the moon. I got it to do another image. However, the original prompt just worked, though it didn't happen to choose the moon, it didn't say the moon.
ChatGPT Images is for me the most impresive of all models, but also the one I hate the most, because of what it's used for: either funny wasteful use, or nefarious use, basically nothing else.
Just like WordPerfect enabled your mom to create a professional-enough cover letter, ChatGPT allows your mom to send you a more or less visually pleasing virtual birthday card of you wearing a party hat.
Coming to HN to read comments on releases like this reminds me that there is a small, vocal minority that is pushing so hard for more and more features like this.
The majority of the world does not want or need any of this, yet the nerds in SF that can hardly hold a conversation with another human being are pumping it out as quickly as they can.
Most depressing thing I have read today
I challenge OpenAI to put a "your carbon footprint" field next to each generation. If you have nothing to hide, more information is surely better, right?
I await OpenAI's carbon footprint field.
Separately--yes, I would love carbon receipts for everything. Such a system would be fascinating and do real good in this world.
Whataboutism?
AI haters have somehow discovered personal responsibility. Awesome. They've never ever applied that to anything else they PERSONALLY do, though. It's only things others do. For any personal choices, suddenly it's all the governments fault or the corporations fault. No personal responsibility to be seen.
My 5-ish years of veganism have easily made up for a life time of prompting. Yet, I still would never generate video or images in part because of their impact.
How many of you non-vegan AI haters are going to be vegan now?
I have gone lacto-ovo vegetarian in the past year, so I'll take your 5 years of veganism as inspiration. Otherwise, I welcome people waking up to the harm they cause, regardless of their past hypocrisy.
People are flawed.
It's really hard to believe that people care about the personal impact of their decisions, when it's only ever other peoples decisions that get talked about.
No one was talking about the ethical part (which you may or may not agree with). Carbon footprint is an actually measurable quantity that we can compare - what's your problem?
Thankfully web browsing is no where near that.
Theres going to be a reckoning.
People are having fun making images.
Like they do all sorts of things on the web.
Making an image is as simple as visiting a web-page.
That's it.
Yeah ignoring the massive environmental, social, and moral implications I suppose it’s not a big deal at all.
Second - there are really no 'social or moral implications' for the most part. This is your dystopian delusion.
People now have 'automated Photoshop' and a slightly more power to express themselves. It's not going to change that much, and it's not a bad thing, it's a good thing.
Concerns about misinformation? Compute/energy use? Creative work Replaced by AI?
Yes, some people have fun, but most of the "fun" nowadays is just endless mindless consumption.
Or do you literally mean the entire sport of golf vs 1 data center?
How is the person supposed to know your opposing point of view if you don't tell it?
Costume generation gimmicks on the internet.
> the fact that there are more ethical tools with which to do those things
More ethical what? For what? Costume generation tricks?
> I'm not sure what to tell you.
I'm not trying to argue. I literally have no idea what you're talking about.
But I see your point too. I too like to generate cute and funny images and share it with my friends and family.
But this being done at scale can lead to overall degradation of the online experience.
>Generating 1,000 images with a powerful AI model, such as Stable Diffusion XL, is responsible for roughly as much carbon dioxide as driving the equivalent of 4.1 miles in an average gasoline-powered car. In contrast, the least carbon-intensive text generation model they examined was responsible for as much CO2 as driving 0.0006 miles in a similar vehicle
If anything now I feel LESS guilty about usage.
* if we trust the numbers in the other comment
For context that means a days worth of image generation emits about the same amount of carbon dioxide as a single transatlantic flight (New York to London). There are approximately 1500 transatlantic flights per day.
Ours cars are atrocious. We can have the most powerful artificial minds imagine 1000 images from simple prompts, and that takes as much energy as moving a human being 4 miles.
And that will get more energy efficient. Gas cars have barely budged.
Gas cars are the problem, not AI.
(Note, however, that OpenAI's image model is much much larger than SDXL; however, we don't have precise numbers for it. Nonetheless, misinformation is misinformation.)
Some people feel the need to attach an image to almost everything the post - even in chat rooms (including slack at work). These images serve no purpose, but the poster feels it is useful - though it was often reaction gifs in the past I'm seeing a lot more AI content than I ever saw reaction gifs, maybe because of novelty or maybe because it's ultra personalised.
Sometimes images help to visualise something or to get a point across, but I see so many people who think it's necessary to reply to a discord message with a cat with human limbs doing a dance, or a photo of "themselves" climbing a mountain with the Rust logo to show them mastering Rust... Ok?
Image models are useful and I'm thankful for much better visual reasoning but it really really frustrates me the constant need to burn money for all of this slop.
The sad part is my mother would love "remixing" my old child photos of me.
What is the point of having a fake picture of your dog in a costume? What’s the point of having a fake picture about being at a party?
The only use case I can think of for this is for someone who likes to make up stories and lie about what they’ve done. Is that really the target market?
Mark where you at the party today? Yes of course look at these pictures I took (╥﹏╥)
Except you need to have taken a photo of the unmade bed, uploaded it to ChatGPT, prompted for the 'fake made bed' version, and downloaded the image.
Surely it's less effort to just make the bed?
Well, that's like the entire point of social media anyway, isn't it? People there will _love_ this. Ugh ...
Can you imagine finding a stack of photos in the basement with timestamps of the Before AI times and wondering whether they are real or just got swapped with generated and printed fakes? Scary!
It’s cool, folks… nothing bad is going to happen. Right? Right?!
How is that sad? Why is your mothers joy sad?
Imagine a time where a small moment is actually a smaller amount experienced then the remixed one. At one point you will have more memories of fake events, and more emotional beats for said fake events than real ones.
AI videos showing you a second life if you just had chosen a different path.
People will be depressed from it.
For thousands of years people have told embellished stories, and humanity has thrived upon it. Tall tales, fish stories. Heck, that's still the average person's experience unless they've really worked their critical thinking muscles.
Today's working adults are used to the short thirty-year "safe space" of smartphones and internet. We grew up in a temporary meta stability where "truth" was "recorded" and could be "relied upon", and now that the fundamentals are shifting, we're complaining that the physical world is amenable to storytelling and imagination once again.
Cry me a river. This is awesome and is a direct consequence of everything I ever wanted the future to be: magical creative superpowers. I wanted to graduate into the world that is emerging now rather than spend the first third of my career in incrementalism and slow progress.
2008 - 2020 sucked. What a total lull. AI is healing these things and putting us back on track for the jet pack future we grew up dreaming about. It's taking us back to a creative world without shitty platforms controlling what we say and do, and without a ceiling on what we can accomplish.
It feels like every day we're unwrapping a new present or several. Not small things, but reality-shattering things that fill me with inspiration to build and explore. It's so much fun.
There are many many people that will take an image as an undeniable fact of the real world. While before you could fake(photoshop) things it took more time then 30 seconds.
I wouldnt regulate these aspects, because I believe pandoras box is already wide open. Regulating anything wouldn't change much.
No thank you
It's wild how much of a difference this is - images are coming in at around 35-40s. Very noticable, and makes a difference when you're iterating quickly: https://jjcm.org/gpt-image-2.5-speed.mp4
Why buy art when you can stare at a blank wall and imagine a beautiful painting?
https://www.reddit.com/r/HelloInternet/comments/f0wuej/are_y...
They probably didn't care to check those images in detail.
https://mordenstar.com/blog/gen-failures
Edit: I feel stupid I didn't see the original OP already mentioning three finger issue. I'll just leave it here.
gpt-image-2.5-sunburst: 1421
gpt-image-2.5-flare: 1399
gpt-image-2 (medium): 1381
mai-image-2.6: 1331
Even with LM Arena being flawed, this is significant. I was planning to do a writeup on the original gpt-image-2 as it crushed every complex image comprehension benchmark I had...I'm glad I procrastinated since ChatGPT Images 2.5 seems like an even better starting point to test out what these models can actually do nowadays.
Many people still think AI images output the wrong number of fingers on a regular basis. (EDIT: this was an ironic comment to make in hindsight and I own it)
There's literally an image of a dude with 3 fingers in the Composite Party Photo.
Last time I used a frontier image model it made me a seal with three hands so…
I had trouble weeding through all the marketing speak.
https://mordenstar.com/other/hobbes-animation
The key is making them one frame at a time, rather than asking for the whole sheet. Adherence frame-to-frame with a reference image is really good, so just prompt with the previous + direction. I finally got Nano Banana to make fairly decent fluid animations that way.
That said for single images the old model was already okayish for prototypes
> Pricing and availability
> Images 2.5 is rolling out today to ChatGPT, ChatGPT Work, and Codex users across all tiers on desktop, mobile, and web.
> GPT‑Image‑2.5 Sunburst and GPT‑Image‑2.5 Flare are available in the API. See pricing details here.
and the link to the pricing details is a page where i am either too dumb to find the pricing details, or they don't exist
Everywhere I go now I see places that blatantly used ChatGPT for their menus or posters and they all look the same.
It feels almost existential to the service, I know its a bit of survivor bias but so many images you can tell immediately are ChatGPT vs other image providers, and I feel many people are sick of them.
They'd pick a kebab from a menu of professionally made kebab pictures the printer has in stock. Or worse, they'd take pics of their own plated food with a dead-centre point flash in a dark cupboard or something and you end up with awful-looking food, no matter how good.
The original photo I took was at https://www.reddit.com/r/Tovala/comments/1pfuwrb/meatloaf_pa... - it's a cheeseburger meatloaf with potato wedges taken with a phone camera. And I was going for a consistent documentation approach for the photograph, not trying for menu proper.
https://chatgpt.com/share/6aa05d33-3e6c-83ea-9e32-f1371c1ad6... ( https://imgur.com/a/fPB5VKP for just the image)
I suspect that someone doing a menu could take a properly plated meal from the kitchen (rather than me photographing on top of my oven) and have it get redone for a good image for a menu without fundamentally changing what is being served.
Edit: finally, I have my horse on the moon. https://chatgpt.com/share/6aa0627e-b168-83e9-98e5-d5ad5a9cad... Though the horse doesn't have a spacesuit, neither does it have one in the Stable Diffusion wikipedia page.
The majority of the world does not want or need any of this, yet the nerds in SF that can hardly hold a conversation with another human being are pumping it out as quickly as they can.
I truly think we're fucked.