To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
The short-term goal of a tool like this is to sell products. The more ambitious long-term goal is to shift cultural norms, blurring the lines between advertising and reality until the question you're asking is no longer consciously asked. At least, not by average people, and not at the point of purchase.
I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
I've been struggling with this question myself. But isn't this a model training/use problem (a.k.a skill issue )? Isn't there a way to make these models be faithful to how people will actually look?
Many product/model shots actually use clips to make the clothing look like it's perfectly fitted to the body. [1] This image is actually over 15 years old at this point. I think there should be laws that prevent this because it veers into false advertising though others believe it's alright because you can theoretically tailor the clothes to fit like this.
Regardless, I'd rather see real clothing on a real person when it comes to my purchasing decisions. I buy a lot of vintage clothes online and I've noticed a dramatic uptick in AI images of models wearing the clothes. I've never once bought from those sellers because it feels disingenuous. Sometimes they have fake runways which is actual false advertising because it makes the item appear more expensive than it really is. I've also noticed that the AI models' body types are always thin even if the item is a L or XL. Needless to say, the AI isn't showing me what an XL looks like on a small model; it's showing what a small model would look like if the item fit perfectly.
I've noticed something similar in Facebook marketplace ads for used furniture. Most of the images are AI generated to look like a Pottery Barn catalog, then the last image will be the actual item, full of scratches and other damage, sitting in a messy garage.
I've started to see this on Etsy and Wayfair too, where there will be a listing that is clearly just MDF flatpack being resold from China, but the AI-generated images wildly exaggerate the proportions of it.
Ironically, ChatGPT is decently good at ferreting these out. Like I sent it a screenshot of that listing and it not only helped me find where the original item was for sale, but also pointed out how the dimensioned diagram shows it as being just 49" tall, whereas the "in real life" image looks like it's at least six feet, based on it coming up over the top of the picture frame.
I ended up engaging a local woodworker to make me a piece like it instead.
Yes and even if all they do is ask it to stage an empty room picture with furniture and decorations it will make improvements like adding a doorway to a room that doesn’t exist.
The worst I've seen changed the view out a window from an alley to a private garden. Most also add natural light that isn't possible, and in some cases I'm convinced the generate image is depicting the space as larger than it actually is, with furniture and spaces that wouldn't fit in reality
I once heard about a company that specialized in making furniture about 10% smaller than normal for use in show rooms to make the spaces look larger. Not surprised they do it with AI if they can.
Wide angle lenses have been a thing in real estate photography since forever, but you would always have the reference of the furniture to ground your perception. Having furniture be slightly downsized is diabolical.
> But these models will always make the clothes fit your body and show you in flattering light and so on
This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.
Good or bad, I’m pretty sure showing things in the best light has always been the point of “marketing”. There is a reason ads aren’t filled with ugly people with misshaped bodies, and it isn’t because the intention is to reflect reality, so what you allude to as problematic is what a lot of businesses call a feature.
I bet if you had a way to easily train a LORA on you trying on various clothing types the models would do pretty well but I agree without that I can't think of a way to get it to work. Reminds me of the flattering mirrors scam https://finance.yahoo.com/news/company-swears-controversial-...
I use nanobanana 2 to test changes in paint and flooring/tiling to great success. And the real pro move is taking those images to a designer to tweak the remaining 20% or so.
The cynic in me says this is just an advancement and logical continuation from the known problem of sycophantic behavior in text-to-text chatbot format LLM, to image generation models.
There's definitely a predatory "everything will look good", but you can also leverage exactly that part for your own purposes - I've done pre-shopping a couple of times by asking for something like "a grid of 9 versions of my photo, wearing different types of X that look good". Definitely helped with choosing a good style.
i'd like it to notice things i would miss, like "this is ring-spun shirt, so it will sit like this on your torso" or "these pleats will require you to iron them" etc
That's hilarious. These keywords are applied globally, even on pages like https://qwen.ai/usagepolicy, where they very kindly ask you not to use their products for sexual content.
who made qwen stefani's dress
spiderman into the spider verse did qwen meet Peter
qwen stacey porn
do blake shelton and qwen stafani have children together
I assume they use a SEO tool that automatically adds these meta keywords to optimize for some search engines.
It apparently adds common search terms that contain words like "qwen". This evidently includes possibly mistyped searches for "gwen" or "ben" in a NSFW context.
Maybe someone knows more about how such SEO tools work, and where they pull the data from.
Wow, crazy things - "shemale porn", "spider fucks venom", and various porn URLs like tushy.com. What the hell did they do to their AI? Is that the training corpus or intended use case?
Surely it was meant to be "Gwen" (the protagonist's cousin and the other main character in the show). I guess it was somehow (incorrectly) assumed to be a misspelling of Qwen and included in the tags. Or perhaps many people were misspelling Gwen as "Qwen" and it was all hoovered up.
GPT Image 1 ended up with a yellow tint without training on another image-generation model's output. It's just that humans like pictures with a soft sunset glow, and this is a very easy global signal for a preference model to pick up on, and for a image-generation model to imitate. So optimizing for aesthetic appeal makes everything slightly tinted by default, unless you make sure to countersteer.
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
that's a very interesting point to test new models. I speak an Indian language called "Tamil" and have always tested new models with Tamil but also have tried little bit of Arabic (quranic verses) with previous GPT image 2 and Nanobanana pro and they have nailed it. Don't know if it was because of extensive training data.
I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
My wife has taken all her recipes and fed them through ChatGPT image gen to make zine pages and they’re really cool! She’s building a recipe book for the kids so they’ll know all the recipes from their childhood.
…but the illustrative diagrams are a simulacrum; if you ask Qwen, or any image-generator, for an “accurate” poster-design featuring a representation of a model of an atom and explaining its constituent parts I expect you’ll get an imitation-airbrush rendering of red, blue, and grey table-tennis balls orbiting in perfect circles; you might get an electron-shell diagram if you’re lucky. What you won’t get is anything remotely related to probability-clouds.
Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…
This is the image I got for the same prompt: https://jumpshare.com/s/mqqBdl7U59FWPiEXjwoM. It's more like high-school level, but not bad. I can imagine a collage professor can improve the prompt the create a more accurate and detailed diagram
This seems like something that could be solved by asking an LLM to write the prompt for the image model. You can also feed in the output of an image model into an LLM and ask it to check it/make improvements.
It is a reasonably non-opinionated model. My usual test is to ask these models to generate comic book and cartoon characters that hosted image generation models refuse to generate. I think that text is just CYA legalese.
The right models and LORAs can get rid of it right now. People apparently really like everyone to look like over exposed over filtered mannequins, so the big companies target that look.
There is an IRL phenomena, glass skin skincare and makeup, where the person's skin is evenly flat and toned and glossy. Did you think they are more plasticky than IRL models?
the natural endpoint of this technology is product photos that look better than the actual product, which is going to make unboxing videos the last remaining source of truth on the internet
> In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
Wow, it displays Korean properly without breaking.
But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
This model isn't supposed to contain all the numerical data. It will give you a (usually) matching graph transformed from one you provide or from a table of information you provide. Or you can pipeline from an LLM doing research on that data first. But expecting an image gen model to get you GDP info has got to be one of the worst possible approaches.
> Especially text rendering
That's true though. I still got some completely fried letters in headings.
I am not expecting it, it's the Qwen team is claimingthey can do much harder tasks than this, like rendering a consistent page of a maths paper, or creating true to fact explainers.
It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.
I guess you should use a traditional graphing library for your presentation slides for now.
I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.
Regardless, I'd rather see real clothing on a real person when it comes to my purchasing decisions. I buy a lot of vintage clothes online and I've noticed a dramatic uptick in AI images of models wearing the clothes. I've never once bought from those sellers because it feels disingenuous. Sometimes they have fake runways which is actual false advertising because it makes the item appear more expensive than it really is. I've also noticed that the AI models' body types are always thin even if the item is a L or XL. Needless to say, the AI isn't showing me what an XL looks like on a small model; it's showing what a small model would look like if the item fit perfectly.
[1] https://www.primermagazine.com/wp-content/uploads/2011/02/St...
Here's a recent example: https://www.etsy.com/ca/listing/4509158065/corner-wall-shelf...
Ironically, ChatGPT is decently good at ferreting these out. Like I sent it a screenshot of that listing and it not only helped me find where the original item was for sale, but also pointed out how the dimensioned diagram shows it as being just 49" tall, whereas the "in real life" image looks like it's at least six feet, based on it coming up over the top of the picture frame.
I ended up engaging a local woodworker to make me a piece like it instead.
This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.
Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.
Me, I’m a text-learner so I don’t get it at all. But I know people who get value.
The result was always someone extremely good looking
There’s going to be an entirely new class of mental disorders that will emerge from people being deluded by AI
Results are mixed, expensive, but it really feels you're few months off the next improvement to really nail it. It's already good enough.
Wonder what Qwen image will provide over nano banana.
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
Run this in console to see all the tags:
(i.e. the porn references)
It apparently adds common search terms that contain words like "qwen". This evidently includes possibly mistyped searches for "gwen" or "ben" in a NSFW context.
Maybe someone knows more about how such SEO tools work, and where they pull the data from.
Am i gregnant?
https://youtu.be/EShUeudtaFg
(See kids, it is possible to fight memes/racism with memes! And well... yeah, this really is racist.)
https://knowyourmeme.com/memes/bobs-and-vegana
> Please don't post comments saying that HN is turning into Reddit. It's a semi-noob illusion, as old as the hills.
https://pastes.io/uenL6X9K
It also seems to have an obsession with this celebrity, based on how many times ctrl-f for "stefani" turns up a result.
https://en.wikipedia.org/wiki/Gwen_Stefani
[whynotboth.gif]
Do you want to know more?
[ ] Yes [x] no
erm... what?
That said it's possible search engines in China or other countries might use it, but it's very easy to game so it doesn't really make sense
https://www.nytimes.com/2019/06/07/us/hate-groups-porn-consp...
porn is as American as apple pie.
So not at all then as Apple Pie is a traditional English desert. :P
https://en.wikipedia.org/wiki/Apple_pie
Apart from that, I made no judgement about porn in general, just about porn tags in Chinese backed AI websites.
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/i...
source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle
It's a shame they didn't share that prompt - it would make that demo more convincing.
I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.
Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
Any suggestions for the best open, non-opinionated model?
But: not open-source/open-weights, and no indication that weights/source will be released either.
As for the image model, wow...
And yet the Korean text is not accurate... [1]
[1] E.g. "드레스 컬렉션 dress collection" has vowels ㅔ mixed with ㅐ, "초웜한" should be "초월한 exceeding", "신키한" should be "실키한 silky", "디자언되다" should be "디자인되다 have been designed", "로얼" should be "로열 royal", and so on.
When I want to emphasize something, I tend to repeat it
Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.
> Especially text rendering
That's true though. I still got some completely fried letters in headings.
They can't.
It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.
I guess you should use a traditional graphing library for your presentation slides for now.