31 comments

  • jarofgreen 2 hours ago
    Also discussed here: https://news.ycombinator.com/item?id=49638353

    Original posts

    https://mathstodon.xyz/@andreasthom/117240535270608201

    https://mathstodon.xyz/@andreasthom/117240536885387540

    https://mathstodon.xyz/@andreasthom/117240537520615623

    This story is someone on bluesky screenshotting someone on twitter without a link. The person on twitter screenshotted the source on the fediverse without a link. This is getting daft.

    • netghost 1 hour ago
      And here we are on HN discussing it.
      • jarofgreen 54 minutes ago
        Discussion is fine, let's just try actually linking back to the source.
      • lubujackson 1 hour ago
        Screenshotted this thread; will repost on Reddit and 9gag for the updoodles.
    • pluc 1 hour ago
      You can pretty accurately fake a video these days, faking a screenshot is 7 year old level of skills.
    • zerozerotwo 1 hour ago
      A toot embedded in a tweet .
      • egypturnash 46 minutes ago
        A toot in a tweet in a... skeet? Is there any consensus on what to call "a post on bluesky"?
        • Qworg 11 minutes ago
          Skeets is what the community wants, but the company does not (for obvious reasons).
        • keeda 19 minutes ago
          Bleat?
  • GodelNumbering 0 minutes ago
    Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
  • sashank_1509 13 minutes ago
    Both things can be true:

    1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.

    2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.

    The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.

  • aaronharnly 33 minutes ago
    Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.

    My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.

    • bitexploder 12 minutes ago
      Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.
  • winfredJa 10 minutes ago
    https://x.com/markchen90/status/2097400166554993041?s=20

    that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

  • jrflo 2 hours ago
    I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

    > Improve the model for everyone

    > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

    It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

    • ColinWright 2 hours ago
      I refer you to this:

      https://news.ycombinator.com/item?id=49643556

      Quoting:

      > "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

      • asimpleusecase 2 hours ago
        Old Facebook trick - likely resetting that box each time the app is updated.
        • morkalork 1 hour ago
          "We've made some updates to improve the security and privacy experience" => "We've changed some of the options available and reset everyone to defaults"
      • unified101 1 hour ago
        In all fairness this is someone saying something. Misremembering happens. Unless we have something with a bit more evidence, the simpler explanation suffices.
        • tomrod 59 minutes ago
          In more accurate assessment rather than assuming no maliciousness nor incompetence, remembering also happens. Unless we fail as a society, the simpler explanation that "OpenAI is training on all data it can and resetting config toggles because it uses the same cohort of engineers that came from Meta and other FANGAMAAMMAM clones" suffices.
      • jrflo 1 hour ago
        I wasn't aware of that, definitely a shady practice if that's the case.
      • ProllyInfamous 1 hour ago
        >"reset this more than once ... to my surprise I found it re-enabled"

        At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.

        What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?

    • nmfisher 2 hours ago
      There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".

      I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

      • fritzo 1 hour ago
        Whoa that's a slippery slope! Next you'll want model runners to cite the data their models were trained on
        • gunalx 1 hour ago
          In fact we should though.
      • jrflo 1 hour ago
        But who gets credit then? Every mathematician who's work was read by an LLM during training? By that logic, we should put every published mathematician's name on the authorship of this paper. Sure, this guy should be higher up the list, but everyone's name should be on it by standard academic convention.

        But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.

        • didroe 40 minutes ago
          Does every mathematician get cited in every maths paper? I think it's pretty clear who should be cited.
    • omnicognate 2 hours ago
      Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
      • rfgplk 2 hours ago
        Under EU rules it doesn't constitute consent.
    • spindump8930 2 hours ago
      "Improve the model for everyone" can be implemented in so many ambiguous ways.

      https://news.ycombinator.com/item?id=49643513

      • ProllyInfamous 1 hour ago
        >>"Improve the model for everyone"

        e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers

        I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).

        ----

        My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."

        Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."

    • tyrabound 1 hour ago
      > We take steps to protect your privacy

      No mention what those “steps” are, success criteria, or whether they are successful by any navies at all … they take steps though… so it’s fine, and if we know one thing it’s that we can really truly trust someone off the likes of Sam Altman.

    • mannanj 16 minutes ago
      Yes but what about “analytical purposes” what does that cover and can you turn it off? I have found out you cannot. It’s the Trojan backdoor to your data.
    • fithisux 2 hours ago
      Ok, you shut it down, or that is what they make you believe. You give the instruction to shut down, you can't know if it has been applied.
  • ColinWright 2 hours ago
  • postalcoder 2 hours ago
    The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).

    People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:

      1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
    
      2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
    
      3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
    
      4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
    
    Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

    If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).

    edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.

    • ahsg17 1 hour ago
      > Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

      Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.

    • larodi 36 minutes ago
      > If it is found that OpenAI and other labs are not respecting the training opt out

      HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!

    • thevillagechief 33 minutes ago
      You know, I don't think I've ever accused anyone of being a shill. I've thought about it maybe a few times (daringfireball). This is going to be as close as I get. I don't know the facts in this case but I cannot believe the argument being made here with a straight face. Is it common sense that tools you use and pay for steal your work and profit off of it at your expense and without recognition? If this isn't the textbook definition victim blaming, I don't know what is.
    • SpicyLemonZest 1 hour ago
      This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.

      I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.

      • cma 1 hour ago
        I think Google does train on anything you put into docs if you aren't careful with the Gemini integration?
        • SpicyLemonZest 24 minutes ago
          Yes, this is a problem with modern AI systems in general. It's not just OpenAI, and if you know any artists you know this is why they're pretty vehemently opposed to all AI.
      • lowbloodsugar 52 minutes ago
        That’s … Googles entire reason for making these “you don’t pay with money” tools. Did you not understand that?
        • SpicyLemonZest 45 minutes ago
          What? I don't understand how you even came up with this idea, much less consider it so obvious to condescend about it. Do you have even a single example of a research project that got scooped because the Google Docs team forwarded their private documents to someone?
    • mittensc 1 hour ago
      Imagine OpenAI Astra model weights were made public because the datacenter they use had T&C that allows them to make them public

      Would that be ok in your mind?

      Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)

      Nobody would care if they provided published research that author made public same as a google search would offer that.

      • aurareturn 39 minutes ago

          Would that be ok in your mind?
        
        It would in my mind. Hopefully companies have looked through the agreement.
    • 1294827 1 hour ago
      [flagged]
    • chunky1994 17 minutes ago
      Why are we being so charitable to trillion dollar organizations here? If OAI keeps re-enabling the train model toggle on every app update to codex, does it also fall under "common sense opsec" to re-disable this toggle every time?

      Arguably you expect that unless you are explicit about providing permissions to these labs to use your data for training then your data is yours, and not theirs. Especially on a paid account (let alone an enterprise one). Why is the opt-out supposed to be "common sense opsec" rather than the opt-in should be common sense regulation?

  • alansaber 2 hours ago
    I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
    • calvbak 1 hour ago
      I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.
    • sigbottle 25 minutes ago
      In general, a lot of moral invariants that natural selection has rendered as "intuitive" to us are no longer intuitive or possible. These natural brakes are not braking.
    • utopiah 1 hour ago
      This is such a naive position though.

      The most successful companies of the last decade have precisely been ... selling usage data.

      Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.

      • alansaber 16 minutes ago
        Yeah but marketing companies are aggressively fingerprinting and stalking you to sell you snacks from japan, or oscilloscopes because they figured out you work in a lab, etc. Not to fuck you over by stealing your livelihood (which is what is happening to these mathematicians). It's on a whole new scale.
  • spindump8930 2 hours ago
    Reminder that there are degrees of "trained on conversations". From John Schulman:

    > pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

    > use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

    > use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

    source: https://x.com/johnschulman2/status/2097440545853637108

    • rfgplk 2 hours ago
      This would cease to be a problem if OpenAI remained true to their founding motto and... actually open sourced their training/inference pipeline.
    • Ydarbleoj 2 hours ago
      This is a reminder based on believing what these companies say.

      I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.

  • rfgplk 2 hours ago
    Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
    • jeremyjh 2 hours ago
      Public domain doesn’t mean anyone can assert copyright. It specifically means no one can.
      • voakbasda 1 hour ago
        No, it means you can use that work in the creation of new works, which can indeed be copyrighted.
      • cyanydeez 2 hours ago
        also, none of it means anything without the lawyers to back it up. Just like you can be a pedophile in the highest office of democracy and escape persecution.
    • krupan 1 hour ago
      What does "with no substantive human input" mean? All of the training data is human input, isn't it?
      • warkdarrior 35 minutes ago
        They also train on synthetically generated data.
  • int32_64 1 hour ago
    Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
    • SpicyLemonZest 1 hour ago
      No, that order was for a defined period that has ended.
  • remywang 1 hour ago
    People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.

    It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.

  • gentlerain 2 hours ago
    So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?

    How do people become that trusting?

    The phrasing itself is guilt tripping

    • ayewo 48 minutes ago
      To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training.

      From their docs[1] (archive copy is at [2]):

      > You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new conversations will not be used to train our models.

      > For a linked teen account, a parent or guardian may manage whether conversations can be used to improve our models through Parental controls.

      > Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.

      [1] https://help.openai.com/en/articles/5722486-how-your-data-is...

      [2] https://web.archive.org/web/20260910151242/https://help.open...

      • the13 19 minutes ago
        "may" = will, unless they screw up
      • ACCount37 33 minutes ago
        I mean, how else would those buttons work? It's explicitly feedback data. And "this is good" or "this is bad" is empty if divorced from what "this" actually is.
        • AlotOfReading 29 minutes ago
          If the buttons are incompatible with the absence of the feature, I'd expect the buttons not to exist when the feature is disabled. Anything else seems like a straight up footgun. I guess it'd also be acceptable to pop up a scary warning box asking "are you sure?"
          • palmotea 23 minutes ago
            > If the buttons are incompatible with the absence of the feature, I'd expect the buttons not to exist when the feature is disabled. Anything else seems like a straight up footgun.

            It's called a "dark pattern." They want you to shoot yourself in the foot, so they'll do their best to aim your gun at your foot and put your finger on the trigger. And then when you do, because you don't have perfect understanding or execution, they'll say "your fault!"

        • ummonk 27 minutes ago
          It could go into personalization / memory. Or they could be A/B testing some system prompt tuning and consider the thumbs up / thumbs down as statistical feedback on the particular flags that are enabled for your account.
    • quentindanjou 2 hours ago
      We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing.

      Or at least: tell me I should be careful/worry about those particular things.

      • the13 18 minutes ago
        No, people need to take responsibility for their actions. We don't need more over regulation.

        Verify, don't trust.

        You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.

        • scuppernong 3 minutes ago
          i'm sure you consult your attorney every time you agree to terms and conditions
        • quentindanjou 6 minutes ago
          So I should verify that my data isn't just shared for product improvement but also to take credit from me?

          I should verify with wireshark and other software that my LG TV isn't listening to me and selling my data.

          I should make sure that whatever product I buy I spend the time to go over every setting page in case there is a switch (defaulted on) that says "I authorize the sell of my data".

          I should make sure to look at every ingredients on the back of each box of food product to make sure it will not kill me.

          I should document myself on the undisclosed growing practices (because no packaging here) of the vegetables and fruit I am buying and make sure that I equal PhD researchers on the dangers of the pesticides used by the specific company I am buying from.

          I should make sure myself that the battery in any device is up to standard and will not blow me and my living place by researching the factory that made it and buying testing equipment.

          I should make sure to educate myself on how my retirement 401k investment strategy works otherwise, I may not have proper retirement.

          ... I could go on and on; it's infinite.

      • cyanydeez 2 hours ago
        The grift economy requires all marks to be responsible for the fraud perpetrated by others.
    • beering 1 hour ago
      Literally every famous open math problem has had >1 mathematicians ask ChatGPT to solve it. Probably greater than >1000 if you count randos. There is no math problem that OpenAI/Anthropic can solve that didn’t have users already try it in Chat/Claude.
      • mettamage 59 minutes ago
        I'm the random that says "solve Riemann make no mistakes" With Fable 5.*

        It's fun! I sometimes have tokens to burn and it's instructive despite knowing nothing about the problem other than a NumberPhile video

    • speak_plainly 2 hours ago
      Coincidentally, a tweet from OpenAI's Tibo yesterday:

      https://x.com/thsottiaux/status/2097746417012166816

    • DrewADesign 1 hour ago
      Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that:

      1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.

      And also:

      2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.

      If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.

      • ianjbutler 1 hour ago
        > To me, dark patterns (like manipulative wording) imply that:

        Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.

        • dylan604 1 hour ago
          That's my take as well. At some point, they will claim that you cannot use their service without contributing back. If you quibble with them using your info in exchange for using their service, you don't get to use the service. Hence, I don't use their service. I do not trust these companies at all.

          At this point, I'm left wondering what is wrong with me that I don't just go with the flow, otherwise, what's wrong with everyone that does.

        • phoghed 59 minutes ago
          Just found out Google didn’t index a googol pages. Lying to me about everything this whole damn time smh
      • michaelmrose 1 hour ago
        [dead]
    • enraged_camel 1 hour ago
      There's also the fact that the setting has been getting turned on by some users: https://news.ycombinator.com/item?id=49643556
  • qg127 2 hours ago
    There are so many naive academics. They still believe an "opt-out" button.

    Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.

    Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.

  • buellerbueller 37 minutes ago
    Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.

    You will not be able to opt out unless you completely isolate yourself from society, tough shit.

  • xbar 1 hour ago
    How can OpenAI figure out how to be trustworthy?
  • techblueberry 2 hours ago
    But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
    • Robotbeat 2 hours ago
      Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
      • techblueberry 2 hours ago
        I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.

        Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.

        • HarHarVeryFunny 1 hour ago
          It seems that in this case OpenAI are suggesting that the researchers whose work they scooped were using OpenAI models with an account setting that allowed OpenAI to train on anonymized prompts.

          It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.

    • faangguyindia 48 minutes ago
      More like, "Who are you going to believe: a multi billionaire, or people competing for a million dollar math prize?"
    • bigstrat2003 2 hours ago
      If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.
      • faangguyindia 47 minutes ago
        Sam Altman has provided people with more generous usage than Claude or Gemini.

        There is no doubt ChatGPT is the most generous LLM provider!

    • cindyllm 2 hours ago
      [dead]
  • mrbluecoat 2 hours ago
    "Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

    Welcome to the party, with the rest of humanity.

  • wslh 1 hour ago
    Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.
  • ChrisArchitect 1 hour ago
  • hn1rig3rak 2 hours ago
    the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
    • spindump8930 2 hours ago
      The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
  • esafak 1 hour ago
    What happens if you use a different harness?? Does opting out online suffice?
  • mannanj 19 minutes ago
    And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.

    Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.

  • seobot_dk1289 2 hours ago
    [flagged]
  • josefritzishere 2 hours ago
    [dead]
  • shevy-java 2 hours ago
    Considering how Apple today announced that its products will contain a spying-on-conversation anti-feature by default, it seems reasonable to assume that all this spying is primarily done to train their AI model; and secondarily also to gather information about The People.

    AI is becoming more evil by the day.

  • mainecoder 1 hour ago
    Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
  • square_usual 2 hours ago
    I think this is stupid, for three reasons:

    1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

    2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

    3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

    • solenoid0937 1 hour ago
      > in this case too they didn't actually have the solution

      Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.

      I would almost expect training to overweight conversations with novel scientific and mathematical implications.

      > the only reason they can't definitively say no is that for privacy reasons

      They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.

      Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

      > that opted-out user data was used for training

      Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.

      • letmevoteplease 1 hour ago
        You quoted the OP saying "in this case too they didn't actually have the solution" and responded with the totally unrelated, "Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data."

        Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.

        >They could 100% definitely say no, if they know they did not train on user data.

        No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.

        >They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

        They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.

    • faangguyindia 26 minutes ago
      If the mathematicians are using ChatGPT, then they themselves are benefiting from the work of other ChatGPT users, so ChatGPT using their work is not wrong!