61 comments

  • teekert 20 days ago
    Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization. This even caused some of our (testing) prompts to be used for training (which Mistral removed after we expressed our disappointment).

    For some time these pages conflicted with what our users reported (they said that in contrast to what I stated to our management they found they were opted into training on prompts by default as per their own privacy page). Mistral just now corrected their docs. I'm not sure how long the conflicting situation has lasted, but at least for several days.

    For contrast: Claude disables training on prompts for organizations starting from the 18 euro tier [0]. As a European I'm disappointed.

    [0] https://claude.com/pricing#team-&-enterprise

    • throwaway89201 20 days ago
      > and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization

      There is a toggle on https://admin.mistral.ai that allows you to disable training for both Vibe and Console/API for your entire organisation. And I'm not on the enterprise plan. I've disabled training the first time I created an account, and it has remained that way.

      You story is also very confusing due to the wording around "opt-in by default" and "disappointed [about] opt-in to training" (most people would be disappointed about an opt-out) and probably conveys the wrong message to most people.

      • teekert 20 days ago
        I don't have that toggle (but could indeed have sworn I saw it earlier).

        Sorry, I have always thought that "opting in" is, "opting for the presented option" and opting out is "opting out of it", so opting out [of sharing prompts for training] is choosing to not share, but apparently I was wrong my whole life. I'm not a native speaker, and I think most people here (in my country) would interpret this the way I do? Weird but TIL.

        • bee_rider 20 days ago
          FWIW I don’t find what you wrote confusing, but I do think it is following a sort of… bad convention that some companies have been pushing.

          Opt-in and opt-out describe the nature of the choice that you make. “Opt” means to choose (apparently it is a French word we stole). Opt-in means you have to proactively choose to be in. Opt-out means you have to proactively choose to be out. “Opt-in by default” is an overly verbose way of saying “opt-out.”

          In either case it describes the choice that you need to proactively make to override the default behavior.

          Edit: I should also say that it is a “known point of contention” where pro-privacy people have been pushing back on this phrasing. So, you have probably accidentally stumbled into an ongoing discussion, which is why some of the comments might be unexpectedly prickly.

          • rcxdude 20 days ago
            >Edit: I should also say that it is a “known point of contention” where pro-privacy people have been pushing back on this phrasing. So, you have probably accidentally stumbled into an ongoing discussion, which is why some of the comments might be unexpectedly prickly.

            Yes, because people will say that something should be opt-in, meaning off by default, and then someone makes it on by default and says 'oh, it's opt-in, you're just opted in by default!'

            • bee_rider 20 days ago
              I 100% agree that “opt-in by default” is bad and that “opted” in that sense is a bullshit incoherent phrase. Unfortunately this “opting as a thing that happens to you instead of a choice” idea has evidently been promoted well, so we have to expect that people will unknowingly use it.
              • therealpygon 19 days ago
                “Opt-in” means to agree to something. If something is “opt-in by default”, it means that it is opt-out and someone is using weasel wording and contract inducements to force you into something you didn’t actually agree to. Any “disagreement” is just people trying to advocate for what they wished it meant.
              • mahirsaid 20 days ago
                And suddenly we have laws directing AI companies to report bad behavior or potential rival candidates.
                • half_fish 19 days ago
                  And some shadowy departments do this without needing any laws at all.
        • Miraltar 20 days ago
          If you opt in it means that by default you're not in and you chose that option. So if you're now included in training by default, then it's an opt-out feature as in you can opt out of it.
          • teekert 20 days ago
            Idk, feels like my brain does not have the model for this, this does not exist in my language afaik haha.
            • bluefirebrand 20 days ago
              Opt-in = "we assume you are out unless you tell us you want to be in"

              Opt-out = "we assume you are in unless you tell us you want to be out"

            • sn0n 19 days ago
              Just remember, good company’s leave them off by default. Bad companies will turn them on for you.
            • blub 20 days ago
              Replace opt with its equivalent “choose” and it becomes choose-(tobe-)in and choose-(tobe-)out.
        • throwaway89201 20 days ago
          > I don't have that toggle (but could indeed have sworn I saw it earlier).

          You don't see "Allow the use of your interactions with Vibe to train Mistral’s AI models" at https://admin.mistral.ai/vibe/privacy ?

          And "Allow the use of your API calls to train Mistral’s AI models" at https://admin.mistral.ai/plateforme/privacy ?

          • teekert 20 days ago
            The toggles I see

            at https://admin.mistral.ai/vibe/privacy:

            Allow public sharing of chats content

            Allow user feedback on model responses

            Chat Retention Policy

            At https://admin.mistral.ai/plateforme/privacy I see

            Allow the use of your API calls to train Mistral’s AI models.

            Enable Labs models

            So indeed, no "Allow the use of your interactions with Vibe to train Mistral’s AI models". I only see that setting here: https://chat.mistral.ai/chat?profile_dialog=privacy And I have to ask every user in my org to go and turn it off.

          • teekert 19 days ago
            Mistral support just confirmed that the org-wide toggle was removed from the Team plan earlier this year and made exclusive to the enterprise tier. I bet that if you signed up earlier, they didn't remove it from your account, but it's not there for new customers. Which explains some of the confusion here.

            Also explains that the community support in Discord insisted that the toggle should be there, moreover they told me the individual toggles on the user's pages wouldn't do anything because it defaulted to "off" for any organization, as per their docs until 2 days ago. But that changed.

          • teekert 19 days ago
            Ok it (org wide toggle) just (24h after posting this) returned and the docs were changed again! They now include the text:

            "Vibe (Teams): Administrators can disable data training usage for the entire organization."

            That’s nice! Except that it was on for my entire org so I’ll be checking if they didn’t store anything first thing tomorrow. I wonder what changed their minds.

        • duskdozer 19 days ago
          No you're not wrong, but while I can't speak for this specific instance, in many other cases, this language is intentionally confusing in order to dark-pattern users into agreeing to the thing they don't want to. Similar is using ambiguous language next to a switch/checkbox that makes it difficult to tell exactly what each state means.
        • quadrifoliate 20 days ago
          I think there is a simple, unambiguous way to describe this, which is to use the passive voice or an explicit subject with the past tense to clarify that you didn't do the opting in.

          Saying "we were disappointed to be opted in to training by default" clarifies the point you're trying to make, which is that the toggle (whatever it may be called) was set to training by someone else, not you.

          Actually in your case you'd say "...after being disappointed that the Pro tier opted us in to training on prompts by default", which gives an explicit subject ("the [Mistral] Pro tier"). No one will confuse that with the opposite "the Pro tier opted us out of training on prompts by default".

          • frereubu 19 days ago
            I think the trouble with "opted in to training by default" is that there was no "opting", i.e. no choosing. I can't come up with a better suggestion off the top of my head though!
        • abhayaya 19 days ago
          Imho, legally, one should always be 'opted out' of 'forced mingling of your ideas' without deliberately opting in. Same as for advertising, click-through tracking, and sociopsycho/marketing metrics. Like the old firewall rule (easier to handle opt-in/out that way than firewalls; like, how many people can run a firewall that way now?). Users should have it right in front of their face when they install something or register, too, not hidden among options. It might take the user five or ten more minutes to get to use the app but worth it.
        • tedggh 20 days ago
          You said it correctly, there’s no confusion.
          • tedggh 19 days ago
            Downvoted for knowing how to read
      • Barbing 20 days ago
        My native-US-English speaker read:

        “disappointed that the Pro tier was opted-in to training on prompts by default” [and required manually opting out]

        “the Team tier was now also opted-in by default” [and required manually opting out]

        In context of each sentence and the larger comment, read smoothly here.

        Also - have seen more than one lively discussion on these phrases, since defaults can stick 95% of the time and Big Tech has done their best to be abusive about what they automatically enable for users by default for some time.

        • rcxdude 20 days ago
          It's understandable but it is technically inconsistent and dilutes the meaning of opt-in. Opt implies an active choice, so you never are opting for the default.
          • wtallis 20 days ago
            It's also simply bad writing regardless of what meaning the reader takes from it. "Enabled by default" or "Allowed by default" are much clearer and more natural phrasings without any connotations of the user having taken an action to express a choice.
      • qwertox 19 days ago
        I also have these toggles, while being on a free plan. I must have set them to disabled at setup because that's how they are now.
      • surcap526 20 days ago
        [dead]
    • lukan 20 days ago
      "For contrast: Claude disables training on prompts for organizations starting from the 18 euro tier "

      In theory also for individuals?

      At least I have that toggle to deactivate that. But how would I ever know if they actually respect that?

      • kccqzy 20 days ago
        If you don’t trust the company to keep their promise, don’t use their products.
        • RussianCow 20 days ago
          I don't trust any company with their word on anything. Luckily, privacy policies are legally binding.
          • lostlogin 20 days ago
            > Luckily, privacy policies are legally binding.

            Companies violate them all the time and massive leaks happen a lot.

            The punishments are trivial.

            • rkangel 19 days ago
              GPDR fines in the EU are NOT trivial.
          • tjwebbnorfolk 20 days ago
            > Luckily, privacy policies are legally binding.

            Laws are violated all the time. The graveyard is full of people who had the right-of-way at a crosswalk ...

          • nullsanity 20 days ago
            Why does that matter? Legally binding just means "slightly more expensive when we get caught"
          • DarmokTanagra 20 days ago
            [dead]
        • LeBit 20 days ago
          How can you trust any company?

          The only company you could think of trusting is one where an external , independent auditor is doing its work.

        • locknitpicker 20 days ago
          > If you don’t trust the company to keep their promise, don’t use their products.

          Your comment is perplexing. No company on earth meets your requirement. What are you expected to do? Move to a hut in the woods?

          • dylan604 20 days ago
            > Move to a hut in the woods?

            I'm so burned out on the bullshit companies that succeed in tech forcing their will upon us serfs that I'm actively looking into that very thing at this time. Tech used to be fun. Now it's just depressing. I'm ready to go back and take the blue pill.

            • lukan 20 days ago
              Been there. Nice and peaceful, but can get lonely, though.
        • roosterIllusi0n 20 days ago
          That is ridiculous. Vote to ensure products have to be legally private and make it a crime to share or reuse your data for anything. This should be the norm that people vote for. The idea that we have to give up privacy so some nerdy pervert can be a billionaire off the backs of people who do the work is absurd.

          Criminalize failure to keep private data private. Arrest CEOs and executives. Put them in jail when it happens.

        • lovich 20 days ago
          What company do you trust to keep their promise?
      • danelski 20 days ago
        My 20 EUR Claude (individual) had the training on by default.
      • teekert 20 days ago
        This post is only about the defaults, and expectations therefore for Team and Enterprise plans and their organizational controls.
      • warkdarrior 20 days ago
        How do you know an LLM provider does not kill puppies every time you send a prompt longer than 14 words?
        • lukan 20 days ago
          Because they have no incentive as a company to do that? But do have a strong incentive to learn from user input. (Most cutting edge features are developed private - very valuable to get that into your tool)

          And getting the data is not hard, they already have it. Risky is indeed a bit making use of that data, as that requires at least some humans (as potential whistleblowers). But you don't even have to tell them, where the data came from.

          Whether they do it? No idea, I assume not, but I see a risk.

          • roosterIllusi0n 20 days ago
            We can easily make it a crime to collect the data and another crime to share it or store it in a way that allows someone else to copy it, authorized or not.

            People always want to overcomplicate everything to benefit 5 super rich people that got rich taking wages from workers and ripping off retail investors. Please stop the pandering. Vote for yourself and your neighbors, don't vote to eliminate privacy so a rich pervert gets to exploit you.

            • lukan 20 days ago
              And what would my vote change about the situation?

              Local hosting will be a solution, once the hardware becomes affordable.

              • roosterIllusi0n 18 days ago
                Vote for people who want to restore the 1950s federal tax rates that we already had in the 1950s. I find so sad that people ignore history and act like solutions don't already exist.

                The 90% tax bracket in the 1950s capped CEO and investor yearly income to the 2026 equivalent of 5mil a year. Use that imaginative brain that likes to speculate instead of learning history to image how our entire society changes if no individual can earn more than 5mil a year in all income sources. We also banned stock buybacks and had more regulations put in place after the 1929 collapse and great depression. Those things were removed slowly over the 60s and 70s until a lot was gutted all at once in the 80s. The final blows came in the 00s.

                All the problems that existed leading up to and during the great depression are back because we reverted all the laws and policies that were put in place to prevent it from happening again.

                None of this is complicated. Stop worrying about culture nonsense or right wing nonsense. Start voting to tax rich like we did in the 1950s.

                Why did we have so many local and regional stores back then, but only a handful of conglomerates today? There is no incentive to merge companies by CEOs when it turns two 5mil a year CEO jobs into one 5mil a year CEO job. In that environment, merging is due to actual company need, not CEO enrichment.

                CEOs will also stop cutting wages, under staffing, and outsourcing. It all stops because they no longer get paid more doing it. This is actual US history, the country boomed economically because of those tax rates and financial regs. This is not something anyone gets to deny, it already happened.

                Voting against a proven solution is madness.

      • tedggh 20 days ago
        Always assume and act as they won’t honor it.
    • summarity 20 days ago
      “Opt in by default” would mean that it is not enabled by default. Do you mean opt out?
      • kevincox 19 days ago
        Yes, I found this comment very hard to understand until I realised they were talking about it being opt-out with no setting to change it. (At first I thought it was opt-in by default, and the "by default" implies that there is a setting to change the opt-in/opt-out setting)
      • teekert 20 days ago
        You are "opting in to sharing your prompts for training", by default in this case. My slider says: "Allow the use of your interactions with Vibe to train Mistral's AI models.", it is on by default for everyone on the Team plan, the admin can't centrally turn it off anymore, and any user can toggle it when they want to. This all changed last week.

        I know I'm naive but I expect that when I pay, this stuff is simply off, so I was already surprised by the Pro plan. But I did look out for it there, because Anthropic made this switch some time ago.

        • KPGv2 20 days ago
          > You are "opting in to sharing your prompts for training", by default in this case.

          The English term for that is "opt out" not "opt in." To opt is to choose. If something is on by default, you have not opted in. You were forced in, and turning it off means you must opt out. (I.e., choose to be out.)

          normally I wouldn't care about a mistake like this, except that opt in/out are very important concepts in software development and hacker culture. And it reversed the meaning of the original comment in a highly confusing, relevant way.

          • jrave 20 days ago
            as another non-native speaker, i think that while the person you're answering to didn't use the default way of expressing this in the english-speaking world, i did understand what they meant, i think they do know what opting means and i think the reasoning is this: when they say data collection is "opt in" by default they mean that by choosing to use a product, you are opting into your data being collected (at the same time). on the other hand, native speakers saying something is "opt out" describes the options or toggles one has available in the default case - when something is toggled "true", you (only) have the choice to toggle it of. so english speakers talk about the controls one has to CHANGE the status quo.
            • kzrdude 20 days ago
              As another non-native speaker, I misunderstood what the person meant.
          • brendoelfrendo 20 days ago
            They said "opt in to training on prompts by default," which is coherent English and perfectly fine. The phrases "opt in" and "opt out" will always be contextual based on what is being opted, so it really falls to the reader to pay attention to that context. Please don't give English language advice as though you are an authority; there is not a rule of the English language that would make their usage unacceptable.
            • F3nd0 20 days ago
              Since the whole idea of ‘opt in’ is choosing to do something, the only reasonable interpretation of ‘opt in by default’ is that the default state requires you to make the choice yourself. However I try to bend my brain, sing the term ‘opt in’ for the exact opposite makes no sense whatsoever.

              Note that I’m not using English grammar as my argument; human language (and English especially) has plenty of exceptions and expressions which make little sense, and if people were to adopt this phrasing en masse, it would naturally become a part of the language, whether it makes sense or not. My argument is that adopting this phrasing is a terrible, user-hostile decision and in my honest opinion we really, really, really ought to avoid it.

              • PotatoPrime 20 days ago
                I read it as though he was trying to say '[you are] opt[ed] in by default'.

                Now granted, he did not say that, and I only reached that conclusion by the rest of this conversation, but I think to have meant that and said '... was now also opt-in by default...' isn't unreasonable.

                Just sharing since you mentioned you couldn't see how using that term for the opposite would make sense.

                • F3nd0 19 days ago
                  Thank you. However, assuming you mean ‘an affirmative choice is made for you’ when you say ’you are opted in’, that use really takes the ‘opt’ out of the ’opt-in’. You can opt in, but you can’t have someone else ‘opt you in’ for you; that seems to go against the very core meaning of ‘opt-in’ (i.e. having to explicitly choose something yourself).

                  In other words, you can explicitly choose something by yourself, but if other people (in this case companies, service providers) choose something for you, it’s no longer your explicit decision made by yourself. Hence why I believe we can’t consider that ‘opting in’.

                • Timon3 19 days ago
                  Even read this way, it doesn't make logical sense. It's a great example of newspeak - corporations know people use "opt-in" as a purchasing signal, so they're trying to redefine the word by abusing it enough times.

                  "Opting someone else in" is essentially the same category of error as "consenting for someone else". People can only consent for others in specific circumstances and even then only for specific people (e.g. legal guardians), trying to argue "but I consented for them!" in front of a judge usually ain't gonna work.

              • KPGv2 19 days ago
                Here's the dictionary: https://en.wiktionary.org/wiki/opt-in

                > The property of having to choose *explicitly* to join or permit something; a decision having the *default option being exclusion or avoidance*; used particularly with regard to mailing lists and advertising.

                You're just using the term wrong. That's okay, I've used words wrong before, and it just means you've been presented an opportunity for learning a new fact.

                But there's no argument to be had except with the dictionary.

                • F3nd0 19 days ago
                  My use seems to be perfectly in line with the given dictionary definition, though? Did you perhaps mean to address a different comment?
            • mikebenfield 20 days ago
              I find the usage bizarre and confusing, as others also clearly do. There’s nothing wrong with pointing this out.
            • TulliusCicero 20 days ago
              > They said "opt in to training on prompts by default," which is coherent English and perfectly fine.

              No, it's really not. It should read "opted in to training on prompts by default".

              "Opt in" means "you have to affirmatively turn this thing on". "Opted in" means "the thing is currently enabled".

              Which is why it's probably better to write it as "enabled by default". Much harder to misinterpret.

            • KPGv2 19 days ago
              > They said "opt in to training on prompts by default," which is coherent English and perfectly fine

              "Opt in" and "by default" are contradictory phrases. They mean literally the opposite of each other. The only reason it's coherent is that most native speakers will elect to believe "by default" was used correctly because it's the easier to use term.

              > there is not a rule of the English language that would make their usage unacceptable

              I think you're using emotionally inflammatory language, possibly unintentionally. no one is refusing to accept what they wrote.

              That being said, the way they used "opt in" is categorically wrong. It's the opposite of the dictionary definition.

              https://en.wiktionary.org/wiki/opt-in

              > The property of having to choose explicitly to join or permit something; a decision having the default option being exclusion or avoidance; used particularly with regard to mailing lists and advertising.

        • nozzlegear 20 days ago
          > You are "opting in to sharing your prompts for training", by default in this case.

          I understand what you're saying here, but maybe "turned on by default" is less confusing for everyone.

    • Exoristos 20 days ago
      > opt-in by default

      That's called opt-out.

      • globular-toast 19 days ago
        Yeah, it was confusing to read the parent before I realised they got the terms the wrong way around.

        To those wondering, "opt" means to choose. "Opt in by default" makes no sense because you didn't choose; this is just "in by default". If they give you an option then it's called opt out. Opt in would be "out by default" with the option to go in.

    • jacquesm 20 days ago
      I don't trust any of these companies with my data, and I assume that whatever data they've got is going to be used, one way or another, no matter what they tell you. It would be nice if you could stick a sentinel in your data that if it ever shows up in the models you know they've broken the rules for sure.
      • bluefirebrand 20 days ago
        Yeah. Unfortunately they know you can't catch them on this so they feel absolutely free to do whatever they please

        If such a data sentinel did exist then we might see them change their behavior

    • rezonant 19 days ago
      Buried by the distraction of the semantics of "opt in" vs "opt out", the more interesting conflict isn't addressed in the sibling threads.

      > and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization

      vs

      > Vibe (Teams): Administrators can disable data training usage for the entire organization.

      Is the document out of date? Or did Mistral reinstate this ability after the fact? What's the story here? Others seem to be stating they have and have had the ability to opt out of training centrally for a long time.

      EDIT: Oh, it is addressed just a bit hard to find with all the opt in/out explanations: https://news.ycombinator.com/item?id=49549102

      • teekert 19 days ago
        You are right! They must have just now switched this back, I also have the toggle now! It was turned on sadly but it’s something.
    • teekert 19 days ago
      Just now the sentence:

      “Vibe (Teams): Administrators can disable data training usage for the entire organization.”

      Was added to tfa. And I now see an org wide toggle where there was none before! Sadly it was on so I hope no users submitted stuff in the mean time, but it’s something!

    • Aldipower 20 days ago
      Today I registered a free account with Grok, because I simply was curious. Man, training is "off" by default even with the free tier. I was positively surprised. As a European I'm disappointed too.
      • blazarquasar 20 days ago
        Grok has possibly the worst ToS of any of the AI providers. They are probably different in the EU, but:

        > In choosing to submit, create, generate, record, post, or display Inputs on or through the Service, you grant an irrevocable, perpetual, transferable, sublicensable, royalty-free, and worldwide right to SpaceXAI to use, copy, store, modify, process, adapt, transmit, distribute, reproduce, publish, upload, download, display in public forums, list information regarding, make derivative works of, and distribute such Content, including anything referenced therein, in any and all media or distribution methods now known or later developed, for any purpose, and to aggregate your User Content and derivative works thereof for any purpose, including but not limited to: (i) maintain and provide the Service; (ii) improve our products and the Service and for our other business purposes, such as data analysis, customer and market research, developing new products or features, or identifying or displaying usage or User Content trends; and (iii) perform such other actions to enforce these Terms, comply with our Privacy Policy, comply with applicable law or governmental, court, and law enforcement requests or requirements or keep our Service safe.

        > To the extent the User Content includes a person’s image, likeness, voice, or other similar attributes, you grant SpaceXAI the same rights to use those attributes as part of the User Content as described above. You represent and warrant that you have obtained all rights, licenses, notices, permissions, and consents necessary for SpaceXAI to use that User Content.

        https://x.ai/legal/terms-of-service

      • _puk 20 days ago
        Lol, I'm fully expecting in 6 months time:

        "A bug in our portal had the setting for training inverted. This means that when you expected us not to be training on your data, we actually were. We know this adversely impacts the trust our users invested in us, so as of today we are crediting all affected accounts with $200 to use on our latest models".

        • Aldipower 20 days ago
          One can expect a lot what happens in 6 months. Maybe the earth will turn clockwise then! Think about it. :-D
          • artwr 20 days ago
            Lol. I would have thought it depended more on your relative position to the equator than on the season.
        • leonidasrup 19 days ago
          Until the risk of large financial penalties or long jail times is high enough, the CEOs of there companies operate using the "It's Better to Ask for Forgiveness Than Permission" model.
    • semiquaver 20 days ago

        > opt-in by default
      
      Sorry to nitpick but the scheme you’re referring to is called “opt-out.”
    • kragen 20 days ago
      > being disappointed that the Pro tier was opt-in to training on prompts by default

      "Opt-in" means that the default is non-participation, for example, not training on your prompts. Is it possible that you intended to say "opt-out", which means that the default is participation? That's what the context seems to suggest.

      See, for example, https://termly.io/resources/articles/opt-in-vs-opt-out/:

      > Data privacy laws like the GDPR and CCPA give individuals the right to opt in or out of different data processing activities.

      > · Opt in consent means the user takes an action to show they agree to something,

      > · Opt out consent is when they take an action to say no.

      Or https://bigid.com/blog/opt-in-vs-opt-out-consent/:

      > • Opt-in consent requires users to actively agree before data collection or processing.

      > • Opt-out consent allows data collection by default unless the user declines.

      This is an important distinction, because confusing the two (as you seem to be doing) can lead you both into unethical fraud and legal liability.

    • surcap526 20 days ago
      [dead]
  • 20k 20 days ago
    You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy

    The idea that they'll steal from everyone except you is just wishful thinking

    • teeray 20 days ago
      This is why these companies paying subscriptions shoveling everything into Claude thinking "oh, they're not training on our stuff" is hilarious to me. Of course they are. They probably are extra super-duper sure not to have Claude admit to that in any way, but there's no way they're giving up training on the sum total of both open and closed source code out there.
      • vorticalbox 19 days ago
        Good example of this is figma. Claude was updated to work with it and now we have Claude design.
    • protocolture 20 days ago
      >You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy

      Have been having a think about this statement. I agree in intent, some of them are probably breaking the agreement for training data. I dont think Microsoft is doing it, Enterprise Data Protection is the plank holding up their entire Copilot line. One whiff and everyone's gone. Copilot isnt actually good at anything except giving some illusion of protection, and preventing users from following a desire path to other LLMs without enterprise data protection. Its the core value proposition. But we only need to wait and see what the next 20 data breaches tell us to find out for sure.

      In detail however, I dont know if they could just claim it as fair use, after exclaiming that they specifically wont do that. I dont think "Fair Use" would be the issue so much as contract law. I know a EULA wouldnt hold up the other way (By reading this you agree not to steal my data and train an LLM with it) for data thats made freely available on the internet. But if you are purchasing the "No Training" contract they would be in breach if they trained with it. Possibly fair use would let them keep the data after paying whatever they owe in terms of contract breach.

    • gitgud 20 days ago
      If you have an enterprise contract with them, the legal protections for the consumer is much higher… is what I’ve been told anyway
      • 20k 20 days ago
        It isn't, in both cases you're protected by exactly the same legal system
        • zamadatix 19 days ago
          It's also the same planet but finding a commonality somewhere in the chain is not the same thing as finding a lack of difference in the matter.
        • gitgud 19 days ago
          Same legal system, but an OpenAI enterprise contract is not the same as a ChatGPT pro subscription... The terms of service agreement would be quite different
      • olejorgenb 20 days ago
        Not sure what you are saying exactly, but

        > the legal protections for the consumer is much higher

        You know you give away the right to file class action law suits against Anthropic when you accept their Terms Of Use, right? (at least the Americans ones)

        • stingraycharles 19 days ago
          Is that the case with enterprise contracts as well? Can’t imagine any decent procurement / legal team accepting this.
      • autoexec 20 days ago
        If you have an enterprise contract with them you're still unlikely to ever know that you've been lied to unless some whistleblower at the AI company comes forward and even if you do somehow find out, it's too late. Once they have your data and have trained on it there's no taking it back. At most they'll pay out some tiny settlement that's a fraction of how much money they make in a week and they'll continue to profit from your data forever.
    • rcxdude 19 days ago
      This only makes sense if your model is 'the labs are just breaking the rules all the time' as opposed to 'the labs have a legal theory of defense for the one aspect of the business which is legally questionable'. The question of whether doing something their terms of service explicitly say they will not do opens them up to civil liability is a much more certain one than anything they are doing around copyright. If they are doing this and anyone can prove it, they are going to be sued very hard, and it will be a pretty straightforward case.

      (To put it another way: if the copyright arguments against training on data scraped from the internet fail, the big AI companies don't really have a business, so they are going to proceed on the basis that they do until someone forces the issue otherwise. They don't need to train on data from their customers, and that is an argument that is almost certain to fail in court. If you're carrying 100kg of cocaine in your car, speeding is a really dumb idea, but the analogous action here would be firing a machine gun into the air from the driver's seat)

      • AlexandrB 19 days ago
        It doesn't matter. Even if they don't train on your inputs now they will "boil the frog" and do so in the future. What the user wants is irrelevant in the tech industry.

        > If they are doing this and anyone can prove it, they are going to be sued very hard, and it will be a pretty straightforward case.

        This is a joke, right? The usual settlement in these cases amounts to a few days worth of revenue.

      • knollimar 19 days ago
        Well maybe they have another suspicipus theory for defense of your data. Like taking it, laundering it through a summarizer, and training on that. Surely that's less questionable than taking copyrighted stuff and saying "don't output over 15 words, never verbatim" pretty please.
        • rcxdude 19 days ago
          Copyright has a lot of leeway for arguments about fair use. 'We won't use your data for training' in a contract has far fewer.
    • WarmWash 20 days ago
      It would be catastrohic for any of the big labs if it came out that they were training on what was sold as private.

      I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.

      • thephyber 20 days ago
        No it wouldn't.

        It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts.

        When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes.

        Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file.

        And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.

        • avianlyric 20 days ago
          The credit reporting agencies lost the data of normal people, they didn’t lose their customers proprietary internal data. The credit agencies didn’t loose or misplace their customers data, so obviously their customers don’t really care that much, and the credit agencies weren’t sued into oblivion.

          But I can guarantee you that if a companies internal data got leaked or misused, then every single enterprise customer of that lab would turn around and start suing them. As an enterprise customer you would be foolish not to, if only if figure out via discovery just how badly you got screwed.

          You want to see how nasty that can get. Just go and look at what Apple is doing to OpenAI at the moment. Do you really think Apple wouldn’t find a way to sue a lab into oblivion if they discovered a lab had secretly started training on their data?

        • jppittma 20 days ago
          that was accidental? this would be straight up fraud.
          • chillfox 20 days ago
            You can structure it so that it becomes accidental.

              1. Ensure security barriers are weak or honor based.
              2. Put individual researchers under a lot of pressure.
              3. If you get caught, blame the weak barriers, or the individual researcher.
            
            Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room with a sign on the door saying please don't enter.
            • thephyber 19 days ago
              Yeah, the steps follow exactly what happened at VW with DieselGate. The diesel emissions lies were found out because some enterprising person set up an emissions testing system and drove the car in real world scenarios with it to verify the claimed emissions.

              There's no reliable way to verify a foundation model has been trained on a particular piece of proprietary data. If an API key is ingested, hopefully the foundation model is wrapped in enough moderation that the raw API key oberserved during training is not recited verbatim in the output.

            • jppittma 14 days ago
              Not trying to be rude, but do you work in tech? I can't imagine presenting this as a plan of record in a design review. And the world runs on good faith. If you call a pharmacy, claim to be some doctor, leave a voice mail, and give their (public) NPI number, there is no validation.

              Why aren't people calling in prescriptions for themselves? I guess it just kinda runs on trust me bro and the threat of being put in prison.

      • r_lee 20 days ago
        exactly, I don't understand how HN doesn't understand this

        by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"

        • 20k 20 days ago
          I mean, they've already violated the law in acquiring all their training data already, why would they be uncomfortable violating a contract to get more training data?
          • vineyardmike 20 days ago
            Because violating the law was the prerequisite to starting their business, without it they're worth $0 and have no models. They've already survived Training on customer input may help the models but it isn't "bet the farm" helpful. Now, they have a thriving business, so they shouldn't risk their business for incremental data that they can buy.

            At this point, the reputation of the business matters too. Fable explicitly didn't support zero-retention usage, and it saw significantly lower adoption vs other flagship models, and their past releases. Being caught abusing enterprise contracts is really hard to dig out of.

            • r_lee 19 days ago
              and they're under no contract with those millions of websites and books, but it's a whole other thing when it's a paying customer, especially a large enterprise with a signed contract with DPAs etc.
      • realusername 20 days ago
        It's going to be the same as PRISM, people will be outraged and business will go back as usual.

        (And they totally won't do it again they swear, the contract says so)

      • ajb 20 days ago
        It would be catastrophic if they violated confidentiality blatantly, but that doesn't exclude learning of any description. For example, an ordinary human being can't fork a subagent for a particular client and wipe it afterwards. Humans can't stop themselves learning, so confidentiality can't ban all learning.

        Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes.

        So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time.

        Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it?

        Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.

      • tripledry 19 days ago
        I think it's arguable that it wouldn't be catastrophic.

        Doesn't have to be outright lying, it can be something like "opt out" being off by default, and them opting you in at the next update without you noticing.

        This is just a gut feeling and goes into the conspiratorial energy, but I'm pretty sure I can find countless examples of companies behaving like this in the past where it wasn't catastrophic (google meta amazon adobe ...)

      • 20k 20 days ago
        I mean the models have literally trained on:

        1. Child porn

        2. Stolen music

        3. Private github repos, before that was 'stopped'

        4. Illegally pirated books

        Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on

        Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?

        • Henchman21 20 days ago
          These folks need to feel repercussions so hard their souls flee to the afterlife leaving only their sad, dead husks behind.
    • addag 20 days ago
      And even if they are not doing it right now, they probably keep the history available for future training, "just in case".
    • j4k0bfr 19 days ago
      I think innocent-until-proven-guilty is the correct approach in general but... It's also immature to assume your data will stay private in the long term.

      These AI companies are immensely valuable targets. When they get breached, their datasets will inevitably penetrate public datasets. And why would any AI company refuse to train on 'public' data?

      • PunchyHamster 19 days ago
        innocent until proven guilty is definitely wrong approach vs tech giants who time and time again prove they will do anything unethical if only they can theoretize how to get away with it or the fine is low enough
    • solaire_oa 20 days ago
      Pinky-promises, embarrassing. I'd point to tinfoil.sh. I'm not a shill for tinfoil, I haven't even used it or looked past its homepage really, but if we are able to legitimately secure privacy, then training data questions are moot (and the market for training data would probably shift/expose).
      • lukewarm707 19 days ago
        agreed. the pricing is a nightmare. either $20 for codex and up to 00 millions of tokens or $0.50 cache in per million in tee.
    • someothherguyy 20 days ago
    • classified 19 days ago
      I really don't get how there could be anyone to whom this is not beyond obvious.

      And yet, software companies are donating their means of production to these crooks left and right, hastening the day where they really will become obsolete. I'm sure others do equally unwise things.

      If you're big enough, crime pays exceedingly well.

    • deadbabe 19 days ago
      I don’t care about what they do, I only care about what they say, that way, when compliance requires us to use LLMs that don’t train on inputs, I can point to that policy and continue on using the LLM.

      If they are secretly scraping input prompts, that no longer becomes my problem. Someday, a massive lawsuit comes down, and everyone gets to play the victim.

      You’re naive for thinking we don’t know how this really works.

      • flakeoil 19 days ago
        What they say is not the same as what they wrote in the ToS you didn't read.

        If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from?

        • rcxdude 19 days ago
          Have you read their ToS? It's not that hard. This feels like this nihilistic 'law is witchcraft and the big company lawyers always have you screwed in the fineprint' when it generally is pretty comprehensible with not that much work.
        • PunchyHamster 19 days ago
          > If they stole your stuff and used in to train their model, what does it help if there is a lawsuit some years later which you most likely have no benefit from?

          The point is that if someone accuses you of leaking data you can sue them for not abiding by contract, not that the data is not used for training

        • deadbabe 19 days ago
          Because I can get on with my life and do my job. Let the legal departments worry about that.
  • rectang 20 days ago
    I pay for a subscription to Duck.ai mainly because I don't want to be constantly fighting my vendor to protect my privacy.

    Microsoft already did a rug pull on me and opted me in to training months after I signed up with Github Copilot. It exhausting and ultimately futile to monitor these companies.

    It's not guaranteed that Duck.ai will continue to uphold its promise of not training on your sessions — if the company gets bought by Microsoft, it's only a matter of time before the switch to "you can opt out at any time". But since privacy is Duck.ai's brand, it will be somewhat harder for them to hide what they're doing should they betray their customers.

    I also don't actually trust that Duck.ai sub-vendors OpenAI and Anthropic will uphold whatever contract they have with Duck.ai — the whole AI business model is built on lawless consumption of others work.

    We'll ultimately have to run our own models locally, because it's impractical to defend against untrustworthy AI vendors.

    • nozzlegear 20 days ago
      I also pay for duck.ai's service. They're my main "chat bot" thing, and while I had already been paying for their other services when I discovered I also had a duck.ai subscription, I choose to use them primarily because privacy is their whole brand. I only have two minor complainst: I want my conversations to sync between mobile and desktop (while being E2E), and I want a way to organize conversations into "projects" or grouped chats.
      • samename 19 days ago
        Have you tried Lumo by Proton? It is E2EE and has projects. A bit more expensive than duck.ai but I found the model to be good enough
  • saaaaaam 20 days ago
    This is a hugely misleading editorialised title.

    The page title is "Can I opt out of my input or output data being used for training".

    Right at the top of the page it says "In certain cases, your input and output data (such as conversations, documents, and other user-provided content) may be included in Mistral’s model training programs. You retain full control over this processing and have the right to opt out of these programs at any time."

    • pixl97 20 days ago
      "You can opt out at any time"

      ....

      "Of course we'll randomly turn that option off for you aka FB style and hope you don't notice. There is zero legal liability for us doing so, so why wouldn't we".

      • indigo945 19 days ago
        This is a European company, there absolutely is legal liability. The GDPR expressly prohibits using customer data for a purpose other than the one the customer intended and consented to.
    • heaney-555 20 days ago
      This is the case for all the major AI providers. Training collection is on by default, but you can opt out.
      • teekert 20 days ago
        "No model training on your content by default" [0]

        It's not the default on any Team plans and didn't used to be at Mistral (until last week or so). The team plan has a central admin role and page, and "seats".

        And if it was the default, then I'd still expect a big button to turn it off for all seats, and not have to ask all user separately. But this changed over night and that button is not there. Although I expect it used to be, because some people here report that they have it.

        https://claude.com/pricing#team-&-enterprise

    • kieranmaine 20 days ago
      Agreed. I read the title as they would start using my data for training and I couldn't opt-out. After looking at the page and checking my app (I have Pro subscription) it seems like I can opt-out and my initial opt-out when I subscribed was preserved.
      • teekert 20 days ago
        When I started my search for an AI "partner" the Mistral TEAM plan had "use my prompts for training" turned off by default, as per their docs in multiple places. I could have sworn I saw a organization switch in my dashboard for this setting org wide, but am not sure.

        I start the subscription, users report it is "on" by default, I ask support what's up, they say "sorry, docs should have been updated earlier but they are now". And they give me a lot of credits.

        I just want to warn people, the Team sub just changed, docs were update too late, there was very little press about this (in my view) very important change. Actually, it is so important that I would not advice Mistral to our management if they'd use our prompts for training, so I take a TEAM sub so this is disabled for sure, or I can disable this org wide. But I can't anymore, now I have to ask each user/seat to disable sharing, and hope they do, I have no way to check. Way to inspire confidence. And their response: "You can still do this with the Enterprise subscription".

        We flamed Anthropic for their switcheroo with their personal Pro plan a year ago [0], now Mistral does it with their business focused Team plan, so they deserve some fire imo.

        [0] https://news.ycombinator.com/item?id=45076274

  • maz1b 20 days ago
    I could be mistaken, but wasn't Mistral openly championing themselves as a company and EU option that wouldn't do this/didn't do this?
    • whizzter 20 days ago
      Maybe they realized that they were falling behind too much? To me it seems like user-feedback on bad descisions by the AI once it's trained to a basic level is among the most important signals in tuning the model to perform better.
      • danelski 20 days ago
        That's what I believe. Once you have scanned every passive source available, using the user conversations to e.g. find common paths to a solution and shortcut them would seem natural.
        • whizzter 20 days ago
          Paths and shortcuts isn't the most important part here I think, rather negative signals about unfitting choices is more important, do we use algorithm/library/etc XYZ in this situation or not, the developers using it would provide the context suitability of options on a more finegrained level than resulting (semi-)public artifacts provide.
    • mhitza 20 days ago
      You must have not been keeping up with the times :)

      Their big play this year was to write a "whitepaper" on the future state of EU economy, which is something that they'd like to hand of to EU leaders and part of that proposal was some kind of mandatory 10% sovereign AI spend, or some other nonsense like that.

      They are, at least, trying to make big enterprise (with tailored models, custom integration) and government policy plays.

      That just goes to show you how ineffective they are as well at making AI click as a usecase.

      And when they don't get ahead by their own terms they copy what they see ongoing with US AI labs. Le Chat, and Vibe.

      • schubidubiduba 20 days ago
        That just sounds like exactly the same as the US companies are doing with their government (trying to get a lot of money out of them)
    • dwedge 20 days ago
      The same mistral that has a patent for "code implemented tool calls"https://news.ycombinator.com/item?id=49243397#49243588 and said "Companies selling artificial intelligence models in Europe should pay a "levy" to support cultural industries".

      They seem to get a lot of slack just because they're European but every new article I see about them makes my opinion a little worse

  • dang 20 days ago
    Submitted title was "Mistral now trains on user input by default, except on enterprise tier". THat's good information (if true) but best suited to a comment in the thread rather than the title (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...).

    We've changed it to the article title now per https://news.ycombinator.com/newsguidelines.html.

  • thoughtpeddler 20 days ago
    For all the people here alleging that AI companies will train on your data even when you as a user explicitly opt out of training and their terms say they will respect that, etc - do you also believe that within a few years of them having harvested your data, you could perform "knowledge probing" on their models by prompting various questions that determine if they can near-verbatim reproduce your unique data, and then have enough other people do the same that you can then just launch a class-action lawsuit? Because if not... I've got a startup idea for you.
  • segmondy 20 days ago
    There are so models that beats all of Mistral models, plus you can run many of them locally. Why would anyone run Mistral?
    • villish 20 days ago
      To have at least a choice of using a European trained model. Downloadable weights is not open source. Mistral is able to respond to any regulatory queries about training data and concerns.
      • woah 20 days ago
        Haha, the answer is "european regulation"?
        • villish 19 days ago
          Pretty much. Europe is basically not at all involved in the AI race and is reliant on US and China.
    • jnurmine 20 days ago
      To be honest, I have a hard time noticing differences with Claude (in Kiro) and Mistral Vibe, at least with what I use them for. They simply feel like talking to the exact same thing.
    • bossyTeacher 20 days ago
      Sovereignty. Not everything is about performance.
    • icantevenhold 20 days ago
      Mistral OCR is really good and their small models are great for simple translations, I use them regularly.

      Of course I’m not always on the most bleeding edge forefront of the latest hyped model so there might be better alternatives, but I’m happy with what they provide

      • dwedge 20 days ago
        Their OCR did worse at reading a 40 page PDF of printed text than tesseract for me, I just paid $2 for useless output
    • redox99 20 days ago
      Qwen 3.8 27B absolutely demolishes Mistral best offerings at coding, and you only need a 5090 or 2x3090 to run it.
    • KPGv2 20 days ago
      Possibly because Mistral is a French-based company, so EU companies can cut the American umbilical cord a bit more.
      • A_D_E_P_T 20 days ago
        EU companies can download GLM or Kimi-K3 and get way better performance, though? I don't see any case for using a Mistral model when much better open models exist.
        • brendoelfrendo 20 days ago
          Because not everyone has the infrastructure to run that with better performance? Anyway, Mistral serves GLM 5.2 through their global and European endpoints, so if you wanted to leverage their services and use a more powerful open model, that seems to be an option.
        • KPGv2 19 days ago
          Most people don't want a second job trying to figure out self-hosted AI stuff.
    • perks_12 20 days ago
      The OCR stuff is quite usable. Their models perform good at some very basic tasks, so I use them to diversify. But their current model lineup is really terrible.
    • bossyTeacher 20 days ago
      Sovereignty.
  • matthieu_bl 20 days ago
    Summary of the different scenarios

    Service: Vibe Plan: Non-Enterprise Default: Opted in Opt-out possible? Yes

    Service: Vibe Plan: Enterprise Default: Opted out Opt-out possible? Yes (admin-managed)

    Service: Mistral Studio/API Plan: Not specified Default: Not stated, I assume opted in Opt-out possible? Yes

  • munksbeer 19 days ago
    I can understand companies not wanting to leak sensitive data or information, so not wanting their inputs used for training.

    But the individual level moral perspective confuses me. You object to your own input being used for training "for free", but you're ok to use models which already slurped the data of millions of other people "for free"?

    I don't get it. Obviously this is going to be a controversial take on here (I am not blind to the sentiment), so, please help me understand. If you're a conscientious objector, why are you using the models in the first place?

    • greggoB 19 days ago
      Private users also have sensitive information which, if leaked, could be used to their detriment (e.g training data getting hacked, etc).

      I don't see any distinction between companies and people when it comes to moral objections re training data - at least, if that's a thing, I'm unaware of it.

      I do agree with your last point re the hypocrisy.

    • dsign 19 days ago
      There's something to this argument. I for one use an API proxy (Kagi Assistant) when I'm submitting anything remotely sensitive to an LLM. But for coding in my hobby project, I have marked "use my conversations for training" in the Claude app. I do it in the off chance that one day we will use stupid amounts of compute for personalized medicine and to fight cancer; code for scientific and technical computing is really complicated and only a tiny fraction of humans can produce it.
      • munksbeer 18 days ago
        If I understand your comment correctly (and highly likely I missed something), I am very similar. Using stuff at work the AI team tries their best to ensure we don't leak our internal data or code.

        But for my personal use, I'm totally fine with helping train the models.

        I'll be called naive, but I'm on the optimist side of the fence here. I hope, and expect, AI actually frees us up and delivers much improved lives for everyone. I know this goes against the current zeitgeist, but it is genuinely what I think will happen rather than the dystopia most people predict.

  • maxdo 20 days ago
    It’s a spyware, but a sovereign one
    • kieranmaine 20 days ago
      How so? I've opted out of using my data for training.
      • pixl97 20 days ago
        How many times with large companies have you found that they removed the old opt-out value and put a new one in that is enabled by default because of course it's opt-out?

        Opt-out doesn't mean dick when the regulatory environment allows them to do things like the above without any recourse. Of course you can go to any other company that follows the exact same rules if you'd like.

        • icantevenhold 20 days ago
          Tbh there is much more regulatory pressure on EU AI companies than US/China - of course you can only use local models which Mistral actually releases unlike the US alternatives.

          So out of all the bad options this still seems to be one of the best

          • inkysigma 20 days ago
            I'm curious what regulatory pressure there is on EU AI companies that wouldn't apply to many of the US firms. The AI Act and GDPR still apply to the tech giants since they operate in the EU (e.g. Anthropic watermarking). There are labor laws that are definitely more strict but many of the US firms have operating presence in the EU as well. I imagine that being directly in the EU probably does increase regulatory oversight but I'm curious how this would manifest in Mistral in particular.

            Regarding local models, Gemma, GPT-OSS, Nemotron, and Inkling (maybe upcoming Muse series as well) all are fairly decent options at a variety of hardware costs.

            • schubidubiduba 20 days ago
              These laws apply to US companies in a more limited way because if we fine them too much they complain to US government that other countries have laws they have to follow. And then the US government intervenes to protect their money and data/intelligence extraction machinery (big tech companies). Usually with at least partial success.
      • teekert 20 days ago
        Defaults matter. The expectation for an organization tier plan is not that you have to go and ask each user in your org to please turn off "Allow the use of your interactions with Vibe to train Mistral's AI models" before starting your work.
    • zoobab 20 days ago
      What do you expect from running on "someone else computer?"
      • r_lee 20 days ago
        a service provided to you for the money you paid? just like with any other business?
        • pixl97 20 days ago
          Sheet, that is so 20th century.

          All the other businesses have been bought up by VC, going to rental models, and looking for new and interesting ways to screw you over too.

  • tensor 20 days ago
    I just checked my admin dashboard and training on my data is (still?) off. Maybe I already set it to off when I signed up, I don't really remember.
    • teekert 20 days ago
      They don’t switch it for you probably, that would be very bad. But the default for new users and orgs on the Team plan just changed. But they didn’t update the docs for a while which left me confused as I started my subscription after they switched but before they updated their docs.
  • olejorgenb 20 days ago
    > In the Admin panel[links to the admin panel], open the Privacy menu in the left-hand navigation bar.

    Why not link to the https://admin.mistral.ai/plateforme/privacy panel directly

  • zelphirkalt 20 days ago
    I remember this being the default for a while already. When I signed up and got an API key, I remember I did turn off the training option specifically, so it must have already been a default to have training on input on.
    • teekert 20 days ago
      This is about the Team plan specifically, which was like enterprise before. And it’s about the removal of the option to disable training on user input for all users from said plan at the org level, all last week.
  • Noaidi 19 days ago
    No one will read this but AI probably. I am done with contributing to AI. I never used AI but even posting here I am contributing to it as AI is known to scrape HN.

    https://alterlab.io/blog/how-to-give-your-ai-agent-access-to...

    So no more data for them. No more social media, no more cloud storage, just no more.

    It's been fun, but I am smarter than AI, so bye.

  • Lucasoato 20 days ago
    Question: when AI companies say they don’t train on your input or output, do they mean that they don’t train on an extremely simple AI rework of your input or output as well?
    • j4k0bfr 19 days ago
      Damn, that's an interesting point. Having a model transform your work into 'independent' IP and then training on that. I would assume any actual trade secrets (e.g. recipes, financial info) are protected but... tech companies have bent the law before.

      I'm assuming they've all at least thought about it quite hard, which is worrying in-and-of itself :).

  • rezonant 19 days ago
    Can anyone shed light on why there's a button for copying specifically for LLMs ("Copy for LLM") and why there's also an "Open in Claude" option on these articles?

    Regarding the former, there's no normal Copy button so presumably this just copies the content of the article. Not sure why it needs to be clarified that its for LLMs.

  • f6v 20 days ago
    Same company that has a partnership with Saudis, btw.
    • icantevenhold 20 days ago
      This means exactly what?

      I think you would be hard pressed to find any relevant tech company that doesn’t have a relationship with the Saudis or is funded by them - or any government for that matter…

      • f6v 19 days ago
        Yeah, let's just sweep it under the rug then!
        • icantevenhold 19 days ago
          Nah but be a little bit more specific than just sprinkling random FUD - what is the relationship and why is it problematic ?
  • __MatrixMan__ 20 days ago
    I'd like to provide extra context to help with training. Like, here's my codebase and the logs and the docs, feel free to ask me questions about it... Whatever makes the next version more applicable to the problems I'm trying to solve would be excellent.

    I just wish I could force them to share it with their competitors also.

  • creativeSlumber 20 days ago
    > Vibe: users are not opted out by default

    > Vibe (Enterprise): customers are opted out of training by default

    So they made it opt in for enterprise (how it should be), but intentionally made opt out for regular user.Basically saying "screw you: to regular users.

    Any self respecting user should stop using them.

  • Diti 20 days ago
    This gets posted literally moments before I was about to pay them after ditching Claude. Thank you! I hate data collection in paid products.

    I think I will be using Kagi Ultimate for the inference UI, so the data is somewhat anonymized before being collected.

    • throwaway89201 20 days ago
      I'm really hoping Mistral will succeed with their open models, so I'm a bit biased, but I don't have any affiliation. This submit is a bit of well-meaning but confused scaremongering however.

      Mistral has had, and continues to have, a toggle in the admin settings that permanently disables training on your data. The option has not been removed, and previous opt-outs are still honored. As far as I know, Mistral always trained on your data by default except for the enterprise plan, with the option to disable it on all plans, and with the option for organizations to make the choice for all your users.

      Kagi Ultimate still uses mostly closed models by OpenAI / Anthropic, etc.

      • teekert 20 days ago
        There is no toggle in my admin settings to disable this for the whole org (when on the Team plan). If it was there before they recently removed it.

        "Mistral always trained on your data by default except for the enterprise plan" - This is not true, until last week all docs stated that the use of your interactions with Vibe to train Mistral's AI models was off by default on the Team plan. Now that is only still the case on the enterprise plan.

    • teiferer 20 days ago
      > I hate data collection in paid products.

      You can opt out in this one.

      • trainingonme 19 days ago
        You can, but you still have to opt-out. It'd be better if it was by default.
    • evanjrowley 20 days ago
      My best experience with Mistral has been with their expensive GLM-5.2 model. It's actually developed by Z.ai, but unlike Z.ai, the GLM-5.2 at Mistral can be used at a low price without training on your prompts - but only if you remember to hit the privacy toggle in their admin settings.

      Planning code changes with GLM-5.2 using a Mistral Studio API key and implementing code changes using the Mistral Vibe API key has worked well for me. At my basic subscription tier, Vibe will share data with Mistral. It works for me because, when it comes to privacy, I care less about the actual code and more about the planning / high-level stuff.

      • nozzlegear 20 days ago
        > My best experience with Mistral has been with their expensive GLM-5.2 model. It's actually developed by Z.ai, but unlike Z.ai, the GLM-5.2 at Mistral can be used at a low price without training on your prompts - but only if you remember to hit the privacy toggle in their admin settings.

        Same here, actually. I'm American but have had a Mistral sub to supplement local models. The Mistral models, IMO, are very hit or miss, and I was about to cancel my sub until I saw they added GLM.

      • Zambyte 20 days ago
        Thanks for the pointer! I had set up mistral a while back using pi to compliment my local Qwen usage, but I found Qwen to just be better for all of my tasks at a substantially cheaper price, and reasonably fast tps. I also tried signing up for z.ai, but they would never send me an email to sign in. Using GLM through mistral to compliment my local models seems like exactly what I want.
        • evanjrowley 16 days ago
          Qwen is leaving a strong impression on me as well. Qwen-3.8-Max via openrouter has been a strong performer for me. I run it in autolith and then have claude code check the commits. Claude highly praises the quality of the work. Lol
    • saaaaaam 20 days ago
      You can opt out.
    • teekert 20 days ago
      You can still easily disable "Allow the use of your interactions with Vibe to train Mistral's AI models." What's annoying is that they switched this to on by default on Team plans with no way to turn it off for your whole team/org, this week.
    • rvz 20 days ago
      Just run your LLMs locally instead of using external providers.

      Using Claude, Mistral and the rest of them is not going to solve your issue with data collection.

  • sbinnee 20 days ago
    I remember mistral provides some free credits for opting in data for training. And for some reason I could not find this option anywhere in my console. Does someone know if this plan still exists?
  • TZubiri 20 days ago
    i don't get the legal aspect of this

    if you sign the contract (click I agree on Terms of Service), and it says that they do not train on the data, by what right can they backtrack on that?

    Maybe it is assumed that they notified you and you have the right to terminate the contract? Relying on some clause where the contract can be updated at any time. Or terminated at any time (with the presumption that a clause change is a termination and an automatic signing of the new contract, but that's weak)

    • teekert 20 days ago
      I think it only goes for new users/orgs. I was caught between the time they changed the default and updated their docs. I don't believe they would alter users/organizations exiting choices. In facts, from some of the responses it seems that some still retain the toggle to opt out their entire org.
  • selicos 19 days ago
    Run local or expect some data to be used for development. These companies did not create these tools by paying fair market value for the content.
  • vb-8448 20 days ago
    My bet is that mistral models will suddenly start to shine.
  • htrp 20 days ago
    LeChat when it was launched as a consumer product was always going to be a data acquisition play.

    This is them just making it very clear and disclosing as per European rules.

  • bigbuppo 20 days ago
    Great news if you want your LLM to be horny, I guess.
  • lukeschlather 20 days ago
    I'd be interested in some legal/GDPR takes on PII handling in prompts. If someone enters PII into a prompt, and Mistral retains it for training, is it sufficient for them to say "don't enter PII into prompts?"

    Of course with Claude and so on this bothers me too, but it doesn't seem like there's any real recourse under US law. But I would hope that "oh you shouldn't enter PII" isn't going to cut it under European law, that if I say "don't store my prompts, they include PII I don't want you storing" should be sufficient here under the GDPR and Mistral shouldn't be able to just store it anyway.

    • TZubiri 20 days ago
      They might circumvent this by adding a prompt like "remove PII from this prompt", so I don't think it's a worthwhile route. The issue exists whether PII is input or not, like IP or secrets.
  • threecheese 20 days ago
    Realistically, what are the risks? I get not wanting info you provide to AI being used against you in the future, like if you are gay and move to a country where that’s illegal (or live in one where it becomes illegal), but is “training” an actual risk here? I would assume that PII is redacted, and so beyond redaction failure, the real risk here seems to be to intellectual property and not to personal privacy.

    Edit: Does “use for training” include “we store all your chat logs forever tied to your identity”?

    • randomblock1 20 days ago
      The retain it for "the duration necessary to achieve the intended purposes", which could mean forever.
    • LtWorf 20 days ago
      Seems in USA you can go to jail for going somewhere else to legally have an abortion. So for example that.
  • perks_12 20 days ago
    I assume they do this no before releasing new models. It feels like they expect higher user influx from their new models.
  • croes 20 days ago
    > users are not opted out by default

    Of course because you can’t be opt out by default that would be called opt in.

  • xtiansimon 19 days ago
    Ha! I was researching local llm’s training yesterday—I _want this_ for my local llm project.
  • _blackhawk_ 19 days ago
    ZDR on OpenRouter + Aperture by Tailscale. And you should be good with the right config
  • DanielHall 20 days ago
    Just to achieve a great ideal: MEGA Make Europe Great Again.
  • daRealDodo 19 days ago
    You can check-out any time, but you can never leave.
  • scotty79 20 days ago
    I wonder how much the progress is slowed down just because most companies don't train on consumer convos. Assuming they really don't.
  • danelski 20 days ago
    Most commenters here clutching their pearls as if Claude and Gemini Pro didn't do it already. In the latter you (as a paying customer) can't even store the chat history unless you agree to their 'improvement of services'. Do you have all your accounts paid for by the enterprise or you never check the settings?
    • rdm_blackhole 20 days ago
      > Most commenters here clutching their pearls as if Claude and Gemini Pro didn't do it already.

      That's not the point though. The problem is that in the minds of lots of people including here on HN or otherwise, a service based in the EU is de-facto "more" respectful of digital privacy.

      You can read the comments on threads related to the EU tech where you will find people defending to the very end that privacy is better in the EU and that EU providers will never stoop as low as their US counterparts.

      As always the truth is a lot murkier than that.

      Yes, some services based in the EU are better in terms of privacy but it's not a given for all of them and it depends entirely on the service. Unfortunately such a nuanced take is not wildly popular in the tech world in this day and age where every US company is labelled as an evil data hungry entity and EU companies are portrayed as saints in this regard.

      That's where the first problem lies.

      The second problem is that for years now, people have been singing the praises of Mistral as a privacy friendly alternative the the US juggernauts because Mistral's headquarters is located in the EU and unfortunately today it seems some people are waking up to the fact that Mistral is doing the same thing than its US counterparts and they are disappointed which is understandable.

      Who is to blame for this dichotomy? Is it Mistral who leaned too much on this marketing angle (the European Chatgpt without the invasive tracking/ better privacy settings) or is it the users who failed to realize that EU or not, Mistral wasn't going to pass on the opportunity to improve its models this way?

      My hunch is that it's both.

      • icantevenhold 20 days ago
        No one is saying that every US company is an evil data hungry entity and EU companies are saints.

        But fact of the matter is that the US is rapidly sliding into totalitarianism and at the same time US tech companies have an hu huge influence worldwide.

        I commend any alternative that comes from outside the US and also offers an “open source” selfhosted option.

  • romanovcode 20 days ago
    I'm sure that the 3 users they have are very pissed about this news.
  • krunck 20 days ago
    Very disappointing. I really get the sense that all AI companies and anti-privacy parasites on society.
  • EFLKumo 20 days ago
    Come on, think how fast Grok evolves after SoaceXAI acquired Cursor. I just believe if you can't get enough real data from practice then you can't train a good model. And for Mistral collecting data is a must step no matter what approach it takes.
  • ylisav 20 days ago
    Very disappointing, but who doesn't train on user data? It's the only moat they have
    • mhitza 20 days ago
      I had the option, and had it disable everywhere, manually. The only one I didn't use, at all, is Anthropics service because of iffy Dario aura and their terms of service. Where is training in user data enforced and not possible to opt out?
  • amelius 19 days ago
    Can we have a unicode character that means "don't use for AI training", and another one for "this is AI generated data", and while we're at it one for "don't use for advertising purposes".
  • xyst 20 days ago
    the rapid enshittification in LLM hype cycle is something else.

    Got to love the private equity parasites ruining everything for the sake of profit.

    • senordevnyc 20 days ago
      I didn’t realize Mistral was owned by private equity.
  • moffkalast 20 days ago
    I'd send them a GDPR request, if I had an reason to use their fifth rate cloud models.
  • ex1fm3ta 20 days ago
    I mean they all do it right ? Mistral is just the first one to publicly say it.
  • camerondennis 8 days ago
    [dead]
  • mjprintz 11 days ago
    [dead]
  • greenjudge 19 days ago
    [flagged]
  • beyondscale-yes 19 days ago
    [dead]
  • alescalaios 19 days ago
    [flagged]
  • gkbrk 20 days ago
    [dead]
  • LoveMistral 20 days ago
    They wouldn’t dare train on language inputs from the wild, but they will happily strip out any code snippets you’re providing and sell/train on that à la GitHub/Copilot. File uploads too
    • TZubiri 20 days ago
      What? They dared train on random websites like reddit, training on language inputs from the wild is the name of the game
  • luciana1u 20 days ago
    [flagged]
  • luciana1u 19 days ago
    [flagged]
  • chris_explicare 20 days ago
    [dead]
  • asyncze 20 days ago
    [dead]
  • floki165 20 days ago
    [flagged]
    • throwaway89201 20 days ago
      This isn't the case. There's a toggle on https://admin.mistral.ai that allows you to disable training for both Vibe and Console/API for your entire organization, I just checked.
      • teekert 20 days ago
        I think they removed the toggle for new customers (on the Team plan) because I don’t have it. If I had, I probably wouldn’t have posted this.

        I mean it’s still annoying that I subscribe my org because they say the toggle is off, then find users telling me in fact it is on. Then to find that I can’t disable it org wide and have to ask each user to “please watch out” is pretty humiliating.

    • shujip 20 days ago
      Yes. If the privacy control only exists per user, it isn't a control for a team.

      Defaults matter more than the blog post. "You can opt out" is not the same product as "the org can actually enforce opt out."

  • ycCantCode 20 days ago
    [dead]
  • DerDerDaIst 20 days ago
    Oh come on. So you should be fine with this because "hey we're the friendly europeans..."?
  • r_lee 20 days ago
    European innovation right here

    you can't make this shit up