AI companies leak data to advertisers [pdf]

(jorgegarciaherrero.com)

109 points | by damaru2 2 hours ago

15 comments

  • Coeur 1 hour ago
    "multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."

    Not good at all.

    • postalcoder 32 minutes ago
      My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.

      Perplexity does this. Visiting a past perplexity search url exposes your full conversation.

      • albert_e 16 minutes ago
        Security by obscurity -- such an age old anti-pattern!

        I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link

        Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped

        There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines

        This is shockingly lax approach to data security and privacy by design

        • postalcoder 11 minutes ago
          Chat UIs are a minefield of “if you accidentally click this your data will be shared or trained without you realizing it!”
      • msdz 21 minutes ago
        Genuinely asking: If you don’t share the UUID-based URL yourself, what makes it not privacy-friendly?

        It’s not like someone’s gonna guess that URL… right?

        • postalcoder 18 minutes ago
          Yes, technically, guessing a url is impossible. But browser histories are stored in cleartext and trivially accessible to sketchy actors.

          I also accidentally paste random stuff into input boxes all the time.

  • drywater2 9 minutes ago
    No, they don't "leak data", data is sold. Leaking data requires a mistake. This is intentional.
  • j4k0bfr 25 minutes ago
    This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!

    My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.

    Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).

    • amarcheschi 19 minutes ago
      I'm taking an onboarding process for an Ai company helping other (much) bigger Ai companies and the amount of vibecoded platforms and documentation is staggering. Like, training process so broken that the platform just doesn't load sometimes, things that have never even been tried are published and you have to use them and they suck so much because it is apparent that no human ever touched that and probably wouldn't want to
    • alansaber 21 minutes ago
      AI companies rushing an implementation? Surely not :).
    • mrweasel 5 minutes ago
      So I only read the abstract, but the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

      If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).

      I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.

  • pluc 23 minutes ago
    You thought... they didn't?
  • Traster 15 minutes ago
    I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

    First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

    But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

    The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.

  • 999ziyadej 2 minutes ago
    data is sold i think
  • gagan2020 36 minutes ago
    All sells but I saw Chinese models are upfront about that most of the time.
  • DrMandalay 42 minutes ago
    The word is "sell" not "leak". This title takes away all agency from the thieves selling private data to advertisers.
  • charcircuit 56 minutes ago
    The paper doesn't say when the app sends the conversion artifact.
  • classified 1 hour ago
    Is it still called a leak if it was the whole point and purpose of the deal?

    Someone should have to investigate, but I suppose it's all "legal"?

    • lava_pidgeon 30 minutes ago
      In the US.

      In EU law it is very likely against GDPR.

  • reedf1 34 minutes ago
    my first guess is always Gboard.
  • robertclaus 21 minutes ago
    Hanlon's Razor given that these tools are almost certainly vibe coded at this point?
  • folkrav 59 minutes ago
    Insert surprised pikachu meme
  • damaru2 2 hours ago
    Entire conversations via exposed permalinks. For Grok: trackers receiving the conversation URL could access the full chat because the link lacked access controls.

    Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.

    Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".

    • Traster 10 minutes ago
      That is a staggering level of incompetence.
  • irregularbowels 2 hours ago
    [dead]