Rampart: Browser native on-device PII radaction

(ndstudio.gov)

41 points | by nateb2022 1 day ago

9 comments

  • dwa3592 2 hours ago
    I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

    A few things jump out since this is done by the government:

    - the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

    - it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

    - technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

    • bberenberg 1 hour ago
      I think all of your points are very valid, but it doesn’t remove the point that the government is trying to make it free and easy for application builders to do a little bit better than they are today. This is commendable on its own, even if it’s not perfect.
  • bob1029 2 hours ago
    I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

    If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

    PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

  • handfuloflight 1 hour ago
    Why did the National Design Studio see the need to put all the readable text on the right column of the page?
  • Onavo 14 minutes ago
    Is this good enough for HIPAA?
  • swiftcoder 52 minutes ago
    What we'd really like is a PII redaction model for video...
  • nhinck2 2 hours ago
    98.4% is nowhere near good enough to call it PII redaction.
  • iAMkenough 2 hours ago
    So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal

    https://www.csoonline.com/article/4046997/whistleblower-doge...

  • sgnelson 2 hours ago
    The skeptic in me really can't trust our current government to protect my information.
    • BowBun 2 hours ago
      Good thing you don't need much trust in this case. Source available here - https://github.com/nationaldesignstudio/rampart

      I suppose the model could be doing stuff, but you can also switch that out for your own with the source.

    • Cider9986 1 hour ago
      You could never trust the US government to protect our information since it moved online.
    • goodmythical 2 hours ago
      I mean, it can be run locally, so you don't necessarily have to trust it, unless there's been any model-as-a-vector CVE.

      That hasn't happened yet, has it? Where running a model from HF directly compromises the machine as opposed to some breakout or exfil done by the model after the fact?

      That said, if you're filing your taxes or have a 'real' ID, the government already has all pertinent information required to fuck you over either deliberately or via a leak, so...

  • Cider9986 1 hour ago
    The source code should be public domain, no?