Microsoft says email spammers are adopting ASCII smuggling

(arstechnica.com)

35 points | by utiiiD 4 days ago

8 comments

  • CrimsonRain 2 hours ago
    Getting bombarded by "Paypal" <[email protected]> spam for months and Microsoft (outlook) can't handle that. This is beyond their level.
  • gschizas 3 hours ago
    Why not sanitize and denormalize (or whatever it's called) the text before feeding it to the spam filter? Or the LLM prompt?
    • gerdesj 37 minutes ago
      MS and the other big buggers are pretty incapable of running an email system that you would really want to use. They also have to support those that you might consider spammers too.

      However, I'm sure that you are also familiar with the HN standard issue meme that it is impossible to run your own email system.

      So, your email provider is crap and you can't do it yourself!

      Bollocks! I use Exim (1) and rspamd (2) and run quite a few bijou email systems. It does require some effort.

      You pays your money and you makes your choice ...

      1. https://www.exim.org/exim-html-current/doc/html/spec_html/in... 2. https://docs.rspamd.com/

    • thephyber 2 hours ago
      Do you have any evidence they didn't?

      My suspicion is that the spam filter programmers didn't do a comprehensive evaluation of every code point on every plane of Unicode because... well that's a massive job. So your "sanitize and denormalize" tasks are actually massive mappings which were likely imperfectly created.

  • jongjong 12 minutes ago
    Currently, my experience is that spam filters often block non-spam messages and allow many spam messages through. Platform notifications are spam! Yes, even if I signed up for it, I agreed to be contacted by the company, I never agreed to be spammed by them 100 messages a day.

    Platforms should aggregate all their notifications into a single daily or weekly email unless the thing is marked urgent! It's not difficult to do!

  • k12sosse 1 hour ago
    As a [email protected] holder, unfiltered spam is less commonplace than a) websites adding me to distribution/mailing lists without verifying I was the person to type it in, b) people who either accidentally transform their address into mine, or forget part of their address.

    These are from all parts of the globe. I have access to bank accounts in South America, Disney employee music royalty earnings tax disclosures and forms, European subscribers online platforms, AWS account recovery options, veterinarian records in Studio City, private school/PTA leadership website access in Mountain View.

    I stopped trying to return unopened mail, nobody cared.

    I don't do anything with any of this because I'm not a giant fool, but people, have your users verify their email addresses before you trust them.

  • perching_aix 2 hours ago
    A few years ago, during a particularly spam heavy period, I got pissed at Google's ineptness at combating spam, and decided to have a think about how I could get rid of the problem for good.

    What I arrived at:

    - I should never hand out my actual email address. As in, should be all proxied, every contact reaching me via a different address only known to them. Ideally, these aliases would be generated using a CSPRNG, and would be of sufficient length. Allows for tracing contact provenance in "space".

    - I should rotate that address periodically, if possible. Allows for tracing contact provenance in "time".

    Together, this would tell me who's the source of any particular influx of strange mails, and when did they leak the contact address provided to them (willfully or otherwise). This would enable me to separate the wheat from the chaff and nuke that address for good, inform the other party, etc.

    I was really taken by this idea, even wondered why this has not been baked into the underlying protocols over the years, to make it readily deployed and available for all, making it effortless.

    Because boy is there an effort involved! Several years later, I now avoid spam mail by simply no longer reading my emails anymore...

    Sometimes I wonder what corporate IT thinks when I pass their anti-phishing test a month after the test campaign.

    • SoftTalker 1 hour ago
      I estimate that I can identify 99.5% of spam by the subject line alone. I don't know why its such a hard problem, and why renewal notices for Norton 360 Premium sent to "Customer" from a random gmail address keep passing as legitimate email.
      • cam_l 41 minutes ago
        Hell, just block all emails mentioning Norton 360. Wouldn't be any great loss to the world.

        But what I don't understand is MS letting these through the spam filter but randomly, in a thread between myself and another person both using MS account emails, sending a single email in that thread to spam.

    • smalltorch 1 hour ago
      >Several years later, I now avoid spam mail by simply no longer reading my emails anymore...

      I realized the other day I haven't even checked my main email in months and I missed absolutely noting important.

  • righthand 1 hour ago
    Microsoft is a spammer and spamming platform themselves so this seems like a distraction.
  • rkagerer 3 hours ago
    Whoever thought it was acceptable to have a string of text that renders something unreadable or not immediately obvious to the human eye, was a complete moron.

    I realize ASCII was limited, but one thing I like about it is I can understand every single character code, and program to handle all the edge cases with certainty.

    • kevin_thibedeau 3 hours ago
      The tags were needed for language indication to control CJK glyph variants. Flag emoji were grafted onto this scheme. The key is that tag sequences have to start with a valid introductory codepoint. Simple enough to strip out anything that isn't a flag.
      • thephyber 2 hours ago
        > Simple enough to strip out anything that isn't a flag.

        But it's not simple enough for every company/app to independently research the entire Unicode code point space (which is MASSIVE) to find out what kinds of "fix ups" that app needs to do to clean the data it consumes.

        It's complicated or at least has difficult tradeoffs. Maybe you invest a lot of time carefully surveying all of the Unicode planes and decide which ones you care to keep unchanged and which ones you filter/strip. For every code point you reject or change, there is going to be some user who is confused or dissatisfied with the limitations of your app.

        • twoodfin 1 hour ago
          “Everyone knew” that in-band signaling was an awful, horribly insecure idea… until we found magic math that couldn’t “think” any other way.

          I don’t think it’s reasonable to blame the Unicode authors for not anticipating this turn of events.

    • embedding-shape 2 hours ago
      > I realize ASCII was limited, but one thing I like about it is I can understand every single character code

      It's great for teaching and other things, but everyday life is filled with many characters, is the suggestion we'd have one ASCII per language where there is more distinct characters, or what would we do? I don't see what else we could have done, that would have worked for the world, but I'm curious to hear ideas.

    • thephyber 2 hours ago
      The top comment (the Staff Highlighted one) explains why this range of code point exists.

      There was a rationale (ISO country codes to modify a flag to display that national flag).

      Maybe the problem wasn't the proposal, but the lack of the ability for others to reject it for being insecure.

    • SoftTalker 3 hours ago
      Even classic ASCII has "unreadable" control codes, but to be fair they would not be confused with text even by an LLM. Well probably not.