Meta's Muse is fantastic for web scraping

(sigh.dev)

22 points | by STRiDEX 59 minutes ago

2 comments

  • dvt 31 minutes ago
    I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

    Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

    • bayindirh 20 minutes ago
      Thanks for letting us know that we need a new layer of detection systems.

      Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

      Humans are interesting creatures.

      • pcthrowaway 5 minutes ago
        There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.
    • kurisufag 3 minutes ago
      the 4get dev did the same thing a little while ago for his metasearch engine: https://git.lolcat.ca/lolcat/4play
    • rjtc 6 minutes ago
      anyone can spin these kind of side projects and do easy talk, but the moment you actually try to use this on signed-in Linkedin or Amazon its going to fail

      the only solution is to drive your regular browser with all your sessions/cookies via an extension

      • dvt 4 minutes ago
        > the only solution is to drive your regular browser with all your sessions/cookies via an extension

        This is exactly what I'm doing, but mocking/randomizing all the sessions/cookies/params (like resolution, OS, WebGL , etc.) in a separate Chromium binary. It's popular these days, but imo using your normal browser for agenting stuff is a very bad idea. These models do dumb stuff all the time.

    • frabcus 20 minutes ago
      Right, but what happens when everyone uses that at scale? Without agreements and standards it isn't pretty.
      • dvt 9 minutes ago
        It's sad, but this train has left the station a quarter of a century ago imo. Many people have since become billionaires scraping the web without anyone agreeing (Google, Yahoo, and now possibly Anthropic and OpenAI).
    • bel8 14 minutes ago
      Interesting, how is the chromium sandboxed? profile dir cli param? or deeper like chromium engine framework embeded in the application?
      • dvt 12 minutes ago
        Just simply a separate Chromium binary (not your usual browser).
    • d0vs 22 minutes ago
      FYI Muse smashes through those captchas natively without prompting.
    • ares623 25 minutes ago
      By describing it here you've already released it no?
  • arjunchint 25 minutes ago
    their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

    Muse's utility has significantly dropped with the blockages.

    To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.

    • sejje 1 minute ago
      they can just use the user ip. i think grok already does this.