Show HN: Shoehorn – Quantize any model down to run on your machine

(notactuallytreyanastasio.github.io)

58 points | by rhgraysonii 3 days ago

6 comments

  • jedbrooke 51 minutes ago
    I gotta laugh at some of the models it suggests, for example:

    > AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

    you’re telling me you managed to fit Fable 5 into just 4B?

    • chompychop 25 minutes ago
      I gotta laugh at your thought process: knowing Fable 5 is a large frontier model, you're telling me that the first thing that came to your mind on seeing that model name is that it's a quantized version of Fable? As opposed to a distillation/fine-tuning on Fable responses?
      • unrented7977 5 minutes ago
        Don't make fun of people you think are ignorant, it's a pretty shitty look
  • hmokiguess 3 days ago
    • rhgraysonii 2 days ago
      LLMFit tells you what can run on something. I built something quite similar to their search into Shoehorn now.
  • akshay_akula 2 days ago
    This is interesting. I wonder how it could work with something like https://github.com/JustVugg/colibri.
  • mbuchel-hn 3 days ago
    does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?
    • rhgraysonii 2 days ago
      Yes that is exactly what this does.
      • kennywinker 2 days ago
        Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?
        • akshay_akula 2 days ago
          Wondering the same thing but for 48gb M5 Max.
  • jaylane 3 days ago
    tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running
    • rhgraysonii 2 days ago
      If you could post an issue if you still have the error around that would be awesome.
  • kelvo_ran 1 day ago
    [dead]