6 comments

  • sgtwompwomp 13 minutes ago
    This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
  • nateb2022 14 minutes ago
    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
  • amirhesham 15 minutes ago
    Oh this is so cool. Curious about the business model, too.
  • kenzic 11 minutes ago
    How long does tuning take (on an M3 MacBook Pro for example)?
  • yolandac 7 minutes ago
    does it allow us to run larger models that weren't possible before?
  • p-e-w 21 minutes ago
    What is the business model?