13 comments

  • apimade 1 hour ago
    https://apimade.com/audio-compare.html

    Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

    • barney54 34 minutes ago
      Is Darwin TTS from Fish Audio? It wasn't clear when I searched for it.
  • rahimnathwani 8 hours ago
    For some reason it switched voices half way through a 33 second clip.

    For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

    • toebee 3 hours ago
      hey, thanks for letting us know! will look into the issue and see what went wrong.
  • asaiacai 8 hours ago
    This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
  • konart 4 hours ago
    All TTS generations are too fast. It's almost I'm listening to a podcast on 1.25-1.5x speed.
    • toebee 3 hours ago
      thanks for the feedback! will investigate and get it fixed
  • iharnoor 6 hours ago
    By next month the competition for TTS will be even more!

    Voice models are not winner take all market unlike LLM APIs

    Coming here as Developer Relations at AssemblyAI

  • yoloakki 7 hours ago
    You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality
  • mowmiatlas 7 hours ago
    Cool, I’ve released something to the same beat of the dr this weekend as well

    https://github.com/loudreader/loudkit

    I think real time natural tts should be possible everywhere soon

    • toebee 3 hours ago
      very cool. will try for local use!
  • meatmanek 8 hours ago
    > and Qwen3-ASR

    Is the ASR inference engine open source as well?

    • nshm 6 hours ago
      Yes, and it is very good one. Leading position on private leaderboard on HF: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
      • meatmanek 3 hours ago
        I meant the Nari inference engine for Qwen3-ASR. I'm aware that Qwen3-ASR is open source, but I don't see a repo under https://github.com/nari-labs for nari-qwen3-asr or similar.

        The Huggingface link on https://narilabs.com/product/stt/ links to https://huggingface.co/Qwen/Qwen3-ASR-1.7B , not anything under https://huggingface.co/nari-labs

        • toebee 2 hours ago
          the qwen3-asr inference repo is not OSSed as of now. we're planning to write a paper or tech report on it as it contains some general techniques for ASR inference.
          • kshmir 2 hours ago
            how do I follow you? I have a small 5090 doing inference all the time and I barely use tts but a lot of asr, mostly whisper, I ported your tech report for tts and implemented some improvements on my whisper inference based on your tech report as well!

            would love to talk sometime!

    • verdverm 6 hours ago
      They have a number of demos and examples in their HF space

      https://huggingface.co/Qwen/spaces

      I saw a local-ai demo (something + gemma), where the person used ASR to get text and gemma to clean it up (like turning "question mark" into a literal "?", bullet points another one). The presenter also showed a gemma only option, that did both in one go, but had a higher WER on average, and even though the formatting statements were handled without a multi-stage pipeline, they preferred the multi-stage overall

  • DylanMerigaud 7 hours ago
    Rooting for you on this one.
  • bilaly 4 hours ago
    [flagged]
  • ipsum2 8 hours ago
    If you're going to announce a TTS model, service, or whatever, you really need demos.
  • nthypes 8 hours ago
    [dead]