Gemini 3.8 Flash

(deepmind.google)

133 points | by bratao 46 minutes ago

26 comments

  • simonw 2 minutes ago
    Pelican (thinking effort high): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - 8.9742 cents

    Here's a 3.7 thinking effort high pelican for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - 8.4387 cents

  • mattlondon 16 minutes ago
    Currently top at https://deepswe.datacurve.ai - beating Opus 5!

    https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5!

    Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

    • markasoftware 4 minutes ago
      On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63.

      Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

    • WarmWash 11 minutes ago
      The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
      • scrlk 5 minutes ago
        Not just speed, also reliability. IME, Gemini's speed or quality doesn't degrade badly during weekday working hours compared to OAI and especially Anthropic.
    • satvikpendem 8 minutes ago
      We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.
      • onlyrealcuzzo 4 minutes ago
        And the benchmarks agreed with you... until now.

        So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.

    • onlyrealcuzzo 5 minutes ago
      The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were.
    • ttul 9 minutes ago
      Crushing it on DeepSWE is a very big deal. Excited to give this a try.
    • Gecko4072 15 minutes ago
      Google - we're so back
    • sunaookami 8 minutes ago
      >shows an intelligence score of 59, the same as Opus 5!

      ...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.

  • meh2frdf 7 minutes ago
    The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.
    • onlyrealcuzzo 3 minutes ago
      > The flash models, for coding are reckless in my experience.

      My experience is that antigravity is awful and reckless - but that the model itself isn't.

    • upcoming-sesame 3 minutes ago
      If by reckless you mean commit, push, deploy without me asking it to, the I agree!
  • mattlondon 24 minutes ago
    Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC?

    I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!

    At this point it is a meme of course, but where is 3.5 Pro :)

  • andai 19 minutes ago
    Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?
    • ipsod 16 minutes ago
      IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High.

      But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".

      Flash is my go-to for prototyping, and basically anything that isn't writing production code.

      • ramon156 8 minutes ago
        The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.
        • ipsod 4 minutes ago
          They've been my bet to win the AI race for a while. I was starting to doubt, but this 3.6, 3.7, and 3.8 arc has anchored me.
      • esafak 2 minutes ago
        Luna is way slow. I don't remember an OpenAI model ever being this slow.
    • realist_not 15 minutes ago
      It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.
  • xnx 18 minutes ago
    Seem like a great, no-compromise, upgrade over 3.7 which is already a bargain, fast, and doesn't have the brain-damaged writing style of Claude.
    • fitsumbelay 16 minutes ago
      that's certainly what it's looking like so far. kind of mind boggling ...
  • f311a 9 minutes ago
    Is the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.
    • ipsod 1 minute ago
      I haven't had any issues lately.
  • satvikpendem 9 minutes ago
    Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
    • rancar2 2 minutes ago
      That was sunset and replaced by Antigravity. FWIW until I abandoned it knowing the sunsetting, I was able to get good behavior out of Gemini CLI with overriding the system prompt. The default prompt crippled the harness with very poor instructions, but there was a hidden ENV to override it. Replacing it with Claude Code like prompts based on the model selected, it ran at a much higher intelligence level full stack with significantly less errors.
    • pshirshov 7 minutes ago
      There is no Gemini CLI anymore, nor you can use Gemini with your own harness unless you pay per-token.
    • zipy124 2 minutes ago
      It was superseded by the antigravity CLI.
  • kelvinjps10 11 minutes ago
    I see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.
  • hmokiguess 7 minutes ago
  • leumon 9 minutes ago
    [delayed]
  • pwython 13 minutes ago
    Is there any reason to even use 3.1 Pro now?
    • bitexploder 1 minute ago
      It is still going to be better at text work, skills, document review, deep reasoning, architecture review, etc. It is only 6 months old, it isn’t like its world knowledge and software knowledge is really out of date. Use it to churn on harder design problems.
  • ASinclair 10 minutes ago
    From personal experience it feels much more capable than 3.7 Flash.
  • sva_ 24 minutes ago
    • Barbing 11 minutes ago

        [1] For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red.
      
      Gemini 3 Flash?! So is Gemini 3.8 Flash less safe than 3.7 Flash in all areas besides Text to Text Safety (and identical on Image to Text Safety)?

      Why bother with a column “Gemini 3.8 Flash vs. Gemini 3.7 Flash” when you’re going to disregard the label for 20% of it? Also is the “Tone” label short for “Tone and Instruction Following”?

      Chartcrime, the major AI lab tradition.

    • mattlondon 22 minutes ago
  • fitsumbelay 17 minutes ago
    shows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way
  • tacomonstrous 24 minutes ago
    Looks like Google's given up on frontier models for external consumption?
    • heyjamesknight 21 minutes ago
      Gemini 4 pre training is underway: https://x.com/OfficialLoganK/status/2079594867161022817

      My guess is we skip 3.5 and go straight to 4 Pro. With the monthly Flash releases, releasing 4.0 Flash and Pro in 6-8 weeks would be a nice buildup.

      (I work at Google but don't know anything that isn't already public)

    • WarmWash 17 minutes ago
      Latest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).
      • VirusNewbie 5 minutes ago
        I found 3.5 pro to be much better than 3.5 flash, but 3.7 flash with high reasoning is comparable and way way faster.
    • iamdelirium 21 minutes ago
      How can you say that when a Flash model is benchmarking close to Opus and Sol?
    • thisisauserid 14 minutes ago
      They don't want to release a frontier model that requires data sharing with the government and right now it looks like they'd have to.
    • ok123456 22 minutes ago
      Given up frontier models for selling compute.
  • deanc 7 minutes ago
    And yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash
    • WarmWash 4 minutes ago
      Google has been doing staged roll outs on all their products since forever.
  • advenn 15 minutes ago
    But where is Gemini 3.5 pro?
  • realist_not 28 minutes ago
    Anyone has a cached page / mirror ? 404
  • OG_BME 28 minutes ago
    What did it say?
  • Mashimo 32 minutes ago
    It's 404 now.
    • freedomben 28 minutes ago
      Came and went in a flash
      • k8sToGo 18 minutes ago
        Because they are preparing Gemini 3.9 Flash
    • kingstnap 23 minutes ago
      The blog post is gone but I can currently use it in the gemini chat website.
  • yipinwong 25 minutes ago
    "Page not found"...
  • mythz 15 minutes ago
    It's just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out.

    And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.

  • shuvrojit 14 minutes ago
    Gemini is getting less useful with each update. I could edit a pdf with the 3-pro model before but 3.1-pro couldn't edit the given pdf nor it could generate one for me.
    • leumon 12 minutes ago
      You probably mean 3.5-flash? Pro is still good for a lot of use cases, but it seems it's still officially in the "preview" phase.
    • ipsod 13 minutes ago
      3.5 pro doesn't exist yet?
      • shuvrojit 10 minutes ago
        Sorry my bad, I messed up the numbers, 3 and 3.1 pro. All of these model numbers have me confused