10 comments

  • aliljet 1 hour ago
    It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
    • tomrod 1 hour ago
      They own their hardware. That vertical integration alone probably saves oodles because they can reconfigure to their needs as opposed to individually negotiating data centre plans. I can't imagine the complexity both OpenAI and Anthropic have to maintain for their deployments.
    • aleqs 1 hour ago
      > Opus 5.5 right now because it's nearly unlimited use

      Anthropic has some of the lowest usage per $ in general, not sure what you're taking about.

      • jjice 1 hour ago
        I don't think the OP and your comments are mutually exclusive. If Anthropic usage is lower, and the OP considers it basically unlimited for their case, that just means that they would have virtually unlimited usage with other plans.
        • aleqs 30 minutes ago
          Nothing about it is unlimited or even close to it. Just because you only use 1gb of your 2gb data plan, doesn't mean you have unlimited data.
    • zozbot234 1 hour ago
      It's not as smart as Claude Opus 5.5 High according to the AA benchmark. Looks like a big fat nothingburger so far, though it's possible that future fine-tuned checkpoints of the same pretrained model will do a lot better.
      • mpyne 36 minutes ago
        > It's not as smart as Claude Opus 5.5 High according to the AA benchmark.

        If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.

  • godbox 1 hour ago
    What a snooze fest. Another model that does not meaningfully improve on intelligence or price compared to its peers. Google has basically announced that they've "caught up" with the rest. I think they've been doing great work in the Flash department so seeing this is... underwhelming?
    • losvedir 1 hour ago
      The only models beating it are on "max", while this is "high". There's no guarantee that those effort / reasoning levels compare, but there's almost certainly an "xhigh" or "max" version later which will score higher there.
    • augment_me 1 hour ago
      I found that Google does not benchmaxx as much as the other providers. Of you look at real-case evaluation like lm-arena, even the 3.8 flash is often near the top despite its benchmark index being worse.
    • netdur 9 minutes ago
      flash on ai studio is my fav model to chat with, by miles
    • piyh 35 minutes ago
      > does not meaningfully improve on intelligence or price compared to its peers

      $10 per million output tokens isn't improving on frontier price?

    • enraged_camel 1 hour ago
      Based on some... rumors I've heard, this is their "Pro" offering. There is supposed to be an Ultra coming as well.
  • yipinwong 1 hour ago
    I am still not sold on Gemini 4 Argon yet from the chart.

    The price is enticing for cost per tasks, but let's see how it goes.

    I have montly (cheapy) sub to gemini models and has been underwelming and lowered the tier.

  • ai-x 1 hour ago
    Note: Google can sell their tokens at cost if they really want to drive out competition, but long term they are better off by everyone making a healthy margin (and Google does a double-dip by also selling compute, services).

    So, like any optimal game theory move, they are better off not starting a price war

    • gradus_ad 9 minutes ago
      Important to remember AI is a direct assault on their actually profitable business: search and ads.

      OpenAI and Anthropic are existential threats to Google and it will operate accordingly.

    • miohtama 29 minutes ago
      Google is losing in the market share - their share is 0. Anthropic revenue was $60B in the last 12 months. Google is not growing the pie, it needs to buy its way to the market share and to a seat in the table.
    • pooper 1 hour ago
      A slightly different question would be can they afford to "sit out" of any price war?
  • jwpapi 41 minutes ago
    Always the last model that gets announced is the best. The labs always know in advance.
  • dang 1 hour ago
    Related ongoing thread:

    Gemini 4 Argon - https://news.ycombinator.com/item?id=49913571

  • algoth1 1 hour ago
    The most impressive jump for me is in the low hallucination rate, which is specially impressive given how bad Gemini current models are on this regard
  • dom96 1 hour ago
    I'd love to run it on my benchmark but alas, Google not making it public prevents this.
  • anuragdaram 2 hours ago
    Will have to check what the pricing would be for this model.
  • A_D_E_P_T 1 hour ago
    Trading punches in the benchmarks with Mimo v2.6 and 6.1-Sol (both very cheap!), and decidedly inferior to Opus 5.5. I'm afraid this looks unimpressive. Rather comical that they're delaying its launch "for safety reasons".
    • asdfasgasdgasdg 1 hour ago
      From the charts, it's similar in intelligence and cost per task to both Opus 5.5 (high) and 6-Astra (max). It would be better if it were more intelligent and less expensive, but I don't see a reason to expect it to have better performance than models released around the same time.
    • godbox 1 hour ago
      To play the Devil's advocate, Claude loves chugging tokens, while Gemini appears to be quite a bit more conservative and efficient. I believe AA's price per task breakdown reflects this.
    • thereitgoes456 1 hour ago
      What are you talking about? It’s comparable to Opus 5.5 on “high” (54 vs 53; $1.82 vs $1.99), crushes every model except the most modern OAI/Ant ones, has way lower hallucination than every existing model and probably broader support for multimodal like existing Gemini models. This is so ludicrously off base.