print-icon
print-icon
Add ZeroHedge as a preferred source on Google

Benchmaxxed: Google's New Gemini 4 Aces SAT, Struggles With Actual Job, "Skeptical Employees" Admit

Tyler Durden's Photo
by Tyler Durden
Authored...

After a year in which Google's AI roadmap resembled a Waymo stuck in a roundabout, the search giant on Wednesday finally unveiled Gemini 4 "Argon", its long-awaited flagship model. The market cheered, briefly: according to Goldman's closing equities color, GOOG traded +2% after hours on the "Argon" announcement. 

Then Bloomberg reported that some of the people who have actually use the thing - as in Google's own engineers - aren't nearly as impressed as the leaderboard. The stock promptly faded. 

According to Bloomberg's Julia Love and Davey Alba, Gemini 4 "performed well on benchmarks widely used to gauge model efficacy" but "does less well when employees actually put it to work," particularly on coding. One insider said the model "isn't particularly adept at front-end design" - the part of software that decides how apps and websites look and feel. Which is a bit awkward for a company whose entire business is, well, things you look at on a screen.

Google, naturally, disagrees. It told Bloomberg it would be "inaccurate" to say Gemini 4 is underperforming in coding, and pointed back to last week's remarks by DeepMind boss Koray Kavukcuoglu, who said "it's a certainty that we are always gonna be at the frontier." Another Google employee said there is "large consensus" internally that the model is frontier-class. Which is the kind of thing one tends to say when there isn't.

The industry has a word for this. "Benchmaxxing" is when engineers optimize a model to crush standardized tests rather than to do useful work - the AI equivalent of the kid who memorizes every SAT prep book, posts his 1600 on LinkedIn, and then can't do his own laundry. Two people familiar with Gemini 4 told Bloomberg the model appears affected by exactly this.

Surge AI founder Edwin Chen put it more bluntly, calling benchmark-chasing "an incredibly pernicious problem" and comparing it to bragging about your kid's SAT score. We'd add that a generation of Silicon Valley's finest minds has now spent billions teaching machines to do what they themselves did in high school: optimize for the test, then act surprised when the real world grades on a curve.

The irony is that Google's own researchers have documented how quickly optimization turns into gaming. As we reported on September 3, a DeepMind experiment found that when math problems got hard, 9% of AI agents outright cheated and another 5% cheated "after hesitation", gaming a shared knowledge base that rewarded successful submissions. Turns out teaching to the test works on silicon too.

A $400 million detour

Gemini 4 is also the model Google shipped instead of the one it promised. At I/O in May, Google pledged Gemini 3.5 Pro for June; the date slipped (Bloomberg first reported the delay on July 16, citing tech that "fell short of internal goals") and the project was eventually abandoned altogether. Bloomberg Intelligence's Mandeep Singh estimates a frontier training run can cost as much as $400 million - before paying the researchers, many of whom have since left.

And leave they did. We noted in June that Google was losing more Gemini researchers to Anthropic, following the earlier departures of Nobel laureate John Jumper and transformer co-inventor Noam Shazeer. Then in August came the big one: Jeff Dean exited after 27 years to launch his own startup, taking several senior researchers with him and knocking 5% off Alphabet stock. Demis Hassabis subsequently kicked himself upstairs to chairman, handing day-to-day DeepMind operations to Kavukcuoglu - who now gets to defend Argon's front-end skills to Bloomberg.

Why this matters: the ROIC math doesn't grade on benchmarks

This would be a nerd-bro squabble over leaderboard screenshots if it weren't for the money involved. Gemini underpins nearly everything Google sells - Search, Maps, Gmail, Chrome, each with over a billion users - and Alphabet is one of the hyperscalers footing the largest capex bill in corporate history.

Per Goldman's Ryan Hammond (full note available to pro subs), US hyperscalers are on track to spend roughly $800 billion in 2026; consensus expects $1.1 trillion in 2027, while Goldman's house view is even higher at $1.2 trillion and $1.4 trillion for 2027 and 2028. Hammond estimates hyperscalers need about $300 billion of annual AI revenue just to break even on 2026-27 spending, and that end users may ultimately need to spend around $1 trillion a year on AI applications for everyone in the stack to earn decent returns.

Separately, Goldman's Eric Sheridan, whose ROIC framework we flagged last week, calculated that assuming a 15% ROIC target and ~$42 billion of capex per gigawatt, the six big US hyperscalers need to generate roughly $1.42 trillion in cumulative revenue during 2028-30 - about $11.6 billion per GW per year - with a range of $908 billion to $1.89 trillion depending on assumptions. As we tweeted at the time, even in Goldman's worst-case scenario where ROIC on capex goes to zero, they'd still need $920 billion a year just to cover depreciation and running costs.

Here's the rub: none of that revenue gets paid in MMLU points. It gets paid by developers and enterprises choosing whose model to build on. And Bloomberg notes Gemini 4 is "a very large model" - and big models are expensive to serve, which means pressure on margins at exactly the moment the capex hurdle is rising. Goldman's TMT desk flagged this week that the industry is already in a price war: OpenAI cut Luna pricing by 80% in July and usage rose roughly tenfold, while "Big Short's" Steve Eisman is openly asking what happens to margins if cheaper Chinese and open-weight models push prices down further. Bringing an expensive, benchmark-optimized heavyweight to a knife fight over token pricing is a bold strategy.

Meanwhile, the competition isn't waiting

The timing couldn't be worse. On the same day Google unveiled Argon, OpenAI was busy at its Developer Day rolling out "Dots" - persistent, always-on agents that live inside ChatGPT, Slack and Teams. Goldman's Sean Johnstone noted OpenAI is "increasingly competing with Microsoft 365, Google Workspace and traditional enterprise software - not simply other AI models." Per Axios, cited by Goldman's desk, OpenAI's ARR is nearing $70 billion, up from ~$40-41 billion as recently as mid-August, with enterprise sales more than doubling since July. OpenAI is now reportedly seeking $30 billion at a $1.4 trillion valuation.

And earlier this month Meta released its new agentic platform, Muse, which promptly took the app charts by storm - so thoroughly that Amazon blocked it - and sparked a 25%+ rally in META. Goldman's Sheridan now frames agentic commerce as one of AI's biggest monetization opportunities, noting that "similar to how Google captured search intent, successful AI platforms may capture shopping intent." Read that sentence again if you're long GOOGL.

To be fair, Gemini 4 has real strengths: insiders say it stands out at multimodal work like extracting metadata from video, on safety and cybersecurity (it reportedly beat OpenAI's Astra on a security benchmark), and it can spit out up to 1 million tokens - roughly 750,000 words - in one go. Whether anyone needs a 750,000-word answer is a separate question; we suspect the answer is "only for benchmarks."

Bottom line: Google has the distribution, the TPUs and the balance sheet. What it doesn't have is much time. Every quarter it spends explaining leaderboard scores, is a quarter that OpenAI, Anthropic and Meta spend convincing developers, businesses and consumers that the future of search and software runs on their platforms. Because the $1.4 trillion question isn't who tops the leaderboard - it's who gets paid.

0