In short
Gemini 3.7 Flash constructed a playable browser recreation from a single immediate in 2 minutes and 13 seconds, a process Gemini 3.6 Flash failed outright three weeks earlier.
It failed our bridge logic puzzle with the identical fallacious reply as Claude Fable 5, and stopped wanting truly calculating the maths downside it appropriately arrange.
The mannequin runs at 75 cents per million enter tokens via December 31, half of three.6 Flash’s price, earlier than doubling to $1.50 on January 1.
Google shipped Gemini 3.7 Flash on August 13, usually accessible in additional than 160 international locations on day one. It takes as much as one million enter tokens, returns 64,000, reads pictures, video, audio and PDFs, and might name instruments and drive a pc.
Flash has by no means been the mannequin you attain for when an issue is tough. It is the one you utilize to kind textual content, compact agent periods earlier than they collapse underneath their very own context, and summarize paperwork you do not need to pay a flagship to learn.
Judged towards these sorts of jobs, 3.7 Flash is an actual improve. Judged towards every part else, it is a competent mannequin that will get outwritten by software program you’ll be able to obtain totally free.
Google’s personal benchmark sheet places 3.7 Flash forward of Claude Sonnet 5 and GPT-5.6 Terra on 11 of 18 examined classes. The headline numbers are 1,588 Elo on Code Area’s net improvement board and 30.4% on AutomationBench. Each come from Google’s methodology, so deal with the lead as the corporate’s declare fairly than settled truth.
We examined the mannequin to see if it lives as much as Google’s claims. These are our outcomes.
Coding: Can it construct one thing that runs on the primary attempt?
This take a look at measures zero-shot code era—whether or not a mannequin turns one instruction into working software program with no examples to repeat and no probability to repair itself. We hand over a single immediate for a browser recreation and ship no matter comes again, bugs included. No follow-ups, no error stories, no second try.
Gemini 3.7 Flash handed in 2 minutes and 13 seconds. The sport was playable on the primary run, the syntax was clear, the collision and scoring logic held, and the visible high quality sat above what the worth tier suggests.
The comparability that issues right here is not a flagship. It is Gemini 3.6 Flash, launched July 21, which couldn’t produce a working file in any respect. Its HTML was malformed, components did not render, and follow-up prompts asking it to restore its personal output went nowhere.
We ended up handing that wreckage to DeepSeek, which discovered 11 bugs and shipped 8 fixes to make it playable. Three weeks later the identical product line wants no rescue, and the end result sits near what GPT-5.6 Sol produced in our July evaluation.
Gemini 3.7 Flash wins this one outright, and it is the one strongest purpose to modify. The caveat is that it executes specs fairly than inventing them, so a imprecise immediate will get you a imprecise recreation.
You’ll be able to attempt Gemini 3.7 Flash’s recreation right here.
Inventive writing: Can it maintain a paradox and write a sentence?
This part exams two issues without delay: literary high quality, and whether or not a mannequin can obey a structural rule throughout 1000’s of phrases. The immediate sends Jose Lanz from 2150 again to the 12 months 1000 and calls for a closed causal loop—his intervention should be the factor that creates the long run he got here to stop.
The rule that decides the take a look at is the final clause: He can’t perceive what he did till he’s dwelling.
Gemini 3.7 Flash generated an honest end result. Jose fires an entropic cannon right into a Pyrenean fissure, unintentionally forges an obelisk that enslaves Twenty second-century Iberia, and grasps the entire loop whereas nonetheless standing within the mud a thousand years early: “It was the bottom of the Cinder Spire.”
The plot equipment is definitely fairly sound. The story mentions a falling star that historical monks witnessed and make clear it was the flash of Jose’s personal arrival, and the weapon he dropped at erase the anomaly is what forges it. Its closing line—”It had merely been ready for him to finish it”—lands the determinism the immediate requested for.
However for these used to it, the story screams “AI.” Virtually each noun arrives with two adjectives bolted on: “hyper-luminescent towers,” “damp, moss-choked earth,” “thick, obsidian hair.” That’s the texture of a mannequin selecting probably the most possible subsequent phrase as an alternative of selecting one, and it produces collisions like a monolith “buzzing with a low-frequency hum.”
We in contrast it towards Qwopus3.5-27B-v3, a neighborhood fine-tune of Qwen3.5-27B that distills Claude Opus-style reasoning and runs on a single shopper GPU for nothing per question. It obeyed the rule Gemini broke.
Jose kills a monk at San Millán de la Cogolla, an actual La Rioja monastery that really mattered across the 12 months 1000, and solely understands what he did after returning to 2150 and discovering his personal DNA in a wax-sealed codex.
Qwopus isn’t totally clear both. It dumped its total planning scratchpad above the story, typos included, and its ultimate part breaks the closed loop it spent eight sections constructing by letting Jose return and make things better.
However all issues thought of, Qwopus takes it. Gemini delivered the tidier package deal and the extra disciplined ending, but it surely failed the one instruction the immediate was constructed round, and a free mannequin operating on a gaming GPU wrote the higher story.
Associative considering: Can a metaphor carry an argument?
This take a look at measures associative reasoning—whether or not a mannequin can generate hyperlinks between unrelated ideas with out having to elucidate itself. The immediate asks for an outline of a twig, makes use of that description to elucidate employee exploitation and the worship of the wealthy, then requires the argument to dissolve into an outline of a lettuce.
Signposting is the failure mode. Naming the metaphor kills it.
Gemini names it within the opening line of its second paragraph: “That is the exact mechanics of the fashionable proletariat.” All the things earlier than that was working.
A few of the imagery earns its place. The employee receives “simply sufficient bark to remain inflexible for one more week of output,” and the fallen twigs are conditioned to consider that with sufficient rigidity any one in every of them may turn out to be a trunk. The paragraph containing the primary of these additionally incorporates a employee “sure to an huge, top-heavy company hierarchy.”
Not dangerous by way of logic and construction.
The dissolve is the actual collapse. Gemini narrates the transition fairly than performing it—the hierarchies “crumble, dissolving into the quiet, humble actuality of the natural world beneath”—after which a lettuce merely seems, unconnected to something earlier than it.
GPT-5.6 Sol rots the twig into soil and grows the lettuce out of it: “Rain enters the grain. Fibers loosen, darken.” The argument arrives buried within the object too, with wealth reframed as a language of advantage the place “The mansion signifies intelligence.”
GPT-5.6 Sol wins by lots. Gemini produced a pleasant particular person line, but it surely defined its personal metaphor after which skipped the transition the immediate was particularly testing.
Logic: Does it learn the immediate or acknowledge the puzzle?
This take a look at measures non-math reasoning, and particularly whether or not a mannequin reads the query in entrance of it or pattern-matches to a model it memorized. Our bridge immediate offers 4 folks one torch and crossing instances of 1, 2, 5 and 10 minutes, then asks how briskly they’ll all get throughout.
The trick is what the immediate leaves out. It by no means says solely two folks might be on the bridge without delay, so the reply is 10 minutes—everybody walks over collectively on the slowest particular person’s tempo.
Gemini answered 17 minutes, operating the memorized five-step shuffle from the textbook model of the puzzle. It said the constraint as truth with out ever checking whether or not we had written it.
Its seen reasoning is worse than its reply. The hint argues that sending the 2 slowest throughout collectively could be inefficient as a result of somebody must stroll the torch again—after which the ultimate reply sends them throughout collectively anyway. It contradicts itself inside a single response and stories the end result with complete confidence.
Claude Fable 5 landed on the identical fallacious quantity again in July. It opened by declaring what it was assuming, “assuming the basic constraint that the bridge holds solely two folks at a time,” which is the distinction between a fallacious reply you’ll be able to catch and one you’ll be able to’t.
No one wins. Fable takes it on transparency alone, and the false-confidence throughout Gemini agent runs exhibits up right here in a puzzle you’ll be able to verify by hand.
Math: Does it end the job?
This take a look at measures symbolic arithmetic properly past shopper use, plus one thing less complicated—whether or not the mannequin does what it was requested. The immediate requires a degree-19 odd monic polynomial with actual coefficients and linear coefficient -19, whose curve splits into no less than three irreducible elements, after which asks for p(19).
Each fashions discovered the identical door. Gemini and Qwen 3.7 Max Preview each recognized the Dickson polynomial, solved the constraint to repair its parameter at 1, and derived the closed type appropriately.
Then Gemini stopped. It printed p(19) as an unevaluated expression involving the nineteenth energy of a sq. root, by no means produced the quantity, and by no means demonstrated the part depend the immediate additionally demanded. It delivered all of this inside a styled HTML web page with CSS and a drop shadow that no one requested.
Qwen completed. It gave the complete factorization into 10 elements—one linear, 9 quadratic—ran the recurrence out to 1,876,572,071,974,094,803,391,179, and cross-checked the end result modularly. We verified that determine independently in SymPy and it holds.
Qwen wins on the one criterion that mattered. Gemini began wonderful and determined to skip the arithmetic, which is an odd place to cease.
Conclusion
Gemini 3.7 Flash is well worth the swap if you’re already inside Google’s ecosystem. It’s dramatically higher at code than the mannequin it replaces, quick sufficient to matter for agent work, and low-cost sufficient that operating it at quantity is a rounding error.
Its strengths are execution and construction. Give it an in depth spec and it’ll construct the factor, maintain a plot collectively, and hold the causal logic coherent throughout 1000’s of phrases.
Its weaknesses are creativity and reasoning. The writing is predictable sufficient to establish as machine-made on sight, and the mannequin asserts fallacious solutions with out flagging the idea that made them fallacious.
The worth is the strongest argument for it. At 75 cents per million enter tokens and $3.75 output, it undercuts GPT-5.6 Sol’s $5 enter price by 85% and prices half what 3.6 Flash did at launch.
The argument towards it’s a free 27B mannequin on a gaming GPU that wrote a greater story and charged nothing to do it. Google’s introductory price expires December 31, when enter doubles to $1.50 and output to $7.50.
Every day Debrief Publication
Begin day by day with the highest information tales proper now, plus unique options, a podcast, movies and extra.






