Technology

Google released Gemini 4 Argon to a handful of security firms. It lost two of its own four coding benchmarks

Google calls Argon its largest and most powerful model, matching OpenAI’s Astra and Anthropic’s Claude Opus on coding and security tests. It trailed rivals on two of four published coding benchmarks.

Google released Gemini 4 Argon to a handful of security firms. It lost two of its own four coding benchmarks

Google has released a new Gemini model to compete with Anthropic and OpenAI. Gemini 4 Argon, built for complex work, is the company's largest and most powerful model, it says.

Who can use it

A Google spokesperson said Argon matched the capability of OpenAI's Astra and Anthropic's Claude Opus on several important coding and cybersecurity tests.

Initially, only selected cybersecurity firms will be able to use it. The spokesperson gave no timeline for a general release.

Google says Argon beat Astra and Opus on several industry benchmarks in its own testing — but not on all of them. Of four published coding tests, Argon came behind its rivals on two.

Google had decided months ago to release the model. It delayed amid concerns about model safety and the resignation of several senior figures, including the chief executive of its AI lab DeepMind.

What it means in Bangladesh

Three things in that short account deserve to be separated, because they point in different directions.

The first is the honesty, and it is unusual enough to credit. A company that publishes four benchmarks and loses two of them is telling you something real. The industry norm is to publish the charts you win; Google published the ones it did not. Take the claim of parity with Astra and Opus as what it is — a vendor's own testing — but take the two losses as the more informative number.

The second is the release pattern, and it is the part that matters here. A frontier model going first to a handful of cybersecurity firms, with no date for anyone else, is a deliberate choice about who gets to probe a new system before the public does. The reasoning is defensible — a model strong at cybersecurity work is a model strong at the offensive half of it too. The consequence is that the capability exists, is being used, and is unexaminable from outside that list. No Bangladeshi institution is on it, and none will be.

The third is the delay, and this is where the week's stories join up. A company postponed a finished model because of safety concerns and leadership resignations at its AI lab — the same pattern of internal argument visible at its two main rivals this month. Whatever one concludes about the models, the people building them are not behaving like people who are confident.

What Bangladesh can act on is narrow. When a model arrives in a product used here — in Search, in Workspace, in a bank's fraud system — the version that arrives will have been shaped by a safety argument settled in California, and the local deployment will not restate it. Asking a vendor which model version a product runs, and what evaluation it passed, is a reasonable procurement question that almost nobody asks.

The jailbreak evidence on a rival model is in the Kimi review.

Source: প্রথম আলো

Written by

Zayed

Zayed writes Tech BD’s artificial intelligence coverage — model releases, AI safety research, and the regulation forming around them. His interest is less in what a system can demonstrate than in what it changes for someone using it in Bangladesh. He writes in both English and Bangla.