Google and Meta released their latest frontier AI models within hours of each other on Wednesday. Google shipped Gemini 3.8 Flash alongside a cybersecurity variant, while Meta pushed out Muse Spark 1.3.
The two launches invite a direct comparison. Independent testing by Artificial Analysis splits the result. Meta leads on agentic knowledge work and scientific reasoning, while Google holds an edge in factual recall and terminal coding.
Two Frontier Releases Land on the Same Day
Gemini 3.8 Flash is Google’s third Flash release in six weeks. The model costs $0.75 per 1 million input tokens and $3.75 per 1 million output tokens.
Follow us on X to get the latest news as it happens
That introductory rate runs through December 31, 2026. Prices then double to $1.50 and $7.50 per 1 million tokens.
Google paired the general model with Gemini 3.8 Flash Cyber. The model is available to a set of trusted defenders. It scored 86.2% on CyberGym, a benchmark for finding vulnerabilities.
The cyber variant also reached 47.2% on CWE-Bench, a patching benchmark. Google said the model produced 2.6 times more correct patches for Chrome vulnerabilities than larger commercial models.
Access sits behind the Fairwind Program, which limits the model to government authorities and critical infrastructure operators. OpenAI drew a similar boundary a day earlier around Astra, the first model it rated at a critical cybersecurity threshold.
Meanwhile, Meta rolled out Muse Spark 1.3 via Muse Code and the Meta Model API. Company engineers measured roughly 20% fewer tool calls than version 1.2.
Gemini 3.8 Flash vs Muse Spark 1.3: Independent Benchmarks Split the Result
Artificial Analysis tested the models, and the results show how they rank. Muse Spark 1.3 in max mode scored 1,754 Elo on GDPval-AA v2. Gemini 3.8 Flash (high) returned 1,545.
Meta also led the Sierra Research banking agent test (52.4% to 44.9%) and CritPt physics reasoning. Gemini 3.8 Flash led Terminal-Bench 2.1 at 87.6%, AA-LCR long context at 81%, and AA-Omniscience accuracy at 55%.
Gemini 3.8 Flash posted the highest GPQA Diamond score among the models tested, at 95%. The two finished within a point of each other on Humanity’s Last Exam.
Meta’s top scorer does not ship today. The company said max reasoning will arrive once further safety testing is complete, leaving xhigh as the available variant.
That version scored 61 on the Artificial Analysis Intelligence Index, four points above Muse Spark 1.2. It trails Claude Fable 5.1 at 66 and Claude Opus 5 at 63.
More releases are already queued. Elon Musk has said Grok 4.7 arrives shortly, which would place four frontier launches inside a fortnight.
Subscribe to our YouTube channel to watch leaders and journalists provide expert insights









