杨澜访谈录
刚刚,Gemini 3.6 Flash 正式发布,但网友笑得更大声了_我的网站

A | 如果你从去年就开始用 Gemini,那感觉就像看着自家兄弟慢慢得了阿尔茨海默症。 Islamabad, Oct 21 (UNI) The Financial Action Task Force (FATF) on Friday removed Pakistan from a list of countries under “increased monitoring”, also known as the “grey list”.
Addressing the press conference at the conclusion of the its plenary, President Raja Kumar noted that Pakistan had been on the grey list since 2018.
“It has two concurrent action plans. After a lot of work by Pakistan authorities, they have largely addressed all of the action plan items,” he said.
He stated that the task force had conducted an onsite visit at the end of August. The onsite team verified that there is a high-level of commitment from the Pakistani leadership, sustainability of reforms and commitment to make improvements in the future, he said.
“As a result of these action plans, Pakistan has made significant improvements to strengthen the effectiveness of this framework for combatting terrorism financing.”
Foreign Minister Bilawal Bhutto-Zardari was quick to congratulate the country on the development moments before the FATF press conference began, according to Dawn reports.
Prime Minister Shehbaz Sharif said Pakistan’s exiting the FATF grey list is a “vindication of our determined and sustained efforts over the years”, according to reports.
He congratulated the civil and military leadership as well as all institutions whose hard work led to today’s success. “Aap sab ko bohat bohat Mubarak,” the premier said, following up with a text smiley.
In a follow-up tweet, PM Shehbaz particularly commended the role and efforts of Foreign Minister Bilawal Bhutto, Army Chief General Qamar Javed Bajwa and their teams and all political parties for putting up a united front to get Pakistan out of the grey list.
Pakistan was included among jurisdictions under increased monitoring list in June 2018 for deficiencies in its legal, financial, regulatory, investigations, prosecution, judicial and non-government sector to fight money laundering and combat terror financing considered serious threat to global financial system.
Islamabad made high-level political commitments to address these deficiencies under a 27-point action plan. But later the number of action points was enhanced to 34.
The country had since been vigorously working with FATF and its affiliates to strengthen its legal and financial systems against money laundering and terror financing to meet international standards in line with 40-recommendations of the FATF.
A 15-member joint delegation of the FATF and its Sydney-based regional affiliate — Asia Pacific Group — paid an onsite visit to Pakistan from Aug 29 to Sept 2 to verify the country’s compliance with the 34-point action plan committed with the FATF.
The authorities that had kept the countrywide visit of the delegation low profile later termed it “a smooth and successful visit”.
The delegation had detailed discussions with relevant agencies pursuant to the authorisation of onsite Visit by FATF Plenary in June 2022.
According to the Foreign Office, the focus of the visit was to validate on ground Pakistan’s high-level commitment and sustainability of reforms in AML/CFT regime and [it] looked forward to logical conclusion to the evaluation process.
Pakistan believed that as a result of strenuous and consistent efforts over the past four years, it has not only achieved a high degree of technical compliance with FATF standards but also ensured high level of effectiveness through implementation of two comprehensive FATF action plans.
In June this year, FATF had found Pakistan “compliant or largely compliant” on all the 34 points and had decided to field an onsite mission to verify it on ground before formally announcing the country’s exit from the grey list that finally took place in August and September.
In terms of technical compliance with FATF standards, Pakistan has been rated by APG as “compliant or largely compliant” in 38 out of 40 FATF recommendations in August this year, which placed the country among the top compliant countries in the world.
UNI GNK。
这是最近网友对 Gemini 的一句调侃。
按理说,一家公司发新模型,通常是来打脸这种调侃的。可就在刚刚,Google 一口气发了三个新模型之后,网友非但没收回这句话,反而笑得更大声了。多少有种他们都不看好你,可偏偏你最不争气的即视感。
三个模型分别是 3.6 Flash、3.5 Flash-Lite,还有一个专搞网络安全的 3.5 Flash Cyber。官方博客的措辞也相当眼熟,「更高效」「更聪明」「为大规模 AI agent 而生」,一句不落。
只不过,实际体感却冰火两重天。
主角登场,3.6 Flash 到底强在哪
先说头牌,3.6 Flash。官方给的定位就俩字——「主力」(workhorse),干活模型。
3.5 Flash 是今年 5 月 I/O 大会上发的,这次算是小版本迭代。
最大的卖点是省 token。
按 Artificial Analysis Index 的数据,3.6 Flash 比 3.5 Flash 少用 17% 的输出 token,在 DeepSWE(Datacurve 出的基准)这类场景上甚至能省到 65%。官方还说它跑多步任务时,推理步数和工具调用都更少。也就是说,干同样的活,废话更少,绕路更少,账单更薄。
价格也确实降了。输入 1.5 美元/百万 token,输出 7.5 美元/百万 token——注意,上一代 3.5 Flash 的输出是 9 美元,这波直接砍到 7.5。又快又省,单个 agent 任务的成本就压下来了。

B | 基准测试上,官方摆出了一整排数据。 代码能力是重点。DeepSWE 从 37% 提升到 49%,官方的说法是多余的代码改动更少、执行循环也更短,生成的代码更贴近生产环境的要求。机器学习研究方向的 MLE Bench 提升更明显,从 49.7% 拉到了 63.9%。 计算机操作(computer use)能力也有长进,OSWorld-Verified 从 78.4% 升到 83%。

C | 并且,「计算机操作」(computer use)也成了Gemini API 和企业版的内置工具,主打一个开箱即用。 知识工作方面,GDPval-AA v2 从 1349 涨到 1421。Hebbia、Harvey 等客户反馈,它在文档解析、图表数据分析、报告起草这类多模态任务上表现尤其突出。Figma、JetBrains 也都出面背书了一番。 哦对了,还有个不起眼但挺实在的更新:知识截止日期终于从 2025 年 1 月推进到了 2026 年 3 月。老黄历总算翻篇了。 安全这块,官方说 3.6 Flash 上了升级版的 Frontier Safety 防护,重点盯着 CBRN(化学、生物、放射性、核)和网络攻击这两类滥用,抗越狱能力更强了,同时又尽量不误伤正常需求、少乱拒绝。 ·以小博大,便宜大碗的 3.5 Flash-Lite 也能更换推理档位了 第二个是 3.5 Flash-Lite,定位比 3.6 Flash 还低一档,主打「快」和「便宜」,专治高吞吐、低延迟的任务,比如 agent 搜索、文档处理。 它是 3.5 系列里最快的,按 Artificial Analysis 的测量是每秒吐 350 个 token。价格更是白菜——输入 0.3 美元、输出 2.5 美元每百万 token。

D | 官方说质量比 3 月发的 3.1 Flash-Lite 好一大截。 有个设计挺灵活:它支持调「推理档位」。低配活儿就开最低档,图个快和省;碰上要多步拆解的子任务,就把思考档位往上拨。电脑操作这次也做成了内置工具。 跟前代比,提升相当明显:代码和 agent 任务的 Terminal-Bench 2.1,从 31% 干到 54%;长上下文的 GDM-MRCR v2,从 60.1% 提到 72.2%;真实任务执行的 GDPval-AA v2,更是从 642 飙到 1140。 最有意思的是,这个「小弟」在不少 agent 和代码任务上,居然反超了辈分更高的 3 Flash。

E | 比如 SWE-Bench Pro(54.2% vs 49.6%)、OSWorld-Verified(74.0% vs 65.1%)。以小博大,属实有点东西——等于说跑 3 Flash 的活儿,现在有了个更快更强的平替。 官方还举了几个用法:从海量电商数据里抽产品特征、给 3.6 Flash 当副手一口气生成 25 个网页设计方案、批量翻译总结收据、边玩边迭代做小游戏。

F | Ashler、Palo Alto Networks、Ramp 这几家客户也来夸了它「速度、智能、成本三合一」。 最后一个 3.5 Flash Cyber,画风突变,专门用来找和修代码里的安全漏洞。 Google 的逻辑挺有意思:现在 AI 找漏洞的速度,已经比现有系统修漏洞的速度快了。既然漏洞越堆越多,那干脆用便宜高效的 Flash 来批量补锅。 这个模型是在 3.5 Flash 基础上微调出来的,搭配自家的 CodeMender 工具用。CodeMender 里跑的是好几个 Flash Cyber agent,协同工作、最后汇总成一份报告。 在业内有名的 CyberGym 基准上,官方称它摸到了第一梯队的水平,还比更大的模型更省 token。

G | 不过这模型我们暂时也用不上。 考虑到「找漏洞」这种能力是把双刃剑,浓眉大眼的 Google 也学 Anthropic 玩起了安全叙事。把它锁得死死的——只对「可信合作伙伴」限量开放,走内测。目的是给一线防守方争取时间,赶在漏洞被人利用前先修掉,同时防着有人拿它去干坏事。 说好的 3.5 Pro 呢?如来 看到这儿你可能会问:发了一堆 Flash,那真正的旗舰 3.5 Pro 去哪了? 官方口径是「正在和合作伙伴测试,准备好就上」。但这套说辞背后的故事,可能没那么体面。 据彭博社等媒体报道,3.5 Pro 的代码生成能力一直没达到内部预期。Google 6 月下旬还专门更新了一版训练数据来补代码这块短板,结果成绩依旧没起色。

H | 更有传闻说,Google 干脆推翻了原来的训练方案,从零开始重训。 而 Google AI 宣传委员 Logan Kilpatrick 更是索性在舆论上直接跳过 3.5 Pro,将预告的重点放在了Gemini 4,声称他们已经启动了史上最雄心勃勃的 Gemini 4 预训练,还放话说进展让人兴奋。

I | 听着虽然挺提振人心,但把它放在「旗舰跳票」的背景下看,多少有点转移视线、画饼续命的味道了。

J | 除了这些内幕,光看今天刚刚发布的 Flash 模型本身,网友们也并不全然买账。 第三方评测机构 Artificial Analysis 表示:两个新模型确实把单任务耗时砍半、token 效率也提升了,Flash-Lite 的智能指数涨了 11 分——但 3.6 Flash 的智能水平,相比 3.5 Flash 基本原地踏步。

K | 价格没有便宜到足以忽略能力差距,能力也没有高到能够解释它的成本。 X 网友吐槽道:3.6 Flash 在 Artificial Analysis 上跟 3.5 Flash 拿了一模一样的分,还不如 Meta Spark 1.1、GLM-5.2、5.6 Luna、Sonnet 5、Grok 4.5、5.6 Terra 这一票对手,直呼简直烂透了(阴阳怪气 doge)。

L | 吐槽的还不止一个。 网友 Angel 则直接从性价比入手:Gemini 3.6 Flash 的使用成本比 GPT-5.6 Sol medium 还高,可智能程度反而更低。又贵又笨,这组合属实有点尴尬——毕竟 Flash 系列立身的根本就是「性价比」,结果被人当场指出性价比不如对手,等于自砸招牌。

M | 作为对比,OpenAI 的 Tibo 刚刚也发话了:由于周活跃用户达到 1000 万,新的一天,Codex 和 ChatGPT Work 的付费用户迎来新一轮额度重置,尽情享用。 这对比,这落差,还有人记得大明湖畔的 Google Gemini 3? 实测派也来补刀。网友 Balder 更是直接开团: 这三个模型性能全比之前的 3.5 Flash 差。他甩出自己的测试吐槽,说问题包括但不限于——基础解码就有毛病、中文选词经常很离谱;生成的图片视频质量极差、跟指令对不上还限额严重;经常记忆混乱、调用完全错误的工具。 而且他还上了个阴谋论,认为 Google 之所以这么搞,是为了把算力腾出来卖给 Anthropic,还阴阳了一句这就是皮查伊想要的「Cloud First」。 此外,X 网友 Conor Dart 用 Google 自家的 Antigravity 跑了个 AI 生成游戏的测试,结果没有出乎意料,木头纹理更差了,大理石看着差不多但还是不能用。

N | 他直言这又是一次翻车,表示自己还是等 Gemini 3.5/3.6 Pro 吧。 简言之,你别问新模型强不强,你就说 Flash 不 Flash 就完事了。一边是官方更高效、更便宜、更省 token的漂亮账本,一边是大多数网友的当场差评,以及模型背后的旗舰跳票、大牛出走、市值缩水等一连串糟心事。

o | 这很难不让人好奇是不是 Google 这次真点歪了科技树,光顾着省成本忘了提智商,只能先靠几个 Flash 稳住场子? 反正 Gemini 4 的预训练都开跑了。这一把要是再翻车,那句阿尔茨海默的玩笑,可就真要成谶语了。
Current article:http://6n91ve9.senchuoshaozaodidia.shop/news/20260826_10950.html
Published on:20:52:02
