OpenAI 正在将 ChatGPT 的默认模型替换为 GPT-5.5 Instant,该模型在医学、法律和金融等高危话题上的模型幻觉减少了 52.5%,同时在数学、科学和视觉推理方面取得了显著的基准测试成绩提升。
一项新的“记忆来源”功能现在可以向用户展示哪些个人背景信息——包括过往聊天记录、已保存的提醒事项或上传的文件——影响了某次具体回复,并且用户能够修正或删除其中的单条记录。
GPT-5.5 Instant 将立即向所有 ChatGPT 用户推送,不过通过过往聊天记录、文件和 Gmail 实现的高级个性化功能最初仅限 Plus 和 Pro 订阅用户使用,更广泛的开放将在未来几周内实现。
OpenAI 正在将 ChatGPT 的默认模型更换为 GPT-5.5 Instant。此次更新减少了模型幻觉并收紧了回复内容,同时一项名为“记忆来源”的新功能会向用户展示哪些存储的背景信息影响了某次具体回复。
GPT-5.5 Instant 取代了 GPT-5.3 Instant,并且也通过 API 以“chat-latest”的名称提供。在 OpenAI 的内部测试中,GPT-5.5 Instant 在医学、法律和金融领域的高风险提示词上,相比前代模型产生的幻觉性陈述减少了 52.5%。OpenAI 声称,在用户之前标记为存在事实错误的棘手对话中,不准确陈述的比例下降了 37.3%。
OpenAI 以一个代数问题为例。用户上传了一张手写方程式的照片,其中包含一个计算错误。GPT-5.3 Instant 最初认同了该解法,随后注意到 x=3 不成立,但错误地得出结论认为没有实数解。GPT-5.5 Instant 最初也认同了用户的数学计算,但随后发现了用户在重新排列方程时的错误,并解出了修正后的二次方程。
基准测试成绩也反映了类似情况。在竞争性数学考试 AIME 2025 上,准确率从 65.4% 跃升至 81.2%。在测试博士级科学推理能力的 GPQA 上,得分从 78.5% 攀升至 85.6%。在用于解读和推理科学图表的基准测试 CharXiv 上,得分从 75.0% 提升至 81.6%。
MMMU-Pro(衡量模型处理跨文本与图像专家级问题能力的基准)从 69.2% 提升至 76.0%。OmniDocBench(一项从复杂文档中提取结构化数据的测试)的错误率从 14.6% 降至 12.5%。
基准测试 基准测试描述 指标 GPT-5.3 Instant GPT-5.5 Instant CharXiv-reasoning 科学图表推理 准确率 75.0% 81.6% MMMU-Pro 专家级多模态推理 准确率 69.2% 76.0% OmniDocBench 文档解析 平均错误率(越低越好) 14.6% 12.5% GPQA 博士级科学 准确率 78.5% 85.6% AIME 2025 竞赛数学 准确率 65.4% 81.2%
更精炼的回答与更智能的个性化
OpenAI 还着力于精简内容。该公司表示,回答更简短但不失实质;模型减少了不必要的追问,去掉了多余的 emoji,并避免了繁重的格式排版。OpenAI 写道:“它能够传递相同的信息,且通常比之前的模型更具实用性,同时减少了导致回答过长的冗词和过度格式化。”
当相关功能开启时,该模型还能更好地利用来自过往对话、上传文件以及已关联 Gmail 账户的上下文。据报道,GPT-5.5 Instant 能更准确地判断何时额外的个性化设置确实有助于回答,并且能更快地检索之前的对话。
OpenAI 还在所有 ChatGPT 模型中推出记忆来源功能。当回答依赖于存储的上下文时,用户现在可以看到使用了哪些信息,无论是保存的笔记还是过去的对话。这些条目可以被标记为相关或不相关,也可以进行编辑或删除。
但 OpenAI 表示,记忆来源并不总会显示回答背后的所有因素。例如,只有模型搜索到的部分对话会作为来源显示。该公司计划随着时间的推移让视图更加完整。当对话被分享时,记忆来源不会被一并传递,而临时对话既不会读取记忆,也不会更新记忆。
按套餐分阶段推出
OpenAI 表示,GPT-5.5 Instant 将立即向所有 ChatGPT 用户推出。付费用户仍可通过模型设置访问 GPT-5.3 Instant,该版本将在三个月后退役。
基于过往聊天记录、文件及 Gmail 的增强个性化功能,将首先面向网页端的 Plus 和 Pro 用户推出,移动端即将上线。免费版、Go 版、Business 版和 Enterprise 版用户预计将在未来几周内获得访问权限。记忆源将首先面向所有个人版网页端用户推出,移动端随后跟进。部分个性化功能可能并非在所有地区都可用。
OpenAI 近期推出了 GPT-5.5 Thinking 作为更高级别的模型,而 GPT-5.5 Instant 则作为 ChatGPT 的日常默认模型。Thinking 版本依然更强大:据报道,它在网络安全任务上的表现与 Claude Mythos 相当,并且取代了专门的 Codex 编程模型。
OpenAI is replacing ChatGPT's default model with GPT-5.5 Instant, which shows 52.5% fewer hallucinations on high-risk topics like medicine, law, and finance, along with strong benchmark gains in math, science, and visual reasoning.
A new "memory sources" feature now shows users which personal context—past chats, saved reminders, or uploaded files—informed a given response, with the ability to correct or remove individual entries.
GPT-5.5 Instant is rolling out to all ChatGPT users right away, though advanced personalization via past chats, files, and Gmail is initially limited to Plus and Pro subscribers, with wider availability coming in the following weeks.
OpenAI is swapping out ChatGPT's default model for GPT-5.5 Instant. The update reduces hallucinations and tightens responses, while a new feature called "memory sources" shows users which stored context shaped a given reply.
GPT-5.5 Instant replaces GPT-5.3 Instant and is also available through the API as "chat-latest." In OpenAI's internal testing, GPT-5.5 Instant produced 52.5 percent fewer hallucinated claims than its predecessor on high-risk prompts in medicine, law, and finance. On tough conversations users had previously flagged for factual errors, inaccurate claims dropped by 37.3 percent, OpenAI claims.
OpenAI offers an algebra problem as an example. A user uploaded a photo of a handwritten equation with a calculation mistake. GPT-5.3 Instant initially agreed with the solution, then noticed that x=3 didn't work but wrongly concluded there was no real solution. GPT-5.5 Instant also agreed with the user's math at first, but then caught the error in how the user had rearranged the equation and solved the corrected quadratic.
Benchmark scores tell a similar story. On AIME 2025, a competitive math exam, accuracy jumped from 65.4 to 81.2 percent. GPQA, which tests PhD-level science reasoning, climbed from 78.5 to 85.6 percent. CharXiv, a benchmark for interpreting and reasoning about scientific charts, went from 75.0 to 81.6 percent.
MMMU-Pro, which measures how well models handle expert-level questions across text and images, rose from 69.2 to 76.0 percent. The error rate on OmniDocBench, a test for extracting structured data from complex documents, dropped from 14.6 to 12.5 percent.
Benchmark Benchmark Description Metric GPT-5.3 Instant GPT-5.5 Instant CharXiv-reasoning Scientific Chart Reasoning Accuracy 75,0 % 81,6 % MMMU-Pro Expert Multimodal Reasoning Accuracy 69,2 % 76,0 % OmniDocBench Document Parsing Average error rate (lower = better) 14,6 % 12,5 % GPQA PhD-Level Science Accuracy 78,5 % 85,6 % AIME 2025 Competition Math Accuracy 65,4 % 81,2 %
Tighter answers and smarter personalization
OpenAI also focused on cutting fluff. Answers are shorter without losing substance; the model asks fewer unnecessary follow-ups, drops superfluous emojis, and skips heavy formatting, the company says. "It can deliver the same information, often with more utility than previous models, while reducing the verbosity and overformatting that can make responses too long", OpenAI writes.
The model also makes better use of context from past chats, uploaded files, and connected Gmail accounts when those features are turned on. GPT-5.5 Instant is reportedly better at judging when extra personalization actually helps a response, and it searches previous conversations faster.
OpenAI is also rolling out memory sources across all ChatGPT models. When a reply draws on stored context, users can now see which information was used, whether that's a saved note or a past chat. Entries can be flagged as relevant or irrelevant, edited, or deleted.
But memory sources won't always show every factor behind a response, OpenAI says. Only some chats the model searches will appear as sources, for instance. The company plans to make the view more complete over time. Memory sources aren't passed along when a chat is shared, and temporary chats neither read from nor update memory.
Staggered rollout across plans
OpenAI says GPT-5.5 Instant is rolling out to all ChatGPT users right away. Paying users can still access GPT-5.3 Instant through model settings for another three months before it's retired.
Enhanced personalization based on past chats, files, and Gmail is launching first for Plus and Pro users on the web, with mobile coming soon. Free, Go, Business, and Enterprise plans are expected to get access over the coming weeks. Memory sources will roll out to all consumer plans on the web first, with mobile to follow. Some personalization features may not be available in every region.
OpenAI recently introduced GPT-5.5 Thinking as the higher-tier model, while GPT-5.5 Instant serves as ChatGPT's everyday default. The Thinking version is still more powerful: on cybersecurity tasks it reportedly matches Claude Mythos, and it replaces the specialized Codex coding models.