Mira Murati 的 Thinking Machines 使 Bridgewater 的专家判断可训练,错误率降低 29.8%

Rohan Paul · @rohanpaul_ai · X·2026-07-03 13:11·60天前
AI 导读

Mira Murati 的 Thinking Machines 与 Bridgewater 合作,利用其金融知识微调模型以复制专家判断。任务为筛选投资者应阅读的金融文章、报告、央行文件及邮件。朴素提示词下准确率仅 46%–50%,专家提示词升至 74%–78%。训练采用交错批次、CISPO 损失(MiniMax 提出)及基于更优教师检查点的在线策略蒸馏。结果击败最佳前沿模型,错误率降低 29.8%,推理成本下降 13.8 倍。非专家标签因任务依赖品味而失效,Bridgewater 将模型争议案例送回专家复核并清洗标签。

Rohan Paul@rohanpaul_ai
54AI 编辑部评分,满分 100

Mira Murati 的 Thinking Machines 使 Bridgewater 的专家判断可训练,错误率降低 29.8%

2026-07-03 13:11· 60天前
AI 导读

Mira Murati 的 Thinking Machines 与 Bridgewater 合作,利用其金融知识微调模型以复制专家判断。任务为筛选投资者应阅读的金融文章、报告、央行文件及邮件。朴素提示词下准确率仅 46%–50%,专家提示词升至 74%–78%。训练采用交错批次、CISPO 损失(MiniMax 提出)及基于更优教师检查点的在线策略蒸馏。结果击败最佳前沿模型,错误率降低 29.8%,推理成本下降 13.8 倍。非专家标签因任务依赖品味而失效,Bridgewater 将模型争议案例送回专家复核并清洗标签。

Mira Murati's Thinking Machines made Bridgewater’s private expert judgment trainable, beating frontier models with 29.8% fewer errors.

With naive prompts, all tested models sit around coin-flip accuracy, roughly 46% to 50%. Expert prompts lift them sharply, reaching about 74% to 78% average accuracy.

The workflow was filtering finance articles, reports, central-bank documents, and emails to decide what investors should read.

This is a serious signal for enterprise AI, that bringing private judgment in the loop beats general intelligence.

The problem was not reading finance documents, because frontier LLMs can already read them.

The harder task was deciding which facts deserve attention inside an investor’s workflow.

A tariff headline can move markets, while another geopolitical headline may add no signal.

The breakthrough came from replacing written rules with high-quality labels from expert investors.

Non-expert labels failed because the task depends on taste, not surface financial language.

Bridgewater cleaned those labels by sending model-disputed cases back to experts for review.

The model then learned patterns that experts could recognize, but could not fully verbalize.

Training used interleaved batches, CISPO loss, and on-policy distillation from stronger teacher checkpoints.

Interleaving helped the model share judgment across tasks without blending them into noise.

CISPO controlled policy updates, so learning stayed aggressive without drifting into brittle shortcuts.

(CISPO is a new reinforcement-learning loss that caps how strongly each generated token can update the model, improving training stability while keeping useful rare tokens active. It was initially proposed by MiniMax team in 2025)

On-policy distillation penalized moves away from better teachers, then promoted stronger checkpoints.

The result beat the best frontier model, with 29.8% fewer mistakes and 13.8x lower inference cost.

Mira MuratiBridgewater used their unique financial knowledge and partnered with us on @tinkerapi to fine-tune a model that helps their analysts focus on what's important. ...

来源:Rohan Paul· x.com