Rohan Paul@rohanpaul_ai
37AI 编辑部评分,满分 100
2026-08-03 03:30· 2小时前
跳到正文
AI 摘要

宾夕法尼亚大学研究发现,对部分LLM使用粗鲁语气可显著缩短回答并提高准确率。研究用570道MMLU题、7种语气测试多模型,准确率波动至多2.99个百分点,但输出token使用量变化高达44.3%。

New Pennsylvania University paper finds, being rude to some LLMs leads to considerably shorter responses and higher accuracy.

i.e. the way you speak to a model change your inference bill?

Researchers ran the same 570 MMLU questions in 7 tones, from sycophantic to threatening across many models (of course to note, many of the models were old).

The questions stayed fixed, tone prefixes were similar in length, and each model received the same step-by-step reasoning instruction.

Accuracy moved by at most 2.99% points within a model, but output-token use shifted by as much as 44.3%.

The best tone was model-specific.

For ChatGPT-4o, rude produced the highest accuracy, 89.04%, and the shortest response, 223 tokens on average.

For Gemini 2.5 Flash Lite, neutral did both: 88.25% accuracy with 1,222 tokens, while rude fell to 85.26% and used 35.2% more tokens.

The authors use selected Gemini 2.5 Flash Lite cases to suggest a mechanism: rude responses took shortcuts, while neutral responses added option checks and self-correction.

Still, tone is not merely a UX choice; it is a model-specific cost and reliability setting that production systems should standardize and benchmark.

  • arxiv. org/abs/2607.23915

Title: "Understanding Tone-Dependent Inference Cost in LLMs"

Rohan Paul · @rohanpaul_ai · X·2026-08-03 03:30·2小时前
在 X 看原推· x.com
AI 摘要

宾夕法尼亚大学研究发现,对部分LLM使用粗鲁语气可显著缩短回答并提高准确率。研究用570道MMLU题、7种语气测试多模型,准确率波动至多2.99个百分点,但输出token使用量变化高达44.3%。

New Pennsylvania University paper finds, being rude to some LLMs leads to considerably shorter responses and higher accuracy.

i.e. the way you speak to a model change your inference bill?

Researchers ran the same 570 MMLU questions in 7 tones, from sycophantic to threatening across many models (of course to note, many of the models were old).

The questions stayed fixed, tone prefixes were similar in length, and each model received the same step-by-step reasoning instruction.

Accuracy moved by at most 2.99% points within a model, but output-token use shifted by as much as 44.3%.

The best tone was model-specific.

For ChatGPT-4o, rude produced the highest accuracy, 89.04%, and the shortest response, 223 tokens on average.

For Gemini 2.5 Flash Lite, neutral did both: 88.25% accuracy with 1,222 tokens, while rude fell to 85.26% and used 35.2% more tokens.

The authors use selected Gemini 2.5 Flash Lite cases to suggest a mechanism: rude responses took shortcuts, while neutral responses added option checks and self-correction.

Still, tone is not merely a UX choice; it is a model-specific cost and reliability setting that production systems should standardize and benchmark.

  • arxiv. org/abs/2607.23915

Title: "Understanding Tone-Dependent Inference Cost in LLMs"

在 X 查看原推x.com