Simon Willison 博客
35AI 编辑部评分,满分 100

不要分类,去幻觉:用向量嵌入为博客打标签

2026-08-15 05:54· 21分钟前· Simon Willison
AI 导读

Simon Willison 分享了一个为旧博客内容打标签的新思路:让模型先自由想象可能适合的标签,再用向量嵌入在现有 1,856 个标签中寻找最接近的具体标签。Doug Turnbull 提出了这一方案,其示例提示词会引导模型生成从未见过的新分类,再通过语义匹配映射到已有标签体系。

Simon Willison’s Weblog

Sponsored by: WorkOS — auth.md by WorkOS: agents register users, no sign-up form. Try it!

14th August 2026 - Link Blog

Don't classify. Hallucinate! I still have quite a bit of older content on my blog that I never got round to tagging. My blog has 1,856 tags - likely too many to feed to an LLM in one go and say "which of these tags match the following content".

Doug Turnbull has a neat solution. Tell the model to output tags without any details of the existing vocabulary, then use vector embeddings against the existing corpus to find the concrete tags that are closest to the ones the model imagined might fit!

His example prompt suggests including an example of the shape of your tags to help the model make a more useful guess:

Your task is to create novel, never seen before, furniture, home goods, or hardware classification that best fit a search query.

Product classifications might look like:

Furniture / Living Room Furniture / Coffee Tables & End Tables / Coffee Tables
Décor & Pillows / Decorative Pillows & Blankets / Throw Pillows
Furniture / Bedroom Furniture / Dressers & Chests
Kitchen & Tabletop / Kitchen Organization / Food Storage & Canisters
School Furniture and Supplies / School Furniture / School Chairs & Seating / Stackable Chairs
Baby & Kids / Toddler & Kids Bedroom Furniture / Kids Beds

Here's the query to generate classifications for:

brown coffee table

14th August 2026

来源:Simon Willison 博客 · simonwillison.net

不要分类,去幻觉:用向量嵌入为博客打标签

Simon Willison 博客·2026-08-15 05:54·21分钟前·Simon Willison
AI 导读

Simon Willison 分享了一个为旧博客内容打标签的新思路:让模型先自由想象可能适合的标签,再用向量嵌入在现有 1,856 个标签中寻找最接近的具体标签。Doug Turnbull 提出了这一方案,其示例提示词会引导模型生成从未见过的新分类,再通过语义匹配映射到已有标签体系。

原文 · 保持原样,未翻译

Simon Willison’s Weblog

Sponsored by: WorkOS — auth.md by WorkOS: agents register users, no sign-up form. Try it!

14th August 2026 - Link Blog

Don't classify. Hallucinate! I still have quite a bit of older content on my blog that I never got round to tagging. My blog has 1,856 tags - likely too many to feed to an LLM in one go and say "which of these tags match the following content".

Doug Turnbull has a neat solution. Tell the model to output tags without any details of the existing vocabulary, then use vector embeddings against the existing corpus to find the concrete tags that are closest to the ones the model imagined might fit!

His example prompt suggests including an example of the shape of your tags to help the model make a more useful guess:

Your task is to create novel, never seen before, furniture, home goods, or hardware classification that best fit a search query.

Product classifications might look like:

Furniture / Living Room Furniture / Coffee Tables & End Tables / Coffee Tables
Décor & Pillows / Decorative Pillows & Blankets / Throw Pillows
Furniture / Bedroom Furniture / Dressers & Chests
Kitchen & Tabletop / Kitchen Organization / Food Storage & Canisters
School Furniture and Supplies / School Furniture / School Chairs & Seating / Stackable Chairs
Baby & Kids / Toddler & Kids Bedroom Furniture / Kids Beds

Here's the query to generate classifications for:

brown coffee table

14th August 2026

来源:Simon Willison 博客· simonwillison.net