新基准 Search Index 评测 AI 智能体搜索 API 的质量、成本与速度

The Decoder:AI News(RSS)·2026-08-19 02:10·7天前·Matthias Bastian
AI 导读

Artificial Analysis 发布 Search Index 基准,在统一智能体设置下用 GPT-5.6 Luna 评测 Parallel、Exa、Firecrawl 等 7 家搜索 API 提供商。

The Decoder:AI News(RSS)
59AI 编辑部评分,满分 100

新基准 Search Index 评测 AI 智能体搜索 API 的质量、成本与速度

2026-08-19 02:10· 7天前· Matthias Bastian
AI 导读

Artificial Analysis 发布 Search Index 基准,在统一智能体设置下用 GPT-5.6 Luna 评测 Parallel、Exa、Firecrawl 等 7 家搜索 API 提供商。

Image description

Artificial Analysis has released the "Search Index," a benchmark that measures how well search API providers work for AI agents across quality, cost, and speed.

The initial lineup includes Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave. Each one is tested with the same model (GPT-5.6 Luna) in a standardized agent setup. Only the search provider changes. The agent runs on Stirrup, an open-source framework from Artificial Analysis, and gets 25 runs per task to search and pull up web pages.

The index combines three equally weighted benchmarks. DeepSearchQA has 900 research questions, each requiring multiple search queries. A BrowseComp subset tests 200 hard-to-find facts that need multi-step browsing. AA-Omniscience covers 600 questions across six knowledge domains. A tool-free baseline, where the model answers on its own, provides the comparison point.

Without search access, the model scores just 33 points. With search, scores range from 65 to 75. Parallel, Exa, and Firecrawl lead with 75, 74, and 73. | Image: Artificial Analysis

Better search quality also lowers total costs. The model uses fewer tokens when it gets good results up front. With Parallel Search (advanced), token use drops by over 40 percent compared to the Basic version. Per-task search costs go up, but total cost comes in lower ($0.084 vs. $0.11).

Raw speed per query doesn't always mean faster results overall. Parallel Search (turbo) clocks the shortest response time per query (0.51 seconds vs. 1.03 seconds for Basic), but its lower quality (67 vs. 73) forces the agent to run more passes. Total time per task winds up about the same.

Artificial Analysis says Parallel, Firecrawl, and Parallel (turbo) hit the best mix of cost and performance. Other providers can apply to join the benchmark. The full methodology is public.

来源:The Decoder:AI News(RSS)· the-decoder.com