Artificial Analysis@ArtificialAnlys
60AI 编辑部评分,满分 100

Artificial Analysis 推出 Optima 自定义基准测试平台

2026-08-13 23:54· 34分钟前
AI 导读

Artificial Analysis 发布 Optima 平台,用户可基于自有数据或 HuggingFace 数据集、agent 轨迹构建自定义基准,一键在主流模型上运行并对比性能、速度与成本效率。Optima 支持复用其 GDPval-AA 等成对评判方法,并追踪每任务成本与耗时,帮助用户找到性能相当但成本或时间降低 10 倍的替代模型。平台现已上线。

We're launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis' leading research and platform

Building and running benchmarks is difficult. We have distilled Artificial Analysis' research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task.

We've integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow:

➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or @huggingface, or import agent traces from platforms including @arizeai, @braintrust and @langfuse. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you

➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released

➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set

➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case

Ahead of launch, here are examples questions our beta testers answered with Optima:

➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent?

➤ Which model best matches the writing style of lawyers for my legal agent?

➤ Which model can best identify different elements in my custom image dataset?

Optima is available today. Build your own benchmark at https://artificialanalysis.ai/optima

来源:Artificial Analysis · x.com

Artificial Analysis 推出 Optima 自定义基准测试平台

Artificial Analysis · @ArtificialAnlys · X·2026-08-13 23:54·34分钟前
AI 导读

Artificial Analysis 发布 Optima 平台,用户可基于自有数据或 HuggingFace 数据集、agent 轨迹构建自定义基准,一键在主流模型上运行并对比性能、速度与成本效率。Optima 支持复用其 GDPval-AA 等成对评判方法,并追踪每任务成本与耗时,帮助用户找到性能相当但成本或时间降低 10 倍的替代模型。平台现已上线。

We're launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis' leading research and platform

Building and running benchmarks is difficult. We have distilled Artificial Analysis' research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task.

We've integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow:

➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or @huggingface, or import agent traces from platforms including @arizeai, @braintrust and @langfuse. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you

➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released

➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set

➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case

Ahead of launch, here are examples questions our beta testers answered with Optima:

➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent?

➤ Which model best matches the writing style of lawyers for my legal agent?

➤ Which model can best identify different elements in my custom image dataset?

Optima is available today. Build your own benchmark at https://artificialanalysis.ai/optima

来源:Artificial Analysis· x.com