Another fantastic evening at our latest Inference, Measured event in San Francisco. We explored how the same model is not always the same product through conversations on serverless inference, Endpoint Accuracy Index, and AA-AgentPerf
Our new Endpoint Accuracy Index reveals another layer. We self-host the released weights as a 100% reference, measure each provider endpoint using the same three evaluations, and score the results relative to that reference. The results range from 73% to 100%.
Quantization, KV-cache compression, and context limits can all affect the quality, even when they don't appear on the pricing page. Visit our website to explore the evaluations and full results.
Thank you to everyone who joined us and contributed to the conversation about measuring the tradeoffs between performance, speed, and cost.