# Perplexity 发布 pplx-embed 模型服务架构，延迟显著低于 vLLM

- 来源：Aravind Srinivas (@AravSrinivas)
- 发布时间：2026-09-05 10:38
- AIHOT 分数：36
- AIHOT 链接：https://aihot.virxact.com/items/cmtns357k0aamroqsfaz57czk
- 原文链接：https://x.com/AravSrinivas/status/2096065348487487832

## AI 摘要

Perplexity 发布博客，介绍如何在 EB 级搜索索引上为 pplx-embed 模型提供 embedding 和 reranker 服务。

## 正文

Join us if you want to work on hard inference and infrastructure engineering problems!

### 引用推文

> Denis Yarats：check out our new blog post on how we serve embeddings and rerankers for our SOTA pplx-embed models over an exabyte-scale search index. embedding models have a ...
