Today we’re open-sourcing Lily, the local inference engine we built for hybrid compute in Perplexity Computer.
Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, built so on-device compute doesn’t bottleneck Computer tasks.
Read more: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon