Prefill-first inference for open models

Starting with real-time voice

80ms TTFT on Qwen3.6-35B-A3B