InFeeo
Language

vLLM prefill paired with TileRT decode(github.com)

×
Link preview GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs A high-throughput and memory-efficient inference and serving engine for LLMs - vllm-project/vllm GitHub · github.com
vLLM prefill paired with TileRT decode through vLLM V1's connector interface: a specialized, latency-optimized decode engine that coexists with native vLLM deco

Comments

Log in Log in to comment.

No comments yet.