High-throughput, memory-efficient LLM inference and serving
Fast and boring in the best way. PagedAttention still magic.
Log in to reply.
Start the thread.