High-throughput, memory-efficient LLM inference and serving
Throughput king. Cut our inference bill almost in half.
Log in to reply.
Start the thread.