Five techniques to reach the efficient frontier of LLM inference
Every dollar that you spend on model inference buys you a position on a graph of latency and throughput. On this plot is a curve of optimal configurations, where you’ve squeezed the maximum possible performance from your hardware. That curve, borrowed from portfolio theory in finance, is the efficient frontier. With the assumption that you […]
Five techniques to reach the efficient frontier of LLM inference Read More »





