Serving request latency mathematical notation
Serving request latency is a recurring research-paper notation family. Measures wall-clock time from request arrival to a chosen completion point such as first token or final token.
Serving request latency: Measures wall-clock time from request arrival to a chosen completion point such as first token or final token. Example: Add waiting, prompt processing, and token-generation phases.