- Hardware
- 1× NVIDIA GeForce RTX 5090
- Workload
- 1024 → 256
- Runs
- 9 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 1,578.6 | 541.1 ms | 4.75 ms | — | 0 | run_Csl1wzmwaa5WbXqq | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 1,563.2 | 525.7 ms | 5.04 ms | — | 0 | run_wffL2aOab4uxofSZ | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=8 | 1,549.7 | 531.6 ms | 5.02 ms | — | 0 | run_Ua9LfctoL62V1zhZ | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=4 | 1,185.6 | 163.2 ms | 3.42 ms | — | 0 | run_TVD-yEbfI7znExhm | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=2 | 718.4 | 147.3 ms | 3.03 ms | — | 0 | run_tQ08GsRht9cMuN-4 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8 | c=1 | 414.7 | 81.8 ms | 2.56 ms | — | 0 | run_F4Kz3F-9yzkw5N-l | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 415.4 | 81.5 ms | 2.56 ms | — | 0 | run_SEz8DAmLekmNP5AB | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 415.5 | 81.5 ms | 2.56 ms | — | 0 | run_qX0tR01cWoRMSr42 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 414.6 | 81.7 ms | 2.57 ms | — | 0 | run_nmXMEBpM2V5devtc |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 1,578.6
- p95 TTFT / TPOT
- 541.1 ms / 4.75 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 1,563.2
- p95 TTFT / TPOT
- 525.7 ms / 5.04 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=8
- Req/s
- Output tok/s
- 1,549.7
- p95 TTFT / TPOT
- 531.6 ms / 5.02 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=4
- Req/s
- Output tok/s
- 1,185.6
- p95 TTFT / TPOT
- 163.2 ms / 3.42 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=2
- Req/s
- Output tok/s
- 718.4
- p95 TTFT / TPOT
- 147.3 ms / 3.03 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8- Test load
- c=1
- Req/s
- Output tok/s
- 414.7
- p95 TTFT / TPOT
- 81.8 ms / 2.56 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 415.4
- p95 TTFT / TPOT
- 81.5 ms / 2.56 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 415.5
- p95 TTFT / TPOT
- 81.5 ms / 2.56 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 414.6
- p95 TTFT / TPOT
- 81.7 ms / 2.57 ms