- Hardware
- 1× NVIDIA GeForce RTX 5090
- Draft model
- incoai/Qwen3.8-27B-DFlash2
- Workload
- 1024 → 256
- Runs
- 9 eligible
Efficient frontierOther eligible run
Runs
All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.
| Runtime | Material parameters | Test load | Req/s | Output tok/s | p95 TTFT | p95 TPOT | p95 E2E | Failures | Run |
|---|---|---|---|---|---|---|---|---|---|
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=8 | 1,126.5 | 1,343.9 ms | 3.5 ms | — | 0 | run_aI4FkZvVY0p1KI45 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=4 | 1,127.4 | 225.2 ms | 3.5 ms | — | 0 | run_okGtXqxwISTiqbue | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=4 | 1,125.2 | 224.7 ms | 3.5 ms | — | 0 | run_lCtZaPCT6by0uscg | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=4 | 1,124.1 | 225 ms | 3.71 ms | — | 0 | run_lVl-eHrXjQFFtVVq | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=2 | 668.1 | 136.2 ms | 2.88 ms | — | 0 | run_HdWZbZ-8Ez77f3Uu | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4 | c=1 | 382.5 | 83.5 ms | 2.68 ms | — | 0 | run_mATdf5hwQdeooXoO | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 385.9 | 83.2 ms | 2.66 ms | — | 0 | run_QPTe584j4T2arbz0 | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 385.9 | 83.2 ms | 2.66 ms | — | 0 | run_T7Riy9CpQJph9INq | |
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 | dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1 | c=1 | 386 | 83.1 ms | 2.66 ms | — | 0 | run_kwpo7O9vGCtX25x8 |
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=8
- Req/s
- Output tok/s
- 1,126.5
- p95 TTFT / TPOT
- 1,343.9 ms / 3.5 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=4
- Req/s
- Output tok/s
- 1,127.4
- p95 TTFT / TPOT
- 225.2 ms / 3.5 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=4
- Req/s
- Output tok/s
- 1,125.2
- p95 TTFT / TPOT
- 224.7 ms / 3.5 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=4
- Req/s
- Output tok/s
- 1,124.1
- p95 TTFT / TPOT
- 225 ms / 3.71 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=2
- Req/s
- Output tok/s
- 668.1
- p95 TTFT / TPOT
- 136.2 ms / 2.88 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4- Test load
- c=1
- Req/s
- Output tok/s
- 382.5
- p95 TTFT / TPOT
- 83.5 ms / 2.68 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 385.9
- p95 TTFT / TPOT
- 83.2 ms / 2.66 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 385.9
- p95 TTFT / TPOT
- 83.2 ms / 2.66 ms
- Runtime
- sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
- Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1- Test load
- c=1
- Req/s
- Output tok/s
- 386
- p95 TTFT / TPOT
- 83.1 ms / 2.66 ms